What we tested
The Model Context Protocol lets any compatible client, such as a desktop assistant, an IDE or a coding agent, call tools exposed by a small server program. The claim is that you write an integration once and every client can use it. We wanted to know what that costs in practice, for the most common real need: let an agent search my own notes.
The setup:
- A server exposing one tool,
search_notes(query, limit), doing case-insensitive full-text search over a folder of Markdown. - A corpus of 221 real README files (2.1 MB) taken from an npm dependency tree. Real prose, real noise.
- A benchmark client that spawns the server over stdio, lists tools, and makes 60 calls (six queries × ten), repeated across 10 fresh server processes.
Results
| Measure | Result |
|---|---|
| Cold start (spawn → handshake complete) | 210–262 ms, mean 238 ms |
| Tool call latency, median | 3.2–4.3 ms per run, mean 3.5 ms |
| Tool call latency, p95 | 5.8–9.9 ms per run, mean 7.7 ms |
Invalid input (query: "a") |
Returned as a tool error, not a crash |
Cold start includes reading and indexing all 221 files into memory. Per-call latency includes JSON-RPC over stdio and the search itself. Neither is noticeable next to a model call, which takes hundreds of milliseconds to many seconds.
The code
The whole server:
import { McpServer } from "@modelcontextprotocol/sdk/server/mcp.js";
import { StdioServerTransport } from "@modelcontextprotocol/sdk/server/stdio.js";
import { readdir, readFile } from "node:fs/promises";
import { join } from "node:path";
import { z } from "zod";
const dir = process.argv[2];
const notes = await Promise.all(
(await readdir(dir)).filter((f) => f.endsWith(".md"))
.map(async (f) => ({ file: f, text: await readFile(join(dir, f), "utf8") }))
);
const server = new McpServer({ name: "notes", version: "0.1.0" });
server.registerTool(
"search_notes",
{
description: "Full-text search over local Markdown notes. Returns matching file names with a short excerpt.",
inputSchema: { query: z.string().min(2), limit: z.number().int().max(20).default(5) },
},
async ({ query, limit }) => {
const q = query.toLowerCase();
const hits = notes
.map((n) => ({ ...n, i: n.text.toLowerCase().indexOf(q) }))
.filter((n) => n.i >= 0)
.slice(0, limit)
.map((n) => `${n.file}: …${n.text.slice(Math.max(0, n.i - 60), n.i + 100).replace(/\s+/g, " ")}…`);
return { content: [{ type: "text", text: hits.join("\n") || "No matches." }] };
}
);
await server.connect(new StdioServerTransport());
To use it from a client, register the command:
{ "mcpServers": { "notes": { "command": "node", "args": ["server.mjs", "/path/to/notes"] } } }
What surprised us
- The API most tutorials show is deprecated. In SDK 1.32,
server.tool(...)is marked deprecated in favour ofserver.registerTool(name, config, handler). Older examples still work but will drift. - Validation errors go back to the model, not up the stack. A query that fails the schema returns a result with
isError: trueand a readable message. That is the right design: the model can correct itself. - Schema details leak into what the model sees.
z.number().int()without a lower bound produced"minimum": -9007199254740991in the advertised JSON Schema. Harmless, but it is noise in the model’s context. Add explicit bounds such as.min(1).
Verdict
Adopt. The protocol cost is a few milliseconds and a few dozen lines. The hard parts are elsewhere:
- Tool design. One sharp tool with a clear description beats five overlapping ones.
- Permissions. Everything the server can do, the agent can do. Read-only first.
- Untrusted content. Whatever the tool returns enters the model’s context. Treat notes from outside sources as potential prompt injection.
What this test does not tell you: how good full-text search is compared with embeddings for your notes, or how well a given model chooses when to call the tool. Those are next.