Earlier this month I piped tools/list out of our own three MCP servers and counted. Octofs, octocode and octobrain together expose 14 tools, and their schemas come to 7,553 tokens. That's what an agent pays before you type a word, on every turn, for tools we wrote ourselves and tuned for exactly this problem.

So I understand why "MCP vs CLI" became an argument this year. I also think most of the argument is aimed at the wrong thing.

Here's the short answer. Use a CLI when the model already knows the tool and the job is a stateless command with text output. Use an MCP server when the tool holds state, needs credentials you don't want in a shell, or has no CLI the model has seen. Use a skill to teach the agent when and how. Good setups use all three.

The rest of this post is the evidence: what our servers cost, what the famous 32x benchmark actually measured, and how we load tools in Octomind so MCP doesn't eat the context window.

What our own MCP servers cost before the first prompt

Each MCP server costs its full schema list in tokens, on every request, whether the agent touches it or not. I measured ours on October 2 by speaking MCP over stdio to the shipped binaries (octofs 0.16.0, octocode 0.26.4, octobrain 0.14.4): initialize, then tools/list, then count the compact JSON.

ServerToolsSchema bytesTokens
octofs69,6162,383
octocode410,1342,455
octobrain411,6472,715
Total1431,3977,553

Tokens are counted with OpenAI's o200k_base tokenizer through tiktoken. Claude's tokenizer will land on different numbers, so treat these as close estimates rather than billing figures.

The per-tool spread is the interesting part. Octocode's structural_search alone is 1,326 tokens, because it documents ast-grep pattern syntax and names the traps models fall into with it. Octobrain's memorize is 1,089. Those descriptions are long on purpose, and I'd rather pay for them than for a failed call.

Now the CLI side of the ledger, measured the same way:

text
octocode search --help     1,248 bytes    302 tokens
octocode --help            1,609 bytes    330 tokens
gh --help                  2,716 bytes    610 tokens
git log -h                   720 bytes    194 tokens

The big difference isn't size. It's timing. A schema is paid up front on every turn. Help text is paid once, only when the agent decides to read it, and for git or gh the agent usually doesn't, because those tools are all over the model's training data.

Is MCP dead? What the 32x benchmark actually measured

MCP isn't dead. Eager schema loading is what deserves the criticism. The number everyone quotes comes from Scalekit's March 11 benchmark: Claude Sonnet 4, five read-only GitHub tasks, 75 runs. On "what language is this repo written in", the gh CLI agent used a median 1,365 tokens and the GitHub MCP agent used 44,026. That's the 32x. The other four tasks came in at 20x, 9x, 7x and 4x.

Read the method and the cause is plain. GitHub's server exposed 43 tools, and every one of those schemas went into every request while the agent used one or two. The MCP agent also succeeded in only 18 of 25 runs, and the failures were TCP timeouts to GitHub's remote server. That's network plumbing, not the protocol.

Two details from the same study get quoted less:

  • A skill made the CLI better, not cheaper. Scalekit added an ~800-token document of gh tips. It raised token use on four of the five tasks, but cut tool calls and latency by a third against the plain CLI.
  • Their recommendation is not "drop MCP". They suggest CLI plus skills for developer tools, and MCP behind a gateway for multi-tenant products that need per-user OAuth, tenant isolation and an audit trail.

Anthropic's engineering team made the same point from the other direction in November 2025. In Code execution with MCP, the agent calls MCP servers from code it writes instead of loading every tool definition up front, and their example drops from 150,000 tokens to 2,000.

Both results say the same thing: the cost is in how tools get loaded. If your client dumps every schema from every server into every turn, you'll get the 32x no matter which protocol carries the bytes. I covered the accuracy side of this in what MCP is and the problem of too many tools.

MCP vs CLI: when a shell command is enough

A shell command is enough when the model already knows the tool, the call is stateless, and the output is text it can read. git, gh, cargo, kubectl, rg, jq: there's no reason to wrap these in a server. The model has seen millions of examples of every flag.

We follow that rule in our own taps. The programming-rust skill doesn't bring an MCP server at all. Its capability declares one dependency, the Rust toolchain, and the agent runs cargo through the shell like a person would. Octomind's shell capability exposes exactly one tool from octofs, octofs:shell, and that schema is 355 tokens. One tool, every CLI on the machine.

The CLI starts to lose when:

  • the output is huge or unstructured, and the agent burns context paging through it;
  • the command needs a secret, and now the secret is sitting in argv, in the environment, or in the transcript;
  • the tool is your internal script that no model has ever seen, so the agent has to read --help every session and guess anyway.

When your agent needs a server

Reach for an MCP server when something has to outlive a single call: an index, a loaded model, an open connection, a logged-in browser. Servers also earn their place when credentials should stay out of the shell, or when results need a shape the model shouldn't have to re-parse.

Octocode is the clearest case in our stack. Semantic search runs against an index of the codebase, and the server keeps that index open and its in-memory code graph built between calls. A CLI would reopen the index and rebuild the graph on every invocation.

I timed the cheaper version of this on octobrain: the same read-only memory search, five times each way.

text
fresh `octobrain memory remember` per call:   1,368  1,432  1,375  1,439  1,522 ms
one long-lived `octobrain mcp` process:       1,289    993    991  1,073  1,006 ms

The first MCP call pays the warm-up. After that, each search is about 400 ms faster, because the process and everything it loaded stay up. An agent that checks memory dozens of times a session feels that.

Then there's the contract. Octofs could have been a set of shell scripts. We made it a server because the interface is the product: edits that return a diff instead of forcing a re-read, fuzzy matching that repairs whitespace instead of rejecting the edit, and line IDs that carry a content hash, so a stale edit fails loudly instead of landing on the wrong line. The story of why we built octofs has the failure data that pushed us there. You can't get that from a CLI unless you rebuild the same contract on top of it.

MCP vs skills is the wrong fight

Skills and MCP servers do different jobs. A skill is knowledge: markdown instructions about when to do something and how your team does it. An MCP server is a capability: typed tools the agent can call. "Claude skills vs MCP" makes about as much sense as "documentation vs API".

Skills have their own context problem, though, and it's bigger than most people expect. Our public tap ships 115 skills. Measured with the same tokenizer:

  • all 115 full SKILL.md files: 304,534 tokens;
  • a one-line name and description for each: 8,971 tokens;
  • the median skill: 2,601 tokens.

You can't preload the bodies, and even the one-line index is more than our three MCP servers' schemas combined. So in Octomind we don't put the list in the prompt. Skills activate on declarative rules checked in-process on every message: file(Cargo.toml), content(rust), env(CI). For everything the rules don't catch, the agent has one skill tool (264 tokens of schema) to list, use and forget skills, 20 per page.

Here's the part that ends the "MCP vs skill" debate in our runtime: a skill can bring its MCP servers with it. A SKILL.md can declare capabilities: in its frontmatter. When the skill activates, Octomind resolves those capabilities to MCP servers and starts them. When the skill is forgotten, the servers are offloaded, with reference counting so a server two skills share stays up until both are gone. The skill decides when; the server does the work. The skills documentation covers the frontmatter.

How we keep MCP from eating the context window

Load servers on demand, cap how many are live, and keep the router out of the model. That's the whole recipe, and here's how it looks in Octomind's source.

  • Capabilities auto-activate per message. Each capability carries hand-written trigger phrases. On every new user message, a small local embedding model (30M parameters, CPU-only) compares the intent against those triggers. On a confident match with a clear margin over the runner-up, the capability's MCP servers are registered right then. No LLM routing turn, no extra tool call. auto_capabilities = true is the default.
  • At most four capabilities are live. MAX_ACTIVE_CAPS is 4. Activating a fifth evicts the least recently used one. With the always-on tools plus four or five per capability, that caps an agent at roughly 35–40 tools instead of every tool from every server.
  • A capability exposes a subset of a server. The shell capability lists one allowed tool from octofs. Octocode is split three ways: codesearch-semantic exposes only semantic_search, so the 1,326-token structural_search schema stays out until codesearch-structural is needed.
  • There's a manual fallback. If auto-activation abstains, the agent can call the capability tool to list, discover and enable what it needs.

That's how we can ship three MCP servers with 7,553 tokens of schemas and still not pay 7,553 tokens on a turn that only needs a shell. The MCP tools documentation shows how servers are configured.

A practical decision table

Here's how I decide, using the numbers above:

CLI through the shellMCP serverSkill
What the agent getsOne shell tool, plus everything it already knows about git, ghTyped tools with schemasInstructions: when, how, and what to avoid
Fixed context cost355 tokens for our shell tool; --help only on demandEvery exposed schema, every turn (2,383–2,715 per server)Zero until activated; median 2,601 tokens once active
Wins whenThe tool is famous, stateless, and prints textState, credentials, structured results, unfamiliar toolsThe agent knows the tools but not your way of using them
Fails whenOutput floods the context; secrets land in argv or envSchemas pile up; remote servers time outIt never activates, or every skill gets preloaded

FAQ

Is MCP dead?

No. What got declared dead this year was loading every tool schema into every request. Scalekit's 32x gap came from 43 GitHub tool definitions in context when the agent needed one or two. Load servers on demand and cap how many are active, and MCP stays the right home for stateful tools and anything that needs credentials.

MCP vs CLI: which uses fewer tokens?

For a single call to a tool the model already knows, the CLI. Scalekit measured 4x to 32x on GitHub tasks. The gap is fixed schema cost: our three servers carry 7,553 tokens of schemas, while octocode search --help is 302 tokens and gets read only when needed. With on-demand loading, the difference shrinks to the servers actually in use.

Agent skills vs MCP: what's the difference?

A skill is knowledge: a SKILL.md file of instructions that enters the context when it's relevant. An MCP server is capability: tools the agent can call. They work together. In Octomind, a skill can declare capabilities, and activating the skill starts the MCP servers behind them.

Should I turn my MCP server into a CLI?

If the tool is stateless and wraps something scriptable, ship both. Octocode and octobrain each have a full CLI next to their mcp command, so people use the CLI in scripts and agents use the server. If you're starting from scratch, our custom MCP server tutorial walks through the server side.

Count before you argue

The MCP vs CLI fight mostly happens between people who haven't measured their own setup. Run tools/list against every server you load, count it, and look at how many of those tools a normal turn actually uses. That number tells you more than any benchmark headline.