In July, every browser task in our cloud machines died the same way. The agent reported that its browser tools weren't available, and that was the whole error. Two problems were stacked under it. The Debian Chromium build in our base image crashed with SIGTRAP on every launch. And once we fixed that, Playwright still went looking for Chrome at /opt/google/chrome/chrome, the hard-coded path its chrome channel expects, which didn't exist in the container.

Getting a browser to launch at all is the first half of giving an agent a browser. The second half is the interface, and that's where most of the current advice is louder than it is measured. So I installed all three popular options in a scratch directory and pointed each one at the same page.

The short answer: use Playwright CLI when a coding agent with a shell uses the browser now and then. Use Playwright MCP when the agent has no shell or the browser is the whole job. Use Chrome DevTools MCP when you're debugging performance, network or console output in Chrome. On the same page, all three return a snapshot of almost the same size. What separates them is fixed overhead.

What is Playwright MCP, and what is Playwright CLI?

They're two front doors to the same engine. Playwright MCP is an MCP server: your agent connects to it and gets browser tools such as browser_navigate, browser_snapshot and browser_click, described by JSON schemas that ride along with every model request. Playwright CLI is a playwright-cli binary plus an agent skill: your agent runs shell commands like playwright-cli open and playwright-cli click e15, and reads a SKILL.md that teaches it the commands.

The CLI doesn't hide where it comes from. Its own --help opens with:

text
playwright-cli - run playwright mcp commands from terminal

Microsoft is also unusually direct about which one it thinks coding agents should use. The Playwright MCP README says CLI invocations "are more token-efficient: they avoid loading large tool schemas and verbose accessibility trees into the model context," and that MCP "remains relevant for specialized agentic loops that benefit from persistent state, rich introspection, and iterative reasoning over page structure."

Chrome DevTools MCP comes from a different family. It's the Chrome DevTools team's MCP server, built on Puppeteer, and it reaches into DevTools itself: performance traces, Lighthouse audits, network requests, console messages and heap snapshots. It can also drive pages, but debugging is what it's for.

If MCP itself is new to you, what the Model Context Protocol actually is covers the client-server model first.

Playwright CLI vs MCP: what I measured

The schemas differ by thousands of tokens; the page snapshots barely differ at all. I ran this on 2026-10-02 on macOS with stable Chrome, using @playwright/mcp 0.0.83, @playwright/cli 0.1.22 and chrome-devtools-mcp 1.10.1.

For the MCP servers, a small script spoke JSON-RPC over stdio, called tools/list, and then made real tool calls. Every number is counted with the o200k_base tokenizer. Claude's tokenizer produces different absolute numbers, but the ratios hold. The test page was the Hacker News front page, which is a dense table of links and a fair stand-in for a real dashboard.

Here's the CLI side, trimmed only where noted:

text
$ playwright-cli open https://news.ycombinator.com
### Browser `default` opened with pid 81070.
### Ran Playwright code
await page.goto('https://news.ycombinator.com');
### Page
- Page URL: https://news.ycombinator.com/
- Page Title: Hacker News
### Snapshot
- [Snapshot](.playwright-cli/page-2026-10-02T05-55-43-914Z.yml)

$ playwright-cli find "Show HN"
### Result
Found 1 match for "Show HN":
- table [ref=e3]:
  ...

And the results:

SetupToolsSchema tokens on every requestHN snapshotfind "Show HN"
Playwright MCP (default)254,41312,301202
Playwright MCP (vision,pdf,devtools)457,032——
Playwright CLI + skill0 (uses the shell)~22 (skill description), 3,688 once used12,302202
Chrome DevTools MCP (default)305,91413,322—
Chrome DevTools MCP (--slim)3223no snapshot tool—

Three things surprised me.

The snapshots are the same snapshot. Playwright CLI returned 12,302 tokens for the page and Playwright MCP returned 12,301. That isn't a coincidence; it's one tool layer behind two interfaces. Much of the "CLI is 4× cheaper" talk compares an agent that dumps the full accessibility tree with one that doesn't. That's a habit, not an interface.

Playwright MCP already stopped inlining snapshots on navigation. In 0.0.83, browser_navigate came back in 78 tokens: the page title, the code it ran, and a link to a .yml file holding the snapshot. Both interfaces also have a find that returns the matching subtree. Searching for "Show HN" cost 202 tokens either way, about 1.6% of the full snapshot.

The real difference is the schema tax. With Playwright MCP connected, 4,413 tokens of tool definitions go out with every request, whether or not that turn touches the browser. The CLI costs about 22 tokens until the agent actually uses the Playwright CLI skill, then 3,688 for SKILL.md. Ten reference files covering tracing, storage state and test generation add another 12,923 tokens, but only when the agent opens them. Over a 40-request coding session that never opens a page, the MCP schema adds up to about 176,000 tokens of definitions. Prompt caching makes the repeats cheaper, but they still occupy the window.

That gives you the decision rule. If the browser is incidental, the CLI wins by thousands of tokens per request. If the browser is the job, the gap is 3,688 against 4,413 tokens: small enough that other things should decide.

When Playwright MCP is still the right call

Pick Playwright MCP when your agent has no shell, or when the browser is the main event. A chat assistant, a scheduled routine or a hosted agent without terminal access can't run playwright-cli, but it can call MCP tools. Typed tool calls are also easier for an orchestrator to log, filter and gate than free-form shell strings.

The schema tax is negotiable, too. Most clients let you expose only a subset of a server's tools. In Octomind it's the tools filter on a server entry:

toml
[[mcp.servers]]
name = "playwright"
type = "stdio"
command = "npx"
args = ["@playwright/[email protected]", "--headless", "--isolated"]
timeout_seconds = 120
tools = ["browser_navigate", "browser_snapshot", "browser_find", "browser_click", "browser_type", "browser_take_screenshot"]

Those six tools measured 1,280 tokens, down from 4,413 for all 25. That covers most read-and-click work. See MCP tools for the full server configuration.

One caveat from the server's own help text: --allowed-origins and --blocked-origins "does not serve as a security boundary and does not affect redirects." Don't treat an origin list as containment.

When Chrome DevTools MCP beats both

Reach for Chrome DevTools MCP when the question is "why is this page slow or broken," not "click through this flow." Its 30 tools include performance_start_trace, lighthouse_audit, list_network_requests, list_console_messages and take_heapsnapshot. Playwright doesn't offer that level of introspection. It can also attach to the Chrome you already have open (--autoConnect, Chrome 144+), which is handy when the bug only reproduces in your own profile.

It had the biggest snapshot of the three, at 13,322 tokens for the same page, because it gives every text node its own uid line. For basic tasks, --slim cuts it to three tools (navigate, evaluate and screenshot) at 223 tokens. You lose the accessibility snapshot, though, so the agent works through scripts and screenshots.

Two defaults are worth knowing before you install it on a work machine. Google collects usage statistics unless you pass --usageStatistics=false or set CHROME_DEVTOOLS_MCP_NO_USAGE_STATISTICS. And performance traces send the page URLs to the CrUX API for field data unless you pass --performanceCrux=false. Both are documented in the server's --help. Neither is a scandal, but you should choose them on purpose.

So the honest answer to Chrome DevTools MCP vs Playwright MCP is that it's not really a contest. One is a debugger, the other a driver. Plenty of setups should have both, with the debugger loaded only when needed.

How we give agents a browser in Octomind

We wrap Playwright MCP in a capability that loads only when a request needs it. Octomind's browser capability starts @playwright/[email protected] --headless (24 tools, 4,016 tokens when I measured that version). Developer and browser roles that don't declare it statically pull it in at runtime when your message matches a trigger like "fill out a web form" or "take a screenshot of a webpage". auto_capabilities is on by default, and out of the box no LLM sits in that routing step: it's embedding similarity against the trigger phrases.

So the schema tax shows up only in the sessions that use a browser. Token efficiency explains the activation and eviction rules.

A second guard sits on the other end: mcp_response_tokens_threshold defaults to 20,000 tokens per tool result. A 12,000-token snapshot fits; a runaway page dump gets cut to the cap instead of flooding the context.

The July failure taught us something the comparison charts skip: in containers, the browser binary is the fragile part, not the protocol. We fixed the missing Chrome path with a symlink to our Chromium. Current Playwright MCP also has --executable-path, which I'd reach for first today. And once the browser launches, logged-in sites raise their own problem: in our tests, a valid session driven headlessly got a 403 where the same session in a headful browser got a 200. Giving an agent a logged-in browser covers that handoff.

Set up Playwright MCP in Claude Code and other agents

Each option takes one or two commands. To add Playwright MCP to Claude Code, the README's line is:

bash
claude mcp add playwright npx @playwright/mcp@latest

Install Playwright CLI globally and drop the skill into your workspace:

bash
npm install -g @playwright/cli@latest
playwright-cli install --skills

That second command printed Skill installed to .claude/skills/playwright-cli and Found chrome, will use it as the default browser. on my machine. The CLI runs headless by default and keeps the profile in memory unless you pass --persistent.

To run Chrome DevTools MCP in Claude Code in slim, headless mode:

bash
claude mcp add chrome-devtools -- npx -y chrome-devtools-mcp@latest --slim --headless

The broader version of this decision, when an agent needs a server and when a shell command will do, is in MCP vs CLI vs Skills.

FAQ

What is Playwright MCP?

Playwright MCP is Microsoft's Model Context Protocol server for browser automation. It gives an AI agent tools to navigate, click, type, take accessibility snapshots and screenshots in a real browser. It works from the page's accessibility tree rather than pixels, so it doesn't need a vision model. The default build exposes 25 tools.

Is Playwright CLI better than Playwright MCP?

For coding agents with shell access, usually yes: it adds about 22 tokens of context until it's used, versus 4,413 tokens of tool schemas on every request for Playwright MCP. Page snapshots come out the same size either way. If your agent has no shell, or browsing is its main task, Playwright MCP is a reasonable choice.

Chrome DevTools MCP vs Playwright MCP: which should I use?

Use Playwright MCP to drive flows: forms, navigation and scraping. Use Chrome DevTools MCP to debug: performance traces, Lighthouse audits, network requests, console errors and memory. Its snapshots ran about 8% larger in my test, and it collects usage statistics by default unless you opt out.

How do I use Playwright MCP with Claude Code?

Run claude mcp add playwright npx @playwright/mcp@latest, then start a new session so the tools load. To keep context small, expose only the tools you need, or use Playwright CLI with its skill instead (playwright-cli install --skills).

The part no table settles

I expected the interface to be the story. It wasn't: the same engine produced the same 12,000-token page either way. The question worth asking before you install anything is whether the browser is a tool your agent picks up occasionally or the job it was hired for. Answer that, and the three options stop competing.