# Playwright CLI vs MCP: give your agent a browser that actually works

> I ran Playwright CLI, Playwright MCP and Chrome DevTools MCP on the same page. The snapshots came out nearly identical; the real cost sits elsewhere.

In July, every browser task in our cloud machines died the same way. The agent reported that its browser tools weren't available, and that was the whole error. Two problems were stacked under it. The Debian Chromium build in our base image crashed with `SIGTRAP` on every launch. And once we fixed that, Playwright still went looking for Chrome at `/opt/google/chrome/chrome`, the hard-coded path its `chrome` channel expects, which didn't exist in the container.

Getting a browser to launch at all is the first half of giving an agent a browser. The second half is the interface, and that's where most of the current advice is louder than it is measured. So I installed all three popular options in a scratch directory and pointed each one at the same page.

**The short answer:** use **Playwright CLI** when a coding agent with a shell uses the browser now and then. Use **Playwright MCP** when the agent has no shell or the browser is the whole job. Use **Chrome DevTools MCP** when you're debugging performance, network or console output in Chrome. On the same page, all three return a snapshot of almost the same size. What separates them is fixed overhead.

## What is Playwright MCP, and what is Playwright CLI?

They're two front doors to the same engine. Playwright MCP is an MCP server: your agent connects to it and gets browser tools such as `browser_navigate`, `browser_snapshot` and `browser_click`, described by JSON schemas that ride along with every model request. Playwright CLI is a `playwright-cli` binary plus an agent skill: your agent runs shell commands like `playwright-cli open` and `playwright-cli click e15`, and reads a `SKILL.md` that teaches it the commands.

The CLI doesn't hide where it comes from. Its own `--help` opens with:

```text
playwright-cli - run playwright mcp commands from terminal
```

Microsoft is also unusually direct about which one it thinks coding agents should use. The [Playwright MCP README](https://github.com/microsoft/playwright-mcp) says CLI invocations "are more token-efficient: they avoid loading large tool schemas and verbose accessibility trees into the model context," and that MCP "remains relevant for specialized agentic loops that benefit from persistent state, rich introspection, and iterative reasoning over page structure."

**Chrome DevTools MCP** comes from a different family. It's the Chrome DevTools team's [MCP server](https://github.com/ChromeDevTools/chrome-devtools-mcp), built on Puppeteer, and it reaches into DevTools itself: performance traces, Lighthouse audits, network requests, console messages and heap snapshots. It can also drive pages, but debugging is what it's for.

If MCP itself is new to you, [what the Model Context Protocol actually is](/blog/what-is-model-context-protocol) covers the client-server model first.

## Playwright CLI vs MCP: what I measured

**The schemas differ by thousands of tokens; the page snapshots barely differ at all.** I ran this on 2026-10-02 on macOS with stable Chrome, using `@playwright/mcp` 0.0.83, `@playwright/cli` 0.1.22 and `chrome-devtools-mcp` 1.10.1.

For the MCP servers, a small script spoke JSON-RPC over stdio, called `tools/list`, and then made real tool calls. Every number is counted with the `o200k_base` tokenizer. Claude's tokenizer produces different absolute numbers, but the ratios hold. The test page was the Hacker News front page, which is a dense table of links and a fair stand-in for a real dashboard.

Here's the CLI side, trimmed only where noted:

```text
$ playwright-cli open https://news.ycombinator.com
### Browser `default` opened with pid 81070.
### Ran Playwright code
await page.goto('https://news.ycombinator.com');
### Page
- Page URL: https://news.ycombinator.com/
- Page Title: Hacker News
### Snapshot
- [Snapshot](.playwright-cli/page-2026-10-02T05-55-43-914Z.yml)

$ playwright-cli find "Show HN"
### Result
Found 1 match for "Show HN":
- table [ref=e3]:
  ...
```

And the results:

| Setup                                  | Tools              | Schema tokens on every request           | HN snapshot      | `find "Show HN"` |
| -------------------------------------- | ------------------ | ---------------------------------------- | ---------------- | ---------------- |
| Playwright MCP (default)               | 25                 | 4,413                                    | 12,301           | 202              |
| Playwright MCP (`vision,pdf,devtools`) | 45                 | 7,032                                    | —                | —                |
| Playwright CLI + skill                 | 0 (uses the shell) | ~22 (skill description), 3,688 once used | 12,302           | 202              |
| Chrome DevTools MCP (default)          | 30                 | 5,914                                    | 13,322           | —                |
| Chrome DevTools MCP (`--slim`)         | 3                  | 223                                      | no snapshot tool | —                |

Three things surprised me.

**The snapshots are the same snapshot.** Playwright CLI returned 12,302 tokens for the page and Playwright MCP returned 12,301. That isn't a coincidence; it's one tool layer behind two interfaces. Much of the "CLI is 4× cheaper" talk compares an agent that dumps the full accessibility tree with one that doesn't. That's a habit, not an interface.

**Playwright MCP already stopped inlining snapshots on navigation.** In 0.0.83, `browser_navigate` came back in 78 tokens: the page title, the code it ran, and a link to a `.yml` file holding the snapshot. Both interfaces also have a `find` that returns the matching subtree. Searching for "Show HN" cost 202 tokens either way, about 1.6% of the full snapshot.

**The real difference is the schema tax.** With Playwright MCP connected, 4,413 tokens of tool definitions go out with every request, whether or not that turn touches the browser. The CLI costs about 22 tokens until the agent actually uses the Playwright CLI skill, then 3,688 for `SKILL.md`. Ten reference files covering tracing, storage state and test generation add another 12,923 tokens, but only when the agent opens them. Over a 40-request coding session that never opens a page, the MCP schema adds up to about 176,000 tokens of definitions. [Prompt caching](/blog/prompt-caching-explained) makes the repeats cheaper, but they still occupy the window.

That gives you the decision rule. If the browser is incidental, the CLI wins by thousands of tokens per request. If the browser is the job, the gap is 3,688 against 4,413 tokens: small enough that other things should decide.

## When Playwright MCP is still the right call

**Pick Playwright MCP when your agent has no shell, or when the browser is the main event.** A chat assistant, a scheduled routine or a hosted agent without terminal access can't run `playwright-cli`, but it can call MCP tools. Typed tool calls are also easier for an orchestrator to log, filter and gate than free-form shell strings.

The schema tax is negotiable, too. Most clients let you expose only a subset of a server's tools. In Octomind it's the `tools` filter on a server entry:

```toml
[[mcp.servers]]
name = "playwright"
type = "stdio"
command = "npx"
args = ["@playwright/mcp@0.0.83", "--headless", "--isolated"]
timeout_seconds = 120
tools = ["browser_navigate", "browser_snapshot", "browser_find", "browser_click", "browser_type", "browser_take_screenshot"]
```

Those six tools measured 1,280 tokens, down from 4,413 for all 25. That covers most read-and-click work. See [MCP tools](/docs/usage/07-mcp-tools) for the full server configuration.

One caveat from the server's own help text: `--allowed-origins` and `--blocked-origins` "_does not_ serve as a security boundary and _does not_ affect redirects." Don't treat an origin list as containment.

## When Chrome DevTools MCP beats both

**Reach for Chrome DevTools MCP when the question is "why is this page slow or broken," not "click through this flow."** Its 30 tools include `performance_start_trace`, `lighthouse_audit`, `list_network_requests`, `list_console_messages` and `take_heapsnapshot`. Playwright doesn't offer that level of introspection. It can also attach to the Chrome you already have open (`--autoConnect`, Chrome 144+), which is handy when the bug only reproduces in your own profile.

It had the biggest snapshot of the three, at 13,322 tokens for the same page, because it gives every text node its own `uid` line. For basic tasks, `--slim` cuts it to three tools (`navigate`, `evaluate` and `screenshot`) at 223 tokens. You lose the accessibility snapshot, though, so the agent works through scripts and screenshots.

Two defaults are worth knowing before you install it on a work machine. Google collects usage statistics unless you pass `--usageStatistics=false` or set `CHROME_DEVTOOLS_MCP_NO_USAGE_STATISTICS`. And performance traces send the page URLs to the CrUX API for field data unless you pass `--performanceCrux=false`. Both are documented in the server's `--help`. Neither is a scandal, but you should choose them on purpose.

So the honest answer to Chrome DevTools MCP vs Playwright MCP is that it's not really a contest. One is a debugger, the other a driver. Plenty of setups should have both, with the debugger loaded only when needed.

## How we give agents a browser in Octomind

**We wrap Playwright MCP in a capability that loads only when a request needs it.** Octomind's browser capability starts `@playwright/mcp@0.0.78 --headless` (24 tools, 4,016 tokens when I measured that version). Developer and browser roles that don't declare it statically pull it in at runtime when your message matches a trigger like "fill out a web form" or "take a screenshot of a webpage". `auto_capabilities` is on by default, and out of the box no LLM sits in that routing step: it's embedding similarity against the trigger phrases.

So the schema tax shows up only in the sessions that use a browser. [Token efficiency](/docs/usage/16-token-efficiency) explains the activation and eviction rules.

A second guard sits on the other end: `mcp_response_tokens_threshold` defaults to 20,000 tokens per tool result. A 12,000-token snapshot fits; a runaway page dump gets cut to the cap instead of flooding the context.

The July failure taught us something the comparison charts skip: in containers, the browser binary is the fragile part, not the protocol. We fixed the missing Chrome path with a symlink to our Chromium. Current Playwright MCP also has `--executable-path`, which I'd reach for first today. And once the browser launches, logged-in sites raise their own problem: in our tests, a valid session driven headlessly got a `403` where the same session in a headful browser got a `200`. [Giving an agent a logged-in browser](/blog/agent-browser-login) covers that handoff.

## Set up Playwright MCP in Claude Code and other agents

**Each option takes one or two commands.** To add Playwright MCP to Claude Code, the README's line is:

```bash
claude mcp add playwright npx @playwright/mcp@latest
```

Install Playwright CLI globally and drop the skill into your workspace:

```bash
npm install -g @playwright/cli@latest
playwright-cli install --skills
```

That second command printed `Skill installed to .claude/skills/playwright-cli` and `Found chrome, will use it as the default browser.` on my machine. The CLI runs headless by default and keeps the profile in memory unless you pass `--persistent`.

To run Chrome DevTools MCP in Claude Code in slim, headless mode:

```bash
claude mcp add chrome-devtools -- npx -y chrome-devtools-mcp@latest --slim --headless
```

The broader version of this decision, when an agent needs a server and when a shell command will do, is in [MCP vs CLI vs Skills](/blog/mcp-vs-cli-vs-skills).

## FAQ

### What is Playwright MCP?

Playwright MCP is Microsoft's Model Context Protocol server for browser automation. It gives an AI agent tools to navigate, click, type, take accessibility snapshots and screenshots in a real browser. It works from the page's accessibility tree rather than pixels, so it doesn't need a vision model. The default build exposes 25 tools.

### Is Playwright CLI better than Playwright MCP?

For coding agents with shell access, usually yes: it adds about 22 tokens of context until it's used, versus 4,413 tokens of tool schemas on every request for Playwright MCP. Page snapshots come out the same size either way. If your agent has no shell, or browsing is its main task, Playwright MCP is a reasonable choice.

### Chrome DevTools MCP vs Playwright MCP: which should I use?

Use Playwright MCP to drive flows: forms, navigation and scraping. Use Chrome DevTools MCP to debug: performance traces, Lighthouse audits, network requests, console errors and memory. Its snapshots ran about 8% larger in my test, and it collects usage statistics by default unless you opt out.

### How do I use Playwright MCP with Claude Code?

Run `claude mcp add playwright npx @playwright/mcp@latest`, then start a new session so the tools load. To keep context small, expose only the tools you need, or use Playwright CLI with its skill instead (`playwright-cli install --skills`).

## The part no table settles

I expected the interface to be the story. It wasn't: the same engine produced the same 12,000-token page either way. The question worth asking before you install anything is whether the browser is a tool your agent picks up occasionally or the job it was hired for. Answer that, and the three options stop competing.
