Animated Octomind icon — purple pixel-art octopus

Every release for four months added a mechanism. This one removed four of them.

Here's the number that caused it. On octobench — 50 real bugs and features harvested from merged pull requests across five languages, graded by each project's own held-out tests — Octomind running GLM-5.3 solved 46 of 50. The same model in opencode solved 43. Good news, and we published it.

The bad news was in the next column. Octomind spent 10.9M non-cache tokens getting there. opencode spent 4.4M for its 43. Codex on gpt-5.6-sol spent 2.8M for 45, at a median of 2.3 minutes per case against our 10.6. One of our runs took 221 minutes and 967 tool calls and still failed.

We were winning on the thing we optimize for and losing badly on the thing we tell users we care about. So we went looking for the weight. Most of it turned out to be machinery we had built, shipped, documented, and never proved.

0.45 is the release where we deleted it. Net −1,061 lines of production code, and +10,404 lines of tests to make sure the deletions were safe. The 0.45 line shipped across August 22 — 0.45.0, then 0.45.1 hours later, then 0.45.2 to pick up a dependency fix. Everything below refers to the line as one release, and 0.45.2 is the one to install.


How we got here

It's worth being honest about the path, because the path is the lesson.

0.42.1 taught the agent's summary to prove its claims. 0.43.0 made compression reversible so a wrong summary could be checked against the original. 0.44 rebuilt the supervisor around evidence: a persisted verification policy, a verify gate that demands per-condition proof, planning taken out of the model's hands.

Each of those was a response to a real failure we had watched happen. And each one added a detector, a threshold, a config key.

By 0.44.2 the supervisor ran six detectors plus a post-compression verification pass. Two of the detectors were cheap counters over information novelty — loop and no-progress. Two were opt-in advisories about wasted work. The remaining three, and the pass, were cleverer:

  • Truncation — count consecutive truncated tool results, on the theory that a model re-querying without narrowing is thrashing.
  • Dedup — count consecutive deduplicated results, same theory.
  • Distraction — embed every sizable tool result, score its cosine against a rolling centroid of recent results, and flag drift when it falls under drift_floor.
  • Compaction fidelity — after every compression, run a second model over the surviving context to check that no binding requirement was lost, and re-inject anything that was.

Every one of those is a defensible idea. Two of them I still think are good ideas. That's exactly the problem: defensible is not measured, and a defensible idea you never measure is how a codebase gets fat while everyone nods along.

What the benchmark actually said

We rebuilt octobench in July specifically so we could stop arguing from intuition. It runs each case at the parent commit of a merged fix, with the network sealed and the upstream repo unreachable, and grades with tests the agent has never seen. The full per-case table is pinned at the commit of this run.

When we went through the traces looking for where 10.9M tokens went, the four clever detectors had a common shape:

Truncation and dedup counted symptoms, not causes. A truncated result in a row is not evidence of thrashing — often it's a large file being read in the obvious way. Both counters fired on healthy work and stayed silent on the runs that actually went sideways, where the failure was never repetition. It was self-referential verification: the agent checking its own belief about the code instead of the code.

Distraction cost real money for a signal we could not tune. It was off by default because we could never find a drift_floor that separated "moved to another subsystem" from "lost the plot" — the docs literally told you to tune it from debug logs before relying on it. A detector that ships off, that nobody can configure without a log-reading session, and that costs one embedding per sizable tool result when enabled, is not a feature. It's an unpaid invoice.

Compaction fidelity was the expensive one. A model call after every compression, on long sessions, forever. It was insurance against a failure mode that 0.43.0's reversible compression and 0.44's signed goal recitation had already closed structurally. We were paying a per-compression tax to re-check a guarantee two other subsystems were already enforcing.

The benchmark table was regenerated on August 20. The detectors came out on August 21.

What 0.45 deletes

Four config keys are gone from [supervisor.detectors]:

toml
# removed in 0.45 — leftovers in your config are ignored, no migration needed
truncation_threshold = 2
dedup_threshold      = 2
distraction_threshold = 0
drift_floor          = 0.7

Along with them: the entire centroid-embedding drift path, the compaction fidelity module, the fidelity call tracking in stats, and the branch of the steering cascade that existed to arbitrate between six competing signals. The steer priority is now two cases — Loop and NoProgress — which is short enough to hold in your head, which is the point.

What survives is what earned it:

DetectorStatusWhy it stayed
LoopOn (loop_threshold = 3)Identical result three times running is unambiguous. Free. No model needed.
No-progressOn (no_progress_window = 5)Zero-novelty churn, fused with the agent's own self-report. The conflict is the signal.
SequentialOpt-in (sequential_threshold = 0)Real token waste when it fires; cheap; you choose.
Re-readOpt-in (reread_threshold = 0)Same deal. Counter resets on mutation, so edit-verify loops never trip it.

The verify gate is untouched — it is the part of the supervisor the benchmark did vindicate, and it got better this release rather than smaller. More on that below.

The 1,260 lines nobody was running

While we were in there, we found src/session/response.rs: a 1,260-line response-processing orchestrator with tool execution, cancellation, event emission, and nested parameter formatting.

It wasn't declared in mod.rs. Nothing imported it. The compiler had never seen it. It had been sitting in the tree being read by humans and by every agent that indexed the repository, costing attention on every code search, doing nothing.

That's the most honest artifact in this release. Dead code does not announce itself — it looks exactly like live code until you check.

There's a smaller one from the same week, and it's more embarrassing. On August 21 at 19:01 we shipped a pinned ledger of satisfied requests into the attention module: 126 lines, tested, in the changelog. Five hours later we deleted it, because nothing read it. Shipping is not the same as being used, and the gap between the two is one grep.

The rule we came out with: a mechanism has to earn its line count against a measurement, not against a story about when it would help. If we cannot name the run where it changed the outcome, it goes.

What 0.45 adds

Now the other half — because a diet is not a product, and the additions this cycle are the kind you notice in the first ten minutes.

@file mentions

Type @ and a path anywhere in your message and Octomind reads that file into the turn:

text
> why does @src/config/merge.rs panic on an unknown role?

The file arrives as a structured <content path="..." lines="1:230"> block appended to your message, with line numbers, so the model can cite it back precisely. The @mention stays where you typed it, so the model knows where in your sentence the file was referenced.

It fails quietly on purpose. A nonexistent path, a directory, a binary file, an email address — anything that isn't a readable text file stays plain text. Files longer than 10,000 lines are clamped rather than refused.

Tab completion comes with it. In 0.45.0 completion was fuzzy over the indexed working directory. 0.45.1 adds real filesystem walking for anything path-shaped — a query starting with ./, / or ~/ gets an actual directory listing, tilde-expanded, case-insensitive on the prefix, directories first with a trailing slash. So @~/Work/ completes, @../sibling-repo/ completes, and the index no longer decides what you are allowed to reference.

MCP tools that show their work

A long-running MCP tool used to be a black box: the spinner spun, the tool's progress notifications went to stderr where they either scrolled past or got swallowed.

0.45.1 routes them into the live spinner phase instead:

text
⠹ [octofs] command still running

Under the hood each MCP progress token is bound to the tool call that owns it, so the notification is attributed rather than floating. The same mapping reaches the other front doors: notifications become ToolCallUpdate events over ACP, and the WebSocket protocol payload now carries tool_id, so an editor or a panel can attach progress to the right row.

This pairs with 0.44's idle-deadline timeouts: a tool that reports progress both stays alive and now tells you it's alive.

The verify gate stops guessing

The gate kept its evidence-conditions design from 0.44 and grew three things that came directly out of watching it fail on benchmark cases:

  • Readback rounds. Instead of inferring what happened from a summary, the gate can now ask for a specific tool action's output by sequence ID. The evidence ledger is keyed, so "show me what command 14 actually printed" is a lookup, not a reconstruction.
  • Stall detection. A verification loop that stops converging is now detected and reported as gate_stall, and blocking advisories are bypassed rather than repeated. The gate gives up honestly instead of grinding.
  • Gaps must close. A reported gap now has to carry an actionable closing observation — what would have to be observed for this to be settled. Findings without one are filtered out. Shape verdicts became three-valued (found / not found / unknown), because "I couldn't tell" is different information from "it isn't there", and collapsing them was costing us honest failures.

Compression stops eating your request

Three fixes here, all from real sessions:

Your request survives the fold. A compression that ran mid-task could drop the bridge carrying your actual request, leaving the agent to continue from a paraphrase. The tail is now checked for a real user task before any continuation wrapper is skipped, with regression tests for both the mid-task and follow-up shapes.

The live exchange survives too. The current assistant step and its tool results are no longer eligible for folding. Compressing the work in flight is how an agent forgets what it is holding.

Compression cannot grow the context. File-context injections inside summaries now have a token budget and a line limit, and a net token increase after a fold is logged as an error rather than absorbed silently. A compression that makes the context bigger is a bug, and now it says so.

Related: tool outputs count as valid grounds for citations, which kills a class of false fabrication flags where the agent quoted a command's real output and got told it made it up.

Smaller things that will bite someone

  • OCTOMIND_DATA_DIR overrides the state directory. See environment variables.
  • Config migration logs moved to stderr. In ACP and MCP stdio modes, stdout carries JSON-RPC — a startup config upgrade printing to stdout corrupted the stream. If you run Octomind as an editor backend, this one was silently breaking your first session after an upgrade.
  • Unknown roles no longer panic. Role config merging moved into its own module with real fallbacks, per-role MCP server and tool filtering, and auto-bind server support. See roles.
  • Queued messages survive re-initialization. Inbox init is idempotent, so a /send sitting in the queue is not overwritten.

The tests are the actual headline

Deleting 2,868 lines of production code from an agent runtime is only responsible if you can prove what still works. So the same release added 10,404 lines of tests, and deleted none.

Not unit tests around the deleted code — end-to-end tests around the product:

  • Interactive PTY tests that drive a real terminal session through a Python driver, synchronized on a prompt-ready marker.
  • CLI run e2e covering the non-interactive path, over 1,000 lines of it.
  • ACP e2e, with agent stderr captured so a timeout tells you why.
  • WebSocket e2e against the live protocol.
  • A Python MCP stub server so handshake, tool listing and connection lifecycle are tested against something that behaves like a real peer instead of a mock.

Plus source-based coverage reporting in CI, so the next deletion argument can be settled with a number.

Upgrading

bash
# Homebrew (macOS / Linux)
brew install muvon/tap/octomind

# or from crates.io
cargo install octomind

Prebuilt binaries for every platform are on the releases page.

Two things to know:

  1. No config migration. The config version did not change. The four removed [supervisor.detectors] keys are simply ignored if they are still in your file — you can delete them at your leisure, and nothing breaks if you never do.
  2. If you had distraction_threshold enabled, you lose that advisory and you stop paying for an embedding per sizable tool result. If it was catching something real for you, we want the trace — that's exactly the evidence that would have kept it.
  3. 0.45.2 is dependency-only — octolib 0.34.2, no agent behaviour change. It carries two fixes worth knowing about. Reasoning tokens now count toward a turn's reported cost across nine providers: they were always billed at the output rate, but providers report them apart from output_tokens, so the number Octomind showed you was under the real one on thinking models. And tool schemas are normalized to scalar types — "type": [T, "null"] collapses to T — because some serving stacks read type as a plain string, fall through to the string branch, and constrain the model to emit that argument as text: "0.6" instead of 0.6, which the MCP server on the other end then refuses to deserialize. Measured on Alibaba Model Studio's Qwen path and on CoreWeave via OpenRouter; the same weights elsewhere were fine, which is why it's a per-serving-stack defect you cannot see from the model id.

The model hub is open to everyone — octomind login, or /login from inside a session, and you are routing.


There is a through-line here, and it is not the one I expected to write four releases ago. 0.42.1 made the summary prove its claims. 0.43.0 kept the originals within reach. 0.44 made "done" a verdict the evidence has to earn.

0.45 points the same standard at us. Every mechanism in an agent runtime is a claim: this makes the agent better. Four of ours could not produce the evidence, so they are gone — and the agent got cheaper without getting worse.

The most valuable thing a benchmark gives you is not a number to put in a headline. It is permission to delete your own work.