The problem with every AI coding tool: they forget.

You've been there. Three hours into a debugging session with Claude Code, you've explored five dead ends, finally found the root cause, and you're halfway through the fix. Then you ask a side question about something unrelated. The context window fills up. Claude starts "compacting" — code for "summarizing away the details you actually needed."

Twenty minutes later it suggests a solution you already tried. It doesn't remember why you rejected it.

You restart the session. Now you spend forty minutes re-explaining the architecture, the bug, the constraints. You're paying to teach the same thing twice.

Octomind 0.23.0 solves this.

This release introduces daemon mode and webhooks — infrastructure for long-running, event-driven agents that don't lose their place. It also adds exponential cooldown compression that keeps context fresh without burning tokens, and per-agent model overrides so specialists use the right tool for the job.

The release breaks down like this.


Daemon Mode: Start Once, Run Forever

The big shift: Octomind now runs as a persistent background service.

bash
# Start a daemon that stays alive
octomind run --name ci-watcher --daemon --format jsonl

# Send it tasks anytime — context persists
octomind send --name ci-watcher "Check the build status"

# Or pipe from stdin
echo "Check the build status" | octomind send --name ci-watcher

Why this matters: Every other AI tool starts fresh. New session, empty context, re-establish everything. Claude Code, Cursor, Gemini CLI — they all treat each invocation as a blank slate.

Octomind daemons keep their memory. The session stays resident. Tools stay initialized. Context accumulates intelligently. You feed it tasks via send (renamed from inject for clarity) and it remembers what you were working on yesterday.

This enables continuous operation: monitoring, background processing, and event reaction. The session persists until you kill it or the system shuts down.


Webhooks: Your AI Reacts to the Real World

Daemon mode is useful on its own. Combined with webhooks, it becomes something more — infrastructure.

Webhooks are HTTP listeners that receive external events, filter them through your scripts, and inject structured tasks into your running daemon.

toml
[[hooks]]
name = "github-push"
bind = "0.0.0.0:9876"
script = "/opt/hooks/process-github-push.sh"
timeout = 30

Activate hooks with --hook when starting a daemon:

bash
octomind run --name ci-watcher --daemon --format jsonl --hook github-push

When GitHub POSTs to your server:

  1. Webhook receives the payload
  2. Your script filters noise (ignore non-master branches, skip bots)
  3. Formatted output injects into the daemon as a user message
  4. AI processes with full tool access — reads actual files, not just the payload
bash
#!/bin/bash
# HTTP body on stdin; env vars: HOOK_NAME, HOOK_METHOD, HOOK_PATH, HOOK_CONTENT_TYPE
payload=$(cat)
branch=$(echo "$payload" | jq -r '.ref' | sed 's|refs/heads/||')

# Skip noise — non-zero exit = ignore
[ "$branch" != "master" ] && exit 1

# stdout → injected as user message into the daemon session
echo "Push to $(echo "$payload" | jq -r '.repository.full_name') by $(echo "$payload" | jq -r '.pusher.name'). Review for issues."

What this replaces:

  • Polling scripts that burn API calls checking if something happened
  • Zapier workflows that can't access your actual codebase
  • Manual code review waiting for CI to fail instead of catching issues immediately
  • Context switching between Slack, PagerDuty, and your terminal

One daemon can listen on multiple hooks — stack --hook flags:

bash
octomind run --name ops-agent --daemon --format jsonl \
  --hook github-push --hook slack-notify --hook pagerduty

GitHub pushes on 9876, Slack mentions on 9877, PagerDuty alerts on 9878. Each event becomes a task for an AI that already knows your codebase.


Scheduled Tasks: Set It and Forget It

Daemon mode plus the built-in schedule tool gives you time-based execution:

text
> Schedule a reminder in 30 minutes to check the CI build

# 30 minutes later, fires automatically
# AI reads CI status, reports failures, suggests fixes

Schedule absolute times (3:00pm), relative delays (in 2h), or datetimes (2026-03-30 10:00). The session stays alive until all scheduled messages fire.

Use case: Long-running development workflows with timed checkpoints. Start a complex refactor, schedule verification steps, let the agent remind you of context you might have forgotten.


Exponential Cooldown: Compression That Doesn't Gaslight You

Long-running daemons amplify the context rot problem. Every AI tool hits this: conversation grows, token costs explode, context window fills. The usual fix? Aggressive summarization that "compacts" away the details you actually need.

Octomind has always compressed intelligently. 0.23.0 makes it smarter.

Previous versions used progressive compression levels — increasingly aggressive summarization as sessions grew. It worked, but sometimes over-corrected. You'd ask about a decision from an hour ago and get "I don't have that detail in my current context."

0.23.0 replaces this with exponential cooldown. Instead of escalating compression aggression, we wait exponentially longer between compressions as sessions stabilize. A busy session compresses frequently; a stable one gets breathing room.

The result: better retention of recent context without the cost explosions. Your daemon can run for days and still remember why you chose that database schema.


Per-Agent Model Overrides: Right Model, Right Job

Not every task needs Claude 3.7 Sonnet. Not every task can get away with GPT-4.1-mini.

Two levels of control. Agent authors set a default model in the manifest — they pick what works best:

toml
# In the agent's [[roles]] section
model = "openrouter:anthropic/claude-sonnet-4"

Users override per-agent in their config without touching manifests:

toml
# In your octomind config
[taps]
"developer:general" = "anthropic:claude-3-7-sonnet-20250219"
"octomind:assistant" = "openrouter:openai/gpt-4.1-mini"

Priority: CLI --model flag > user [taps] override > agent manifest > global default. The agent author picks what works best; you override when you know better.


Also in 0.23.0

  • Shell completions — tab-complete agent names and flags in bash, zsh, and fish
  • Windows named pipe support — cross-platform daemon messaging (\\.\pipe\octomind-<name>)
  • Session ID support for HTTP MCP servers — stateful server isolation
  • stderr capture from MCP servers — server errors surfaced for easier debugging
  • Task-aware compression — drains completed requests before compressing
  • Migration to rmcp SDK — official MCP spec compliance (v1.3.0)
  • Dynamic MCP servers — servers can be added/removed during a session
  • Inbox notification system — external tasks surface as in-session banners
  • inject renamed to send — clearer semantics for daemon messaging

Breaking Change

Configuration migrated from provider-based to layered architecture. The old provider setup guides have been replaced with a new documentation structure covering roles, layers, and the tap system. The inject command was renamed to send.

If you were importing session modules directly or relying on the old doc paths, you may need to update. See the migration guide for details.

For standard usage — octomind run, config files, tap agents — everything works as before.


What's Next

Daemon mode and webhooks are infrastructure. On top of this foundation:

  • Persistent memory across restarts — agents that survive reboots (since shipped: see the Octobrain integration)
  • Multi-agent coordination — daemons delegating to specialist sub-agents
  • Production observability — metrics, tracing, health checks
  • Pre-built webhook handlers — common services without custom scripts

But now: more agents in the tap. The runtime supports continuous operation. Time to populate it with specialists that run 24/7.


Upgrade Now

bash
# Install or update
curl -fsSL https://octomind.run/install.sh | bash

# Verify
octomind --version  # 0.23.0

Full changelog: github.com/muvon/octomind/blob/master/CHANGELOG.md

Documentation: octomind.run/docs

Community tap: github.com/muvon/octomind-tap


Octomind is an open-source AI agent runtime. Specialist agents, zero setup, any provider, no lock-in.

GitHub | Website