The problem with every AI coding tool: they forget.
You've been there. Three hours into a debugging session with Claude Code, you've explored five dead ends, finally found the root cause, and you're halfway through the fix. Then you ask a side question about something unrelated. The context window fills up. Claude starts "compacting" — code for "summarizing away the details you actually needed."
Twenty minutes later it suggests a solution you already tried. It doesn't remember why you rejected it.
You restart the session. Now you spend forty minutes re-explaining the architecture, the bug, the constraints. You're paying to teach the same thing twice.
Octomind 0.23.0 solves this.
This release introduces daemon mode and webhooks — infrastructure for long-running, event-driven agents that don't lose their place. It also adds exponential cooldown compression that keeps context fresh without burning tokens, and per-agent model overrides so specialists use the right tool for the job.
The release breaks down like this.
Daemon Mode: Start Once, Run Forever
The big shift: Octomind now runs as a persistent background service.
# Start a daemon that stays alive
octomind run --name ci-watcher --daemon --format jsonl
# Send it tasks anytime — context persists
octomind send --name ci-watcher "Check the build status"
# Or pipe from stdin
echo "Check the build status" | octomind send --name ci-watcherWhy this matters: Every other AI tool starts fresh. New session, empty context, re-establish everything. Claude Code, Cursor, Gemini CLI — they all treat each invocation as a blank slate.
Octomind daemons keep their memory. The session stays resident. Tools stay initialized. Context accumulates intelligently. You feed it tasks via send (renamed from inject for clarity) and it remembers what you were working on yesterday.
This enables continuous operation: monitoring, background processing, and event reaction. The session persists until you kill it or the system shuts down.
Webhooks: Your AI Reacts to the Real World
Daemon mode is useful on its own. Combined with webhooks, it becomes something more — infrastructure.
Webhooks are HTTP listeners that receive external events, filter them through your scripts, and inject structured tasks into your running daemon.
[[hooks]]
name = "github-push"
bind = "0.0.0.0:9876"
script = "/opt/hooks/process-github-push.sh"
timeout = 30Activate hooks with --hook when starting a daemon:
octomind run --name ci-watcher --daemon --format jsonl --hook github-pushWhen GitHub POSTs to your server:
- Webhook receives the payload
- Your script filters noise (ignore non-master branches, skip bots)
- Formatted output injects into the daemon as a user message
- AI processes with full tool access — reads actual files, not just the payload
#!/bin/bash
# HTTP body on stdin; env vars: HOOK_NAME, HOOK_METHOD, HOOK_PATH, HOOK_CONTENT_TYPE
payload=$(cat)
branch=$(echo "$payload" | jq -r '.ref' | sed 's|refs/heads/||')
# Skip noise — non-zero exit = ignore
[ "$branch" != "master" ] && exit 1
# stdout → injected as user message into the daemon session
echo "Push to $(echo "$payload" | jq -r '.repository.full_name') by $(echo "$payload" | jq -r '.pusher.name'). Review for issues."What this replaces:
- Polling scripts that burn API calls checking if something happened
- Zapier workflows that can't access your actual codebase
- Manual code review waiting for CI to fail instead of catching issues immediately
- Context switching between Slack, PagerDuty, and your terminal
One daemon can listen on multiple hooks — stack --hook flags:
octomind run --name ops-agent --daemon --format jsonl \
--hook github-push --hook slack-notify --hook pagerdutyGitHub pushes on 9876, Slack mentions on 9877, PagerDuty alerts on 9878. Each event becomes a task for an AI that already knows your codebase.
Scheduled Tasks: Set It and Forget It
Daemon mode plus the built-in schedule tool gives you time-based execution:
> Schedule a reminder in 30 minutes to check the CI build
# 30 minutes later, fires automatically
# AI reads CI status, reports failures, suggests fixesSchedule absolute times (3:00pm), relative delays (in 2h), or datetimes (2026-03-30 10:00). The session stays alive until all scheduled messages fire.
Use case: Long-running development workflows with timed checkpoints. Start a complex refactor, schedule verification steps, let the agent remind you of context you might have forgotten.
Exponential Cooldown: Compression That Doesn't Gaslight You
Long-running daemons amplify the context rot problem. Every AI tool hits this: conversation grows, token costs explode, context window fills. The usual fix? Aggressive summarization that "compacts" away the details you actually need.
Octomind has always compressed intelligently. 0.23.0 makes it smarter.
Previous versions used progressive compression levels — increasingly aggressive summarization as sessions grew. It worked, but sometimes over-corrected. You'd ask about a decision from an hour ago and get "I don't have that detail in my current context."
0.23.0 replaces this with exponential cooldown. Instead of escalating compression aggression, we wait exponentially longer between compressions as sessions stabilize. A busy session compresses frequently; a stable one gets breathing room.
The result: better retention of recent context without the cost explosions. Your daemon can run for days and still remember why you chose that database schema.
Per-Agent Model Overrides: Right Model, Right Job
Not every task needs Claude 3.7 Sonnet. Not every task can get away with GPT-4.1-mini.
Two levels of control. Agent authors set a default model in the manifest — they pick what works best:
# In the agent's [[roles]] section
model = "openrouter:anthropic/claude-sonnet-4"Users override per-agent in their config without touching manifests:
# In your octomind config
[taps]
"developer:general" = "anthropic:claude-3-7-sonnet-20250219"
"octomind:assistant" = "openrouter:openai/gpt-4.1-mini"Priority: CLI --model flag > user [taps] override > agent manifest > global default. The agent author picks what works best; you override when you know better.
Also in 0.23.0
- Shell completions — tab-complete agent names and flags in bash, zsh, and fish
- Windows named pipe support — cross-platform daemon messaging (
\\.\pipe\octomind-<name>) - Session ID support for HTTP MCP servers — stateful server isolation
- stderr capture from MCP servers — server errors surfaced for easier debugging
- Task-aware compression — drains completed requests before compressing
- Migration to rmcp SDK — official MCP spec compliance (v1.3.0)
- Dynamic MCP servers — servers can be added/removed during a session
- Inbox notification system — external tasks surface as in-session banners
injectrenamed tosend— clearer semantics for daemon messaging
Breaking Change
Configuration migrated from provider-based to layered architecture. The old provider setup guides have been replaced with a new documentation structure covering roles, layers, and the tap system. The inject command was renamed to send.
If you were importing session modules directly or relying on the old doc paths, you may need to update. See the migration guide for details.
For standard usage — octomind run, config files, tap agents — everything works as before.
What's Next
Daemon mode and webhooks are infrastructure. On top of this foundation:
- Persistent memory across restarts — agents that survive reboots (since shipped: see the Octobrain integration)
- Multi-agent coordination — daemons delegating to specialist sub-agents
- Production observability — metrics, tracing, health checks
- Pre-built webhook handlers — common services without custom scripts
But now: more agents in the tap. The runtime supports continuous operation. Time to populate it with specialists that run 24/7.
Upgrade Now
# Install or update
curl -fsSL https://octomind.run/install.sh | bash
# Verify
octomind --version # 0.23.0Full changelog: github.com/muvon/octomind/blob/master/CHANGELOG.md
Documentation: octomind.run/docs
Community tap: github.com/muvon/octomind-tap
Octomind is an open-source AI agent runtime. Specialist agents, zero setup, any provider, no lock-in.



