# Octomind 0.23.0: Your AI Agent That Actually Remembers

> Octomind 0.23.0 adds daemon mode, webhooks, exponential cooldown compression and per-agent model overrides — infrastructure for agents that keep their context.

**The problem with every AI coding tool: they forget.**

You've been there. Three hours into a debugging session with Claude Code, you've explored five dead ends, finally found the root cause, and you're halfway through the fix. Then you ask a side question about something unrelated. The context window fills up. Claude starts "compacting" — code for "summarizing away the details you actually needed."

Twenty minutes later it suggests a solution you already tried. It doesn't remember why you rejected it.

You restart the session. Now you spend forty minutes re-explaining the architecture, the bug, the constraints. You're paying to teach the same thing twice.

**Octomind 0.23.0 solves this.**

This release introduces **daemon mode** and **webhooks** — infrastructure for long-running, event-driven agents that don't lose their place. It also adds exponential cooldown compression that keeps context fresh without burning tokens, and per-agent model overrides so specialists use the right tool for the job.

The release breaks down like this.

---

## Daemon Mode: Start Once, Run Forever

The big shift: Octomind now runs as a persistent background service.

```bash
# Start a daemon that stays alive
octomind run --name ci-watcher --daemon --format jsonl

# Send it tasks anytime — context persists
octomind send --name ci-watcher "Check the build status"

# Or pipe from stdin
echo "Check the build status" | octomind send --name ci-watcher
```

**Why this matters:** Every other AI tool starts fresh. New session, empty context, re-establish everything. Claude Code, Cursor, Gemini CLI — they all treat each invocation as a blank slate.

Octomind daemons **keep their memory**. The session stays resident. Tools stay initialized. Context accumulates intelligently. You feed it tasks via `send` (renamed from `inject` for clarity) and it remembers what you were working on yesterday.

This enables **continuous operation**: monitoring, background processing, and event reaction. The session persists until you kill it or the system shuts down.

---

## Webhooks: Your AI Reacts to the Real World

Daemon mode is useful on its own. Combined with **webhooks**, it becomes something more — infrastructure.

Webhooks are HTTP listeners that receive external events, filter them through your scripts, and inject structured tasks into your running daemon.

```toml
[[hooks]]
name = "github-push"
bind = "0.0.0.0:9876"
script = "/opt/hooks/process-github-push.sh"
timeout = 30
```

Activate hooks with `--hook` when starting a daemon:

```bash
octomind run --name ci-watcher --daemon --format jsonl --hook github-push
```

When GitHub POSTs to your server:

1. Webhook receives the payload
2. Your script filters noise (ignore non-master branches, skip bots)
3. Formatted output injects into the daemon as a user message
4. AI processes with full tool access — reads actual files, not just the payload

```bash
#!/bin/bash
# HTTP body on stdin; env vars: HOOK_NAME, HOOK_METHOD, HOOK_PATH, HOOK_CONTENT_TYPE
payload=$(cat)
branch=$(echo "$payload" | jq -r '.ref' | sed 's|refs/heads/||')

# Skip noise — non-zero exit = ignore
[ "$branch" != "master" ] && exit 1

# stdout → injected as user message into the daemon session
echo "Push to $(echo "$payload" | jq -r '.repository.full_name') by $(echo "$payload" | jq -r '.pusher.name'). Review for issues."
```

**What this replaces:**

- **Polling scripts** that burn API calls checking if something happened
- **Zapier workflows** that can't access your actual codebase
- **Manual code review** waiting for CI to fail instead of catching issues immediately
- **Context switching** between Slack, PagerDuty, and your terminal

One daemon can listen on multiple hooks — stack `--hook` flags:

```bash
octomind run --name ops-agent --daemon --format jsonl \
  --hook github-push --hook slack-notify --hook pagerduty
```

GitHub pushes on 9876, Slack mentions on 9877, PagerDuty alerts on 9878. Each event becomes a task for an AI that already knows your codebase.

---

## Scheduled Tasks: Set It and Forget It

Daemon mode plus the built-in `schedule` tool gives you time-based execution:

```
> Schedule a reminder in 30 minutes to check the CI build

# 30 minutes later, fires automatically
# AI reads CI status, reports failures, suggests fixes
```

Schedule absolute times (`3:00pm`), relative delays (`in 2h`), or datetimes (`2026-03-30 10:00`). The session stays alive until all scheduled messages fire.

**Use case:** Long-running development workflows with timed checkpoints. Start a complex refactor, schedule verification steps, let the agent remind you of context you might have forgotten.

---

## Exponential Cooldown: Compression That Doesn't Gaslight You

Long-running daemons amplify the [context rot](https://octomind.run/blog/what-is-context-rot) problem. Every AI tool hits this: conversation grows, token costs explode, context window fills. The usual fix? Aggressive summarization that "compacts" away the details you actually need.

Octomind has always compressed intelligently. **0.23.0 makes it smarter.**

Previous versions used progressive compression levels — increasingly aggressive summarization as sessions grew. It worked, but sometimes over-corrected. You'd ask about a decision from an hour ago and get "I don't have that detail in my current context."

**0.23.0 replaces this with exponential cooldown.** Instead of escalating compression aggression, we wait exponentially longer between compressions as sessions stabilize. A busy session compresses frequently; a stable one gets breathing room.

The result: **better retention of recent context** without the cost explosions. Your daemon can run for days and still remember why you chose that database schema.

---

## Per-Agent Model Overrides: Right Model, Right Job

Not every task needs Claude 3.7 Sonnet. Not every task can get away with GPT-4.1-mini.

Two levels of control. Agent authors set a default model in the manifest — they pick what works best:

```toml
# In the agent's [[roles]] section
model = "openrouter:anthropic/claude-sonnet-4"
```

Users override per-agent in their config without touching manifests:

```toml
# In your octomind config
[taps]
"developer:general" = "anthropic:claude-3-7-sonnet-20250219"
"octomind:assistant" = "openrouter:openai/gpt-4.1-mini"
```

Priority: CLI `--model` flag > user `[taps]` override > agent manifest > global default. The agent author picks what works best; you override when you know better.

---

## Also in 0.23.0

- **Shell completions** — tab-complete agent names and flags in bash, zsh, and fish
- **Windows named pipe support** — cross-platform daemon messaging (`\\.\pipe\octomind-<name>`)
- **Session ID support for HTTP MCP servers** — stateful server isolation
- **stderr capture from MCP servers** — server errors surfaced for easier debugging
- **Task-aware compression** — drains completed requests before compressing
- **Migration to rmcp SDK** — official MCP spec compliance (v1.3.0)
- **Dynamic MCP servers** — servers can be added/removed during a session
- **Inbox notification system** — external tasks surface as in-session banners
- **`inject` renamed to `send`** — clearer semantics for daemon messaging

---

## Breaking Change

Configuration migrated from provider-based to layered architecture. The old provider setup guides have been replaced with a new documentation structure covering roles, layers, and the tap system. The `inject` command was renamed to `send`.

If you were importing session modules directly or relying on the old doc paths, you may need to update. See the [migration guide](https://octomind.run/docs/troubleshooting/02-migration-guide) for details.

For standard usage — `octomind run`, config files, tap agents — everything works as before.

---

## What's Next

Daemon mode and webhooks are infrastructure. On top of this foundation:

- **Persistent memory across restarts** — agents that survive reboots (since shipped: see the [Octobrain integration](https://octomind.run/blog/octobrain-octomind-integration))
- **Multi-agent coordination** — daemons delegating to specialist sub-agents
- **Production observability** — metrics, tracing, health checks
- **Pre-built webhook handlers** — common services without custom scripts

But now: **more agents in [the tap](https://octomind.run/blog/introducing-taps)**. The runtime supports continuous operation. Time to populate it with specialists that run 24/7.

---

## Upgrade Now

```bash
# Install or update
curl -fsSL https://octomind.run/install.sh | bash

# Verify
octomind --version  # 0.23.0
```

Full changelog: [github.com/muvon/octomind/blob/master/CHANGELOG.md](https://github.com/muvon/octomind/blob/master/CHANGELOG.md)

Documentation: [octomind.run/docs](https://octomind.run/docs)

Community tap: [github.com/muvon/octomind-tap](https://github.com/muvon/octomind-tap)

---

**Octomind is an open-source AI agent runtime.** Specialist agents, zero setup, any provider, no lock-in.

[GitHub](https://github.com/muvon/octomind) | [Website](https://octomind.run)
