Cross-Session Learning
Use cross-session learning to carry your rules and grounded project knowledge into later sessions. This guide covers
Use cross-session learning to carry your rules and grounded project knowledge into later sessions. This guide covers configuration, inspection, retrieval, and retention for session users.
Get started
With the shipped configuration, learning is enabled. State a durable rule in a session, then use /done to finish the
task and start background extraction:
octomind run developer:generalIn this repository, preserve public API names unless I explicitly request a rename.
/done
/learning list
/learning show 1Extraction may produce no record if the evidence does not support a reusable memory. It runs in the background, so
repeat the list after it finishes. show 1 requires at least one listed record.
The learning system has two phases:
- Extraction — after
/done(or during auto-compaction), an LLM analyzes the conversation and extracts a small number of lessons from your corrections and stated rules. - Active packing — before each genuine user turn, relevant stored lessons are selected into one bounded runtime pack that accompanies specialist requests for that turn.
Each lesson has a scope that decides where it lands and how it is retrieved:
scoped(the default) — tied to a single project and role. Stored underlearning/{project}/{role_base}/and retrieved by relevance to what you're working on right now.global— a durable, user-wide preference that applies in every project and role. Stored underlearning/_/and hot records are considered for every replacement pack by importance; cold records require lexical matches.
So scoped lessons are organized project first, then role (project knowledge stays within the project, the role filters it further), while global lessons deliberately cross both boundaries. See Lesson Scope for details.
Configuration
Learning is one mechanic of the supervisor — the out-of-band control plane around the agent loop — so its config
lives under [supervisor.learning]. See [supervisor] in the config
reference for the sibling sections (gate, plan, condense).
[supervisor.learning]
enabled = true
[supervisor.learning.evolution]
enabled = false| Field | Description | Default |
|---|---|---|
enabled | Enable the learning system. | true |
Learning does not own a separate model. Extraction, recall, verification, retention, and evolution all use
[supervisor.model], which itself inherits omitted fields from [model].
[supervisor.learning.evolution] is off by default (enabled = false). When enabled, the detached learner may
compile one highest-value grounded memory per extraction into a machine-local skill or guardrail candidate. Candidates
move through shadow, bounded trial, active, and retired states; generated behavior never overwrites authored skills or
.agents/guardrails.toml.
Promotion is measured against a concurrent counterfactual rather than by counting wins. While a candidate is in
shadow, its trigger is evaluated but the behavior is not applied; after min_samples matched turns with a verify-gate
verdict, this screen opens a live trial. During the trial each session is randomized, half to the treatment
arm (behavior applied) and half to the control arm (trigger observed, behavior not applied), so both arms run over
the same period, models, and task mix. The control arm restarts when the trial opens; the shadow screen proves only
that the trigger fires. A generated skill stays loaded once its trigger activates it, so skill exposure is sticky in
both arms: from a control skill's first trigger match, or a treatment skill's trigger activation, every later verdict in
that session is a sample, together with the turn's API-call count. Neither arm depends on the model's own report of
what helped. A verdict counts as a pass only when it is execution-verified (see Supervisor);
a pass resting on the verifier's reading alone is not a sample. With [supervisor.evaluate] evolution = true, samples
are graded (the probability the answer fulfils the request) and turns where the artifact does not apply are dropped;
see Supervisor. At most one trial runs among artifacts whose scopes can bind in
the same session, so a verdict is never shared between two trials.
Each arm's pass rate is a Beta(1,1) posterior, and the trial is judged after every treatment sample once both arms
hold min_samples verdicts:
- Promoted when the treatment beats the control by more than
noise_marginand by more than 2.5 posterior standard deviations, with extra API calls withincost_allowance + cost_per_gain × gain; or when the pass rate is withinnoise_marginat that same confidence while API calls drop by more thancost_allowance, with the saving significant on per-turn log calls. - Regressed (retired) when the treatment falls below the control by more than 2.5 posterior standard deviations.
- Inconclusive (retired) after
max_trial_useslive uses without either decision.
The 2.5 bar accounts for checking after every sample: with the defaults, a behavior with no real effect is promoted
in under 5% of trials across base pass rates from 0.5 to 0.95, and one that lowers the pass rate from 0.85 to 0.70
in under 1% (a seeded simulation over the runtime's own decision code checks both). Older configurations with
min_samples = 3 and max_trial_uses = 8 stay below the same false-promotion bound but rarely promote anything.
An active behavior must keep earning its place. It is pruned once it no longer beats its control at all (the same rule with the margin relaxed to zero, so one unlucky verdict near the promotion edge does not remove it). Any shadow, trial, or active artifact with no trigger match or use for 90 days retires as stale. Rejected and retired records stay in the registry with their reason, and synthesis receives them as falsified hypotheses: a candidate drawn only from memories that already produced a record is dropped before verification. Registries written before this test (schema 1) return their trial and active records to shadow on load, so the concurrent test re-validates them.
| Field | Description | Default |
|---|---|---|
min_samples | Verify-gate verdicts in the shadow screen, then in each concurrent arm, before comparison | 10 |
noise_margin | Smallest pass-rate gain worth promoting; also the non-inferiority bound for a cheaper behavior | 0.15 |
cost_allowance | Relative API-call increase tolerated at negligible gain; also the saving that counts as cheaper | 0.10 |
cost_per_gain | Extra relative API-call increase allowed per unit of pass-rate gain | 2.0 |
max_trial_uses | Live uses before an undecided trial retires | 40 |
Evolution needs [supervisor.gate] enabled: without verdicts no arm collects samples, and trials end inconclusive.
Two synthesis modes share that pipeline. Session mode runs after each extraction on the memories that session just
stored, with the transcript as evidence. Store mode runs at most once per day (or on /learning evolution distill)
over the whole hot store: short rules that were materially used or came from a direct correction, plus verified
experiences, are clustered by wording and paraphrase, and the cluster that recurs across the most projects becomes one
candidate. Its scope follows the recurrence: two or more projects make the project dimension global, two or more role
domains make the domain dimension global, so a rule you keep restating in every repository can become one skill that
binds everywhere. Every generated skill must pass a deterministic trigger screen before the verifier runs: the
proposal's replay cases are executed against the rendered activation rules, quoted arguments and && are rejected,
and a body that embeds session evidence handles or machine-local paths is rejected as an experience dump rather than a
procedure.
The 2,000-token active-pack cap and its 512-token global-rule sub-cap are fixed constants, not knobs.
Strict config, template-provided values.
[supervisor]and its nestedlearning,gate,plan, andcondensetables are required by deserialization.LearningConfig::default()hasenabled = false, while the shipped template explicitly enables it. There is no[supervisor.learning.model]; all learning calls inherit[supervisor.model], whose omitted fields inherit main[model].
Memory types
Orientation memory
Alongside lessons (the procedural "do / avoid"), the supervisor stores orientation — durable, descriptive
understanding of the subject: how it works, key decisions, constraints. It rides the same backend under memory_type = "orientation" and is recalled as working assumptions to verify, never as truth, in the pack’s orientation group. It
is part of learning — on whenever [supervisor.learning] is enabled, with fixed injection and decay bounds.
Every orientation record cites 1–4 real user or tool messages and must pass the same grounding verifier as experience before it is stored: a citation that exists is not yet a citation that supports the claim, and a rejected record fails closed. A later record that restates an existing one (word-set Jaccard above 0.6) takes its hot slot; the restated record moves to the cold archive rather than being deleted. The comparison is symmetric, so a short fact contained in a richer record never displaces it.
Long-lived experience memory
A separate detached learner may emit one memory_type = "experience" record when a trajectory contains substantial
non-obvious knowledge that would save several searches or failed attempts. The extra call is value-gated:
verified/failed work needs real user plus tool evidence, while an outcome-unknown trajectory must also be large (at
least eight tool results and 8,000 bounded transcript tokens). Routine sessions pay only for the existing short learner.
Generic advice, activity logs, transient status, secrets, exact line numbers, and facts recoverable with one obvious
search are rejected.
An experience is 150–600 words with Objective, Durable knowledge, Outcome and evidence, and Reuse conditions sections. It carries:
- the external trajectory outcome:
verified(the verify-gate passed a turn that changed state and then ran a recognized check),failed, or honestlyunknown(including a pass on the verifier's reading alone); - 1–6 addressable
session://<session>/message/<n>evidence handles, including real user/tool evidence; - stable IDs of related short lessons or prior memories;
- a separate grounding-verifier verdict before storage. A rejected candidate gets at most one issue-driven repair and one final verification, then fails closed.
Failed trajectories may therefore produce failure-labelled experience records, while short user-backed lessons retain their existing quote-first verification contract.
Managing Lessons (/learning)
The interactive /learning command lets you browse and prune lessons for the current role and project:
The list header summarizes hot/cold item and token totals, local/global scope counts, and per-type hot/cold counts.
Individual rows stay compact; use show for full provenance and retention metadata.
| Command | Effect |
|---|---|
/learning | List lessons (page 1). |
/learning list [page] | List a specific page. 15 lessons per page. |
/learning list *pattern* | Filter by a glob pattern matched against content, title, and tags (e.g. /learning list *auth*). Combine with a page number. |
/learning show <index> | Inspect the complete memory body, file path, outcome, evidence handles, and related IDs. Alias: get. |
/learning delete <index> | Delete a lesson by its 1-based index in the current unfiltered hot list. Aliases: rm, remove. |
/learning clear | Delete all hot and cold lessons for the current role + project scope; global rules are untouched. |
/learning evolution | List evolved behavior matching the current project/domain. |
/learning evolution show <id> | Inspect scope, provenance, native artifact, trials, and history. |
/learning evolution approve|reject|rollback <id> | Explicitly control a generated behavior lifecycle. |
/learning evolution distill | Run cross-store synthesis now, in the background, ignoring the daily stamp. |
The unfiltered list covers current scoped hot records followed by global hot records, each sorted by importance. show
and delete reload that unfiltered list: filtered row numbers are not safe to reuse, and indices may change when
background learning updates the store. Re-list without a filter and inspect the entry before deleting it. clear only
wipes the current role+project scope. See Session Commands for the full command
reference.
For example, browse matches, then return to the unfiltered list before inspecting or deleting an entry:
/learning list *auth* 1
/learning list
/learning show 1
/learning delete 1To inspect generated behavior, use the ID returned by the evolution list (replace CANDIDATE_ID):
/learning evolution
/learning evolution show CANDIDATE_ID
/learning evolution approve CANDIDATE_IDapprove moves only a shadow candidate into trial. reject rejects a record; rollback moves a trial or active record
back to shadow and clears its live-trial evidence (the control arm is kept). To remove all scoped hot and cold memories:
/learning clearTo reject or roll back a listed behavior, substitute its ID:
/learning evolution reject CANDIDATE_ID
/learning evolution rollback CANDIDATE_IDCommon questions
Why is a lesson missing? Extraction is asynchronous and quote-backed rules must pass verification. Recall is relevance- and budget-limited, so a stored item need not appear in every pack. Inspect the store, then enable debug logging before a new request to see the actual pack:
/learning list
/loglevel debug
Review the API authentication rules for this repository.
/loglevel infoWhy did a filtered delete target another item? show and delete use the current unfiltered list. Always run
/learning list and inspect the matching unfiltered index immediately before deletion.
Does exiting guarantee extraction finishes? Exit starts a child process and returns immediately. Closing the terminal can terminate that child before it stores the memories; a script that starts the next run on the same project at once may recall before that child has stored them.
Retrieval and storage reference
Lesson Scope
Every lesson is classified as either scoped or global, and the extraction LLM picks the scope for each one. It is
instructed to be conservative: most lessons are scoped, and a lesson only becomes global when it is clearly about
how you work in general rather than this task, project, or role.
| Scope | Stored in | Retrieved how |
|---|---|---|
scoped (default) | learning/{project}/{role_base}/ | By relevance to your current request (hybrid keyword + embedding search) |
global | learning/_/ | Reconsidered for each active pack, ranked by importance, within the pack budget; cold records require lexical matches |
A worked example: you tell the agent "always open a single PR" while working in project octofs as
developer:general. That is a general working preference, so the extractor may classify it as a global lesson
stored in learning/_/. Later you tell it "in this repo, all API endpoints require bearer auth" — that is specific to
this project, so a grounded extracted rule is scoped and lands in learning/octofs/developer/ (note the role is
truncated at : to its base, developer).
Storage (File Backend)
Scoped lessons are stored as markdown files with YAML frontmatter, one file per lesson, in a project/role directory;
global lessons go in the shared _ directory:
~/.local/share/octomind/learning/
├── octofs/developer/ # scoped: {project}/{role_base}
│ ├── 20260405143000-bearer-auth-required.md
│ └── 20260405143001-custom-error-types.md
└── _/ # global: cross-project, cross-role
└── 20260405150000-always-single-pr.mdThe role component is the base part before : — a lesson from role developer:general is stored under
developer/, while the project component is the working directory’s basename.
On macOS and Linux the default data root is ~/.local/share/octomind; on Windows it is %LOCALAPPDATA%/octomind.
OCTOMIND_DATA_DIR overrides it, including config and learning storage:
OCTOMIND_DATA_DIR="$HOME/octomind-personal" octomind runEach file carries the full frontmatter the backend writes, in this exact order:
---
title: "Bearer token auth required for all API endpoints"
content: "Bearer token auth required for all API endpoints"
memory_type: learning
importance: 0.9
confidence: high
tags: [auth, api]
source: "260405-142040-octofs-25e37715"
role: "developer:general"
project: "octofs"
scope: scoped
created: "2026-04-05T14:30:00Z"
related: []
evidence: ["session://260405-142040-octofs-25e37715/message/1"]
outcome: unknown
last_used: ""
use_count: 0
---titleis a short summary auto-derived from the first 80 UTF-8 bytes of short-rule content, cut safely at a character/word boundary.scopeisscopedorglobaland determines which directory the file lives in.last_usedanduse_countchange only when the specialist reports that the memory materially affected its work. Recall exposure alone is neutral.
Files are human-readable and editable. Delete a file to remove a lesson — or use the /learning
command.
Extraction
Extraction is triggered by:
/done— extracts (ifsupervisor.learning.enabled) regardless of the compression result, and marks the session so/exitand Ctrl+D don't extract a second time.- Auto-compaction — every successful fold first hands a snapshot of the turns it is about to discard to the learner, so long autonomous runs learn from work that no longer fits in context.
- Session end — a detached
octomind distillchild performs extraction when an interactive session ends naturally via/exit,/quit, or Ctrl+D, and when a one-shot run (octomind runwith piped input or--format) finishes, so autonomous runs learn exactly as interactive ones do. Skipped if/donealready extracted during the session; daemons learn on/doneand auto-compaction.
Extraction runs detached (an in-process task for /done/compaction, a child process at session end, with in-process
model costs folded into session spending; exit-child spending is separate) and is deliberately strict about what counts
as a lesson:
- Decision gate. The LLM first emits
<decision>LEARN</decision>or<decision>NONE</decision>. OnNONE, short-lesson parsing stops; orientation and experience are evaluated independently. - Mandatory evidence. Every
<lesson>must carry anevidenceattribute quoting the user verbatim. Missing evidence is dropped; the quote must match a real user turn and pass a separate support verifier. - At most 3 lessons per extraction — one strong lesson beats three weak ones.
- Only user corrections and user-stated rules qualify — explicit corrections, declared project conventions/preferences/constraints, or a repeated correction of the same mistake. Things the AI figured out on its own, one-off debugging steps, generic developer knowledge, and anything derivable by reading the codebase do not qualify.
Long-lived experiences are evaluated independently from that short-lesson decision. Their cited message handles are checked structurally, system-managed recall/steer messages are excluded from the transcript, and a separate verifier rejects unsupported or outcome-inflated records.
Confidence drives importance: confidence=high (a direct correction) → importance 0.9; anything else (a stated
preference, confidence=medium) → importance 0.6.
Dedup and supersede. The extraction LLM receives a bounded, ID-labelled view of existing scoped and global lessons.
Identical content is skipped. A refinement or reversal removes an older lesson only when the new quote-backed candidate
explicitly names its ID through supersedes and both records have the same scope. Similarity alone never deletes a
short user rule.
Long-run retention
File-backed learning uses a two-watermark hot store with fixed internal token budgets per scope and memory type. The soft watermark is 80% of the hard bound:
| Memory type | Scoped hard bound | Global hard bound |
|---|---|---|
| Short user-backed rules | 16,000 tokens | 4,000 tokens |
| Orientation | 24,000 tokens | 8,000 tokens |
| Experience | 48,000 tokens | 16,000 tokens |
Maintenance runs after detached extraction, never in the user-response hot path. Crossing a bucket’s hard watermark
selects at most one similar orientation/experience pair per scope and memory type as a candidate and asks the
supervisor model for a shorter consolidation. Similarity only chooses what to review; it never proves equivalence. A
separate verifier must confirm that the replacement adds no claim, hides no contradiction, preserves
applicability/outcome boundaries, and retains all non-duplicate constraints. Only then is the replacement stored and the
source records archived through atomic file writes and moves. The replacement keeps the source IDs in related, unions
their evidence, inherits the lower importance, and does not strengthen confidence or outcome.
Short user-backed rules are never synthesized by this pass because a generated merge would break their quote-first
contract. They continue to change only through explicit, separately verified extraction and supersedes.
Maintenance also reviews recurrence. When a scoped short rule appears near-verbatim (word Jaccard about 0.6 or higher)
in three or more projects, recurrence only nominates it: three projects that share a language or toolchain produce the
same stack-specific rule. A separate supervisor-model review must confirm the rule governs how you work in every
project and role; otherwise the instances stay scoped. A cluster is reviewed only when the extraction that triggered
maintenance just restated it, so a declined rule is asked again on new evidence, not on every extraction. On
confirmation the highest-importance instance moves to learning/_/ with its content unchanged, its evidence unioned,
and the other instances listed in related; those instances are cold-archived in their project scopes, never deleted.
After that single consolidation attempt, the lowest-utility records move to .archive/<memory_type>/ until the hot
store is back at 80%. Utility combines bounded importance, direct-use count, confidence, and last-use recency:
U = 0.55I + 0.15C + 0.15 min(1, ln(1+uses)/ln(11)) + 0.15/(1+age_days/180)
Here I is outcome-adjusted importance in [0,1], C is 1 for high confidence and 0.5 otherwise, and age is
measured from last_used (falling back to creation time). The logarithm rewards repeated demonstrated use without
letting frequency dominate correctness. Task relevance is deliberately absent from eviction utility because maintenance
has no current task; relevance stays the admission signal during recall.
Cold files are retained losslessly and are excluded from hot embedding recall. A compact append-only catalog keeps their title, tags, and a short preview; exact lexical matches can page at most two cold records per scoped/global retrieval into consideration without embedding the archive. Long cold experiences carry their real archive path in the Active Memory Pack, so the specialist can open the full record. A cold record reported as materially used is automatically promoted back to its hot scope before its use/outcome metadata is updated. This hysteresis prevents maintenance from moving one record on every extraction.
Independently of the hard budget, a scoped record that is both weak (importance <= 0.4) and unused for 90 days
(counted from its last material use, or creation when never used) also moves to the same cold archive. Repeated negative
outcome credit that lowers importance to 0.1 does the same immediately. Automatic retention never permanently deletes
a file; explicit /learning delete and clear remain destructive user actions.
Active Memory Pack
Every genuine user turn replaces the previous runtime pack:
- First message of the session — global rules plus a full hybrid scoped recall are considered.
- Each subsequent new user message — global rules are reconsidered and scoped recall is embedding-only, with no retrieval-prep LLM call.
- Tool follow-up rounds — reuse the same pack without another retrieval.
The file backend may rank up to 20 scoped candidates and expands explicit relationships one hop in either direction, but
only items fitting the 2,000-token pack budget under Octomind’s token estimator reach the specialist; global rules may
consume at most 512 of those tokens. Each selected item gets a short pack-local ID (M1, M2, …). The specialist
reports only IDs that materially affected its answer or action in the hidden supervisor status, and verify-gate outcomes
reinforce or weaken only those used items. Mere exposure receives no credit.
Long experience bodies are represented by a compact card (up to 320 inline tokens) plus the exact .md file path,
outcome, evidence handles, and related IDs. The specialist can inspect the full file with its normal local reader when
the card is insufficient; selected records keep their full body on disk even when only the card fits.
The pack is materialized as a system-managed user-role message only around the provider request. It is removed immediately afterwards, never appended to the session log, never accumulated across turns, and rebuilt automatically on the next genuine request. If the bounded pack alone would cross the model's usable context ceiling, it is dropped for that turn rather than blocking the user's task.
Retrieval (File Backend)
Scoped recall is a hybrid search: LLM-extracted keywords and short phrases (sparse) are fused with embedding cosine
similarity (dense) via Reciprocal Rank Fusion (RRF, k=60), then reweighted by recency and learned importance. An exact
sparse phrase receives strong credit; when the phrase is absent, at least two selective constituent terms must match,
preventing one generic word from admitting a memory.
Sparse and dense normally receive equal RRF weight. If one of the first three sparse hits has learned importance below
0.4, that query is treated as correction-conflicted and sparse weight becomes 0.25; dense outage always restores
full sparse ordering. One highest-importance sparse candidate may be reserved at rank five when fusion buried it,
preserving identifier and indirect cue recall without letting lexical noise control ranks one through four.
Dense retrieval keeps a short memory as one unchanged embedding input. A long heterogeneous memory is divided at semantic line/paragraph boundaries into bounded 128-token chunks, with title and tags attached to every chunk; the memory's dense score is its strongest chunk match. This late interaction keeps small facts from being diluted by unrelated sections while preserving the legacy score exactly for ordinary one-chunk lessons.
Recency uses a 30-day half-life with up to a +50% boost; importance contributes a bounded 0.75x–1.25x multiplier so
relevance remains primary. Embedding candidates below a 0.2 cosine floor are dropped as noise, and if the embedding
model isn't ready yet the cosine signal is silently skipped. The query-rewrite output is accepted only as 3–5 short
keyword lines; malformed or answer-like responses fail safely to retrieval without rewritten patterns. The rewrite call
runs on the first retrieval and after /done resets recall; an empty scoped store skips it; follow-up messages use
embedding-only recall.
With /loglevel debug, retrieval prints the accepted query-rewrite keywords and the exact final Active Memory Pack
after context-headroom checks, immediately before it is materialized for the provider request. Normal and info logging
keep showing only compact retrieval and pack totals.
Relationship to Memory
Learning is separate from external memory MCP tools:
- External memory tools may provide broad context storage — code patterns, architecture, project state, references.
- Learning is narrow and structured — actionable facts scored by confidence, extracted from outcomes, with deduplication.
Both can coexist. Supervisor learning is always file-backed and owns its verified retention lifecycle; external memory tools remain independent MCP tools the specialist may use directly. Learning focuses on the corrections and rules you gave the agent, and surfaces relevant ones automatically.
Source reference
| Surface | Source |
|---|---|
| Defaults and model ownership | config-templates/default.toml, src/config/model.rs |
| Extraction, evidence, and exit child | src/supervisor/learning/extract.rs |
| Pack and retrieval | src/supervisor/learning/inject.rs, src/supervisor/learning/backend/file.rs |
| Retention and evolution | src/supervisor/learning/retention.rs, src/supervisor/learning/evolution/runtime.rs |
| Commands and paths | src/session/chat/session/commands/learning.rs, src/directories.rs |