Skip to content

Memory ​

Each agent has a private memory inbox. Memory is plain Markdown files on disk, indexed by SQLite (sqlite-vec for vector + FTS5 for BM25) for fast hybrid retrieval. Inboxes are short-term — Deep promotes stable knowledge to the shared wiki and deletes the source.

somora unifies three sources behind one retrieval pipeline. Every search (auto-injection or explicit memory_search) ranks across all three:

SourceWhereLifecycleSource-tag in hits
memory~/.somora/agents/<name>/memory/*.mdper-agent, short-term, deleted by Deep on promotionmemory/<slug>
wiki<vault>/<wiki-subfolder>/**/*.mdshared, long-term, single source of truthwiki/<path>
vaultrest of <vault> outside the wiki subfolderuser-managed, read-only from somoravault/<path>

A single memory_search("garten") call returns hits from all three, ranked by hybrid score with per-source boosts (default: wiki 1.4×, memory 0.85×, vault 0.65×). Curated wiki content ranks above raw memory inbox content; both rank above unstructured vault notes.

This document covers the memory inbox — the per-agent layer. See wiki.md for the wiki layer and agents.md for how vault binding works.

Mental model — memory inbox ​

~/.somora/agents/<name>/
├── memory.db (+ -wal, -shm)      ← derived index of THIS agent's notes, rebuilt from .md if deleted
└── memory/
    ├── *.md                      ← un-consolidated notes the agent has now
    ├── .deep-skip-cache.json     ← Deep's hash-cache (skipped files)
    └── .dreams/                  ← REM extraction findings
        ├── <id>.dream.md         ← pending review
        └── processed/            ← resolved findings (audit trail)

The .md files are the source of truth. The SQLite index is derived — delete memory.db* and it rebuilds from the .md files on next agent init. Files survive git, vim, rsync, anything. The agent reads them through the same pipeline regardless of who wrote them (you, the agent itself via tools, REM extraction, or a sync from another machine).

The vault and the wiki are indexed once per instance, not once per agent, in ~/.somora/index/shared.db. The index belongs to the source, not to the reader: every agent would embed exactly the same chunks, so one copy serves them all, and one file-watcher on the vault replaces one per agent. A new agent's first turn therefore costs nothing beyond its own notes. shared.db is derived too: delete it and the server rebuilds it. When an agent DB still holds vault/wiki rows for the same embedding model, those rows are copied over first (seconds, no model call) and the vault is swept in the background; agents keep answering from their own DB until the shared index is ready. GET /health → sharedIndex shows building during that phase and ready after (api.md).

The inbox is volatile by design. Files come in via REM or memory_write; Deep consolidates them into the wiki and deletes the source on Promote/Merge. A clean inbox means everything substantive that's been observed is now in the wiki.

How retrieval works ​

Two paths flow into every chat turn:

1. Auto-injection (always on) ​

The runtime builds an embedding query from the user's current message plus the last few turns, runs hybrid search over the agent's own memory index and the shared vault/wiki index as one candidate pool, takes the top-N hits above a configurable score threshold, and hands them to the engine as a <memory-context> block. The block is per-turn context: every engine places it in front of the user message of that turn, not in the system prompt (see cache-strategy.md).

<memory-context>
Background notes recalled from your memory for this turn. Source tags:
[memory/...] = your own short-term notes; [wiki/...] = shared long-term
wiki; [vault/...] = read-only vault content. These notes are recollection,
not observation: they can be outdated and say nothing about the current
state of the system. Call `memory_search` or `memory_get` to recall more.
Having these notes never replaces using a tool: if the user asks you to
check, run, read or change something, do it with the appropriate tool
rather than answering from these notes.

## Relevant hits for this turn

### [wiki/orte/main-house · score=0.42]
<chunk content>

### [memory/note-x · score=0.31]
<chunk content>
</memory-context>

The wording of that header is deliberate and load-bearing. Phrasing such as "no tool call required" or calling the wiki "authoritative" measurably lowers the tool-call rate of smaller models: it reads as a general "you do not need tools here", and it puts recall above the actual state of the system. Framing the notes as recollection rather than observation, with an explicit "never replaces a tool", keeps the rate up. If you customise this block, keep that distinction.

The agent sees relevant notes (from any source) without having to call a search tool.

Only query-dependent content goes in this block. The wiki topology header is not part of it — it is stable for a whole session, so it sits in the system prompt instead, where the provider's prefix cache holds it. See wiki.md.

2. Tools (agent-driven, on demand) ​

When auto-injection isn't enough, the agent calls tools:

memory_search(query, limit?, minScore?, source?)
memory_get(reference)                              # full content of one item
memory_list(tag?, source?, pathPrefix?)             # browse own memory (or wiki/vault/all)
memory_write(slug, content, frontmatter?)           # write to own inbox
memory_edit(slug, content, frontmatter?)            # modify existing
memory_delete(slug)                                 # remove

memory_search ranks across all three sources by default; pass source: 'memory' | 'wiki' | 'vault' to constrain. memory_get accepts a reference like wiki/personen/familie-klein returned by search, fetches the full file content (vs. the snippet shown in search hits).

Search snippets are chunks (~400 tokens, with overlap). Full files go through memory_get — the agent decides when a snippet is enough vs needing the whole page. Adjacent chunks share one paragraph of overlap on purpose and may both match; a chunk whose lines lie entirely inside another hit from the same file is folded into the wider one before results are ranked, so the same section never lands twice in the injected block.

Hybrid retrieval mechanics ​

  • Vector — local embeddings via ONNX (Xenova/all-MiniLM-L6-v2 by default, 384-dim). Configurable in config.yaml. The model downloads once per machine to ~/.somora/models/transformers/ (a stable location that survives somora update) and is shared by every agent. Until that first download finishes — or if it ever fails — retrieval degrades gracefully to BM25-only; it upgrades to hybrid automatically once the model is present, and the next reindex backfills embeddings for anything indexed while it was unavailable. Whether the model is actually loaded is visible on GET /health as memoryEmbedder (state: idle | loading | ok | failed, plus the error); a failed load is also logged once at boot as memory.embedder_boot_failed.
  • BM25 — SQLite FTS5 over chunk text. Tokenizer drops punctuation, lowercases everything (so [[wiki-link]] tokenizes to wiki and link).
  • Fusion — min-max normalize each modality independently, weighted sum (default 0.7 vector + 0.3 BM25), apply per-source boost.
  • Auto-inject minScore — default 0.35. Hits below this score don't appear in the inject block. Configurable.

Tunables live in config.yaml:

yaml
memory:
  embedding:
    provider: local
    model: all-MiniLM-L6-v2
  chunking:
    targetTokens: 400
    overlapTokens: 80
  autoInject:
    queryTurns: 3            # current message + the 2 turns before it
    maxResults: 5
    minScore: 0.35
    maxTokens: 1500
    historyWeight: 0.3       # how much those turns steer the vector query
    historyWeightShort: 0.55 # … when the message has only 1–2 content words
    historyWeightEmpty: 0.8  # … when it has none ("das solltest du wissen oder?")
    historyTurnChars: 800    # head of each turn that goes into the blend
    shortQueryBm25Weight: 0.5 # BM25 share for a 1–2-word question (null = hybrid default)
  rescanMinutes: 10          # full sweep of the vault/wiki index every N minutes
                             # (0 = off) — catches files saved from another
                             # machine onto a network share, which the file-
                             # watcher cannot see; unchanged files are skipped
                             # by hash, so a sweep of ~1000 files takes ~1 s
  hybrid:
    vectorWeight: 0.7
    bm25Weight: 0.3
    slugMatchBoost: 1.5      # page whose slug names a query word (1 = off)
    slugFullNameBoost: 1.5   # … and extra when the query names the page in full (1 = off)
    logDemotion: 0.5         # the wiki's monthly change logs rank behind the pages (1 = off)
    pageSupport: 0.3         # a page matching in several sections gains support (0 = off)

How the query is built. The current message is the query. The previous queryTurns - 1 turns are context: embedded separately and blended into the message embedding at historyWeight — so the question decides and the conversation nudges. Concatenating everything into one text would let two long answers about something else outvote a short question. Refinements, measured on replayed real sessions:

  • The weight adapts to how much the message says. Three or more content words (everything that is not a filler word like "ok", "kannst", "mir", "so" — a built-in German + English stopword list, also dropped from BM25 queries): historyWeight. One or two ("und seine frau?"): historyWeightShort. None ("das solltest du aber wissen oder?"): historyWeightEmpty — the conversation is the topic.
  • A message with a content word also runs on its own, and a chunk keeps the better of the two vector scores — each query vector's candidates normalised on their own first, because a long blend scores every page higher than a five-word question does. A page the question names outright is never pushed down by the history; a follow-up without a topic word still finds its page through the history.
  • For a one- or two-word question ("wer ist karl?") the exact word match is the question, so BM25 gets shortQueryBm25Weight of the fusion instead of the hybrid default.
  • hybrid.slugMatchBoost (all searches, not only auto-inject): a chunk whose slug contains a content word of the query is multiplied. The page ABOUT a person or thing rarely repeats its own name — the Karl page says "Karl" once, the family page four times — so BM25 alone ranks the mentions above the page; the name in the slug marks the canonical page.

Which page wins — the three rules after the fusion ​

Measured on 26 real questions against a 400-page wiki (2026-10-01; the report that started it: an agent asked "how did we set up acme.com on dmz-host?", the project page came 7th behind the change log and four neighbouring pages, and the agent answered that it knew nothing). With all three rules the page is 2nd, the other 25 questions are as good or better, and the questions that ARE about the log still find it.

  • logDemotion (default 0.5) — Deep writes a monthly change log (logs/2026-09): one dense line per page it created or changed, every keyword of the topic. For "how did we set up X" that line is a pointer; the page about X is the answer. Hits in logs/ are multiplied by this factor — unless the question is about the chronicle: a month name (september, march), a year or 2026-09, or a word like changed, when, promoted, created, log in the question switches the demotion off, so "what changed in September?" and "when was the page created?" keep the log on top. Set 1 to rank logs like any page; a lower value (0.3) if your logs keep winning on questions about the thing itself — check with memory_search before and after.
  • pageSupport (default 0.3) — hits are chunks, not pages. A project page answers "how did we set it up" across several sections (DNS, deployment, timeline), each matching only part of the question, while a page with one dense paragraph scores higher on that paragraph. The page's best chunk gains this share of the scores of its next two chunks, counting only chunks that reach half of the best one — so a page that mentions every word somewhere does not outgrow a single exact note. Raise it (0.5) when long pages keep losing to short mentions; lower it (0.1) or 0 when short memory notes lose to long wiki pages. Measured: 0.3 lifted the project page from unranked to 2nd and cost one memory note one rank; 0.5 cost three.
  • slugFullNameBoost (default 1.5) — slugMatchBoost fires for any page whose name shares a word with the question, so "acme website dmz-host setup" boosts projekte/acme-website and infrastruktur/hosts/dmz-host-vm alike. When the question contains every word of a page's name (two words or more), that page is multiplied again — the question means that page. A tried alternative, scaling the boost by the share of name words matched, made longer names lose to shorter ones (elevenlabs-agents-preise behind elevenlabs-agents for "wie teuer sind elevenlabs agents") and was dropped.

None of the three changes what is indexed; they only reorder hits, so a new value takes effect at the next search (config reload or restart). When you tune them, keep a handful of your own questions with the page you expect and compare ranks before and after — the ranking is sensitive to wording, and a value that fixes one question can cost another.

The BM25 side sees the message only, with filler words removed (FTS_STOPWORDS in src/memory/retrieval.ts, German and English) — otherwise every page that says "was", "du" and "so" a lot outranked the one page that says "karl". POST /agents/<name>/memory/recall-preview runs exactly this path for a message plus a supplied history, which is how recall tuning is measured (api.md).

Writing memory ​

Three ways content lands in the memory inbox:

  1. You edit ~/.somora/agents/<name>/memory/<slug>.md directly. A chokidar file-watcher re-indexes within ~1.5 s of save. The agent sees your edit on the next turn. Works with vim, VSCode, Obsidian, any editor that does atomic-rename writes.

    The same watcher covers the vault and the wiki — with one limit: it only sees writes made on the somora host. A vault on a network share (SMB/NFS) that you edit from another machine gets no file events here, so those files are picked up by the periodic sweep instead (memory.rescanMinutes, default every 10 minutes) and, as before, by the full sweep at every server start. When the share is unreachable while the server boots, the watcher retries on its own with a growing delay (30 s, 60 s, … up to 10 min) instead of staying silent until the next restart (memory.watcher_retry / memory.watcher_recovered in the log).

  2. The agent writes it via tool. memory_write (create or replace), memory_edit (modify existing, fail if missing), memory_delete (remove). Slugs are limited to lowercase [a-z0-9_-] so the agent can never accidentally write outside the memory directory.

  3. Via REM (the per-agent dream phase). REM extracts facts from session transcripts and proposes memory_write / memory_edit / memory_delete findings. You approve via dream_apply; the underlying tool call writes the file. Nothing lands in memory without your say-so.

File format ​

markdown
---
slug: garten
description: Notes about the garden
tags: [home, places]
created: 2026-04-15
updated: 2026-05-01
---

# Garten

Der Garten ist ca. 2000 m², aufgeteilt auf vier zusammenhängende
Grundstücke …

Frontmatter is optional. description (if present) is shown by memory_list. The write tools manage created/updated automatically. Add an wiki_promote: false field to opt a single memory file out of Deep evaluation:

yaml
---
slug: scratch
wiki_promote: false        # stays in memory inbox forever; Deep ignores
---

Useful for scratchpads or transient state you don't want consolidated.

How memory inboxes get drained ​

The inbox is not where things accumulate forever. Deep runs every 12h (or via dream_run({phase:'deep'})) and decides per file:

  • Skip — too thin, transient, already in wiki. File stays.
  • Promote — new wiki topic. New page is created in the wiki. Source memory file is deleted.
  • Merge — wiki page exists, new content integrated. Source memory file is deleted.

After a few Deep runs, your inbox typically contains only:

  • Files Deep has skipped (cached by hash so they're not re-evaluated next run unless content changes)
  • Recent additions since the last Deep run

See dream-phases.md for the full Deep mechanic. The inbox is intended to look mostly empty most of the time — that's a sign Deep is working.

Obsidian vault as a read source ​

Configure server-globally:

yaml
obsidian:
  vault: ~/Documents/Vault/
wiki:
  enabled: true
  vaultSubfolder: somora    # → ~/Documents/Vault/somora/ becomes the wiki
  language: de              # de | en — headings, folders, index/log wording (wiki.md)

All agents share this single vault, and so do they share its index (~/.somora/index/shared.db, see the mental model above). Vault notes are recalled alongside each agent's own memory. Hits return them as vault/<path> (slugs use -- as path separator: Projects/Personal/Travel.md → Projects--Personal--Travel).

Agents CANNOT write to the vault from somora — memory_write is hard- scoped to per-agent memory directories; the wiki is written only by Deep/Lucid (server-side workers, not agent-direct).

A few notes on vault integration:

  • Dotfile directories (.obsidian/, .trash/, .git/) are skipped.
  • The same hybrid retrieval ranks across both memory + wiki + vault.
  • The wiki subfolder of the vault gets source: 'wiki'; the rest of the vault gets source: 'vault' so retrieval can boost differently.

Tool surface ​

memory_search(query, limit?, minScore?, source?)
                              hybrid recall across memory + wiki + vault.
                              minScore defaults to 0 (agent gets best top-N).
memory_get(reference)         full content of a hit; reference like
                              'memory/<slug>' or 'wiki/<path>'.
memory_list(tag?, source?, pathPrefix?)
                              list own memory inbox notes; source: wiki | vault | all
                              browses the other layers.
memory_write(slug, content, frontmatter?)
                              create or replace own-inbox note.
memory_edit(slug, content, frontmatter?)
                              modify existing inbox note; fails if missing.
memory_delete(slug)           remove inbox note (idempotent).

memory_* write tools refuse non-memory paths by construction (slug regex rejects /, uppercase, special chars). Agents cannot write to wiki or vault directly — those go through Deep/Lucid.

The tools are exposed three ways:

  • For claude-cli (and grok-cli): via a local stdio MCP server (src/mcp/server.ts).
  • For codex-cli: as Codex dynamic tools on the app-server; somora serves each call from the in-process registry.
  • For openai-compatible engines: as in-process function definitions via the agent's tool-call loop.

Debug endpoints ​

When recall feels off, query the raw index directly:

bash
# How many notes are indexed for this agent (across all sources —
# memory from the agent's own DB, wiki/vault from the shared index)
curl 'http://127.0.0.1:18737/agents/<name>/memory/notes' | jq '.count'

# State of the shared vault/wiki index: building | ready, files, chunks
curl 'http://127.0.0.1:18737/health' | jq '.sharedIndex'

# Is the embedding model loaded? "failed" = BM25-only for every agent
curl 'http://127.0.0.1:18737/health' | jq '.memoryEmbedder'

# Raw search — see exactly what would be auto-injected for a given query
curl 'http://127.0.0.1:18737/agents/<name>/memory/search?q=garten&minScore=0' \
  | jq '.hits[] | {source, slug, score, vecScore, bm25Score, text: .text[0:80]}'

This returns per-modality scores (vecScore, bm25Score) and the chunk text that matched. Helpful when a recall feels off ("the agent didn't see X even though I wrote about it yesterday") to confirm whether the issue is at the index level (chunks not present) or higher (score below threshold).

See also ​