Skip to content

somora docs — start here ​

A short guide so you don't have to read every page to get started. somora is a local-first gateway for personal AI agents with persistent memory. It runs as a long-lived service on your own machine, exposes an HTTP+SSE API plus three first-party clients (TUI, web desktop, mobile PWA), and lets you route conversations across Claude / ChatGPT / any OpenAI-compatible model — all without giving up ownership of the chat log, the memory layer, or your knowledge base.

What is somora — in one paragraph ​

You define agents (personas with their own model, memory inbox, and session). They come in two kinds: chat agents — the ones you talk to, that remember and coordinate — and builders, coding harnesses that take a repository, plan, build, test and report (builder.md). You talk to both from the TUI, the browser, or your phone. The server orchestrates the actual LLM call against the engine you configured (subscription or API), persists every message as JSONL, and runs a three-phase background dream system that turns raw session content into curated memory notes and a shared wiki over time. All of that stays on your machine.

Choose your path ​

You probably don't need every page below right away.

I just want to try it. Start with the project README quickstart — one line installs everything and starts the setup assistant, which connects a model, creates your first agent and sends it a test message. ~10 minutes.

I want to chat from the browser or my phone. Read setup.md → HTTPS via Tailscale (or run somora setup access, which does it for you) to get a real cert, then open https://<your-tailnet>.ts.net:18737/web/ in the browser. The mobile PWA is the same URL with /mobile/ — see mobile.md.

I want to add a new provider / model. setup.md → Configuring providers plus the comment block in config.example.yaml is the full reference.

I want an agent to build software. Create one of kind builder (builder.md): it plans in a repository, you press Go, it builds, tests and reports; another agent can hand it the order with the bundled builder-handover skill.

I want my agent to do real work — read files, run commands, search the web, edit notes. The tools overview lists the tool families. files.md, resources.md, and tmux.md cover the heavy ones. Skills are markdown how-tos you install to teach an agent multi-step recipes.

I want to decide what each agent may use. Tools and skills are shared by all agents; the web client's Abilities window (or agent.yaml) hides any of them for a given agent. mcp.md → Per-agent tool control and skills.md → Per-agent visibility.

I want my agents to make video. videogen.md — a render takes minutes, so the agent starts one and is woken when it is ready rather than holding its turn open. Off until videoGen is configured; verified against a self-hosted endpoint, the hosted providers are untested.

I want my agents to make images. imagegen.md — configure a model, then generate from the Media window or let an agent call image_generate. Results land in your workspace and show up in the chat.

I want long-term memory + a shared knowledge base. Read memory.md for the per-agent inbox, wiki.md for the shared layer, and dream-phases.md for how sessions automatically flow into both.

I want to build a third-party client / integration. api.md is the HTTP+SSE contract — same surface the built-in clients use.

Core concepts in one sentence each ​

  • agent — a persona with its own AGENTS.md, memory inbox, default model, and chat sessions. Lives at ~/.somora/agents/<name>/.
  • session — a single chat thread with one agent. Persisted as JSONL, resumable, switchable mid-conversation.
  • engine — the adapter that talks to an LLM backend. Four exist: claude-cli (Claude subscription), codex-cli (ChatGPT subscription, bundled), grok-cli (SuperGrok subscription), openai-compatible (any /v1/chat/completions endpoint — OpenRouter, Ollama, oMLX, LM Studio, …).
  • provider / model / alias — config.yaml declares providers (with baseUrl + apiKey if needed) and models on each. An alias lets you refer to a model by short nickname anywhere.
  • memory — three layers an agent reads from: its own per-agent inbox (~/.somora/agents/<name>/memory/*.md), the shared wiki, and the read-only Obsidian vault. Hybrid retrieval: vector + BM25.
  • wiki — a curated long-term knowledge base shared across all agents. Lives under your Obsidian vault. Written by the dream system, editable by hand.
  • dream system — three background phases. REM turns finished sessions into memory-inbox notes (per-agent). Deep promotes high-value notes into the shared wiki. Lucid cleans up the wiki on a slower cadence. See dream-phases.md.
  • project — an explicit pin of "files / paths / vault notes / research artifacts that belong to this chat session". Auto-injected into the agent's system prompt. See projects.md.
  • resource — a configured SSH host. file_*, exec, and tmux tools dispatch against it via target=<resource>. See resources.md.
  • skill — a markdown how-to (agentskills.io format) installed at ~/.somora/skills/<slug>/SKILL.md. The agent loads the body on demand via the skill tool when it recognizes the situation.
  • tools — the things an agent can call. Memory reads, wiki edits, file I/O, shell exec, tmux sessions, web search/fetch, sub-agent spawning, etc. See tools.md.
  • voice — two separate features that share only the word. Dictation and spoken replies: a mic button on web/mobile, optional TTS for the answer, plus a generic /voice/turn audio-in/audio-out endpoint, all on the OpenAI-compatible audio shape with no extra credentials (voice.md). And realtime voice: a standing, interruptible call with an agent, where a second model does the talking and the agent does the knowing (realtime-voice.md).

Where to next ​

If you've installed somora and chatted with the default agent, the high-value next steps in rough order:

  1. Create a second agent with a different persona — copy the default ~/.somora/agents/<name>/AGENTS.md, edit, restart.
  2. Configure a real provider if you started on the subscription defaults — see setup.md → Configuring providers.
  3. Enable the wiki by pointing somora at your Obsidian vault in config.yaml. The dream system then has somewhere to consolidate memory into. wiki.md.
  4. Install your phone as a PWA for chatting from anywhere on your tailnet. mobile.md.
  5. Add an SSH resource so an agent can read files and run commands on a remote machine. resources.md.
  6. Pin a project to a session so the agent has a persistent working set. projects.md.

Reference docs ​

DocReads likeWhen to open it
setup.mdOperator runbookFirst install; adding providers; HTTPS; mobile; ops
agents.mdPersona + agent.yaml referenceCreating an agent; per-agent model, REM, tool/skill visibility
builder.mdThe builder kindAn agent that is a coding harness: plan → Go → build, task panel, questions, hand-over from an orchestrator
lsp.mdLanguage serversA builder's writes come back with the compiler's errors; install, config, routes
api.mdHTTP/SSE contractBuilding a third-party client; debugging streaming
tools.mdTool catalog overviewWondering what an agent can do
mcp.mdExternal MCP serversPlugging third-party MCP tools into your agents
files.mdFile-tool deep diveWorking with file_read/file_write/analyze_file
memory.mdConcept + mechanicsTuning recall, writing notes by hand
wiki.mdConcept + mechanicsShared knowledge base, hand-curation
dream-phases.mdBackground workersREM/Deep/Lucid cadence + triggers
projects.mdConcept + workflowPinning a working set to a session
resources.mdSSH-target configAdding a remote machine
skills.mdMarkdown skill formatWriting or installing skills
tmux.mdMulti-turn shell sessionsDriving long-running CLIs from agents
web.mdBrowser clientWeb-UI specifics + HTTPS notes
mobile.mdMobile PWAiOS/Android install, scope
voice.mdDictation + spoken repliesPress-to-talk in web/mobile, optional TTS answers
realtime-voice.mdTalking to an agentA standing, interruptible call; the voice talks, the agent knows
imagegen.mdText-to-imageGenerating images from the web app or an agent
videogen.mdText-to-videoJob-based renders, and how an agent gets its result without waiting
browser.mdShared browserA Chromium per agent profile that agents drive and you can take over for logins
team.mdOrg chart → prompt blockTelling every agent who is who and who to involve
sentinel.mdTrigger runtimeScheduling proactive agent work
models.mdModel referenceModels known to run with somora, per engine, with the config values that work and why
compaction.mdContext managementWhen and how a session is summarised, which model does it, what contextWindow controls per engine
thinking.mdReasoning depthPer-engine thinking levels, session overrides
sampling.mdSampling parameterstemperature, top_p and friends per model, agent and session
display.mdTUI togglesWhat the terminal client shows, and /queue
cache-strategy.mdPrompt-cache mechanicsWhy the system prompt is ordered the way it is
security.mdTrust modelNetwork posture, sandbox stance