somora docs — start here
A short guide so you don't have to read every page to get started. somora is a local-first gateway for personal AI agents with persistent memory. It runs as a long-lived service on your own machine, exposes an HTTP+SSE API plus three first-party clients (TUI, web desktop, mobile PWA), and lets you route conversations across Claude / ChatGPT / any OpenAI-compatible model — all without giving up ownership of the chat log, the memory layer, or your knowledge base.
What is somora — in one paragraph
You define agents (personas with their own model, memory inbox, and session). They come in two kinds: chat agents — the ones you talk to, that remember and coordinate — and builders, coding harnesses that take a repository, plan, build, test and report (builder.md). You talk to both from the TUI, the browser, or your phone. The server orchestrates the actual LLM call against the engine you configured (subscription or API), persists every message as JSONL, and runs a three-phase background dream system that turns raw session content into curated memory notes and a shared wiki over time. All of that stays on your machine.
Choose your path
You probably don't need every page below right away.
I just want to try it. Start with the project README quickstart — one line installs everything and starts the setup assistant, which connects a model, creates your first agent and sends it a test message. ~10 minutes.
I want to chat from the browser or my phone. Read setup.md → HTTPS via Tailscale (or run somora setup access, which does it for you) to get a real cert, then open https://<your-tailnet>.ts.net:18737/web/ in the browser. The mobile PWA is the same URL with /mobile/ — see mobile.md.
I want to add a new provider / model. setup.md → Configuring providers plus the comment block in config.example.yaml is the full reference.
I want an agent to build software. Create one of kind builder (builder.md): it plans in a repository, you press Go, it builds, tests and reports; another agent can hand it the order with the bundled builder-handover skill.
I want my agent to do real work — read files, run commands, search the web, edit notes. The tools overview lists the tool families. files.md, resources.md, and tmux.md cover the heavy ones. Skills are markdown how-tos you install to teach an agent multi-step recipes.
I want to decide what each agent may use. Tools and skills are shared by all agents; the web client's Abilities window (or agent.yaml) hides any of them for a given agent. mcp.md → Per-agent tool control and skills.md → Per-agent visibility.
I want my agents to make video. videogen.md — a render takes minutes, so the agent starts one and is woken when it is ready rather than holding its turn open. Off until videoGen is configured; verified against a self-hosted endpoint, the hosted providers are untested.
I want my agents to make images. imagegen.md — configure a model, then generate from the Media window or let an agent call image_generate. Results land in your workspace and show up in the chat.
I want long-term memory + a shared knowledge base. Read memory.md for the per-agent inbox, wiki.md for the shared layer, and dream-phases.md for how sessions automatically flow into both.
I want to build a third-party client / integration. api.md is the HTTP+SSE contract — same surface the built-in clients use.
Core concepts in one sentence each
- agent — a persona with its own
AGENTS.md, memory inbox, default model, and chat sessions. Lives at~/.somora/agents/<name>/. - session — a single chat thread with one agent. Persisted as JSONL, resumable, switchable mid-conversation.
- engine — the adapter that talks to an LLM backend. Four exist:
claude-cli(Claude subscription),codex-cli(ChatGPT subscription, bundled),grok-cli(SuperGrok subscription),openai-compatible(any/v1/chat/completionsendpoint — OpenRouter, Ollama, oMLX, LM Studio, …). - provider / model / alias —
config.yamldeclares providers (with baseUrl + apiKey if needed) and models on each. Analiaslets you refer to a model by short nickname anywhere. - memory — three layers an agent reads from: its own per-agent inbox (
~/.somora/agents/<name>/memory/*.md), the shared wiki, and the read-only Obsidian vault. Hybrid retrieval: vector + BM25. - wiki — a curated long-term knowledge base shared across all agents. Lives under your Obsidian vault. Written by the dream system, editable by hand.
- dream system — three background phases. REM turns finished sessions into memory-inbox notes (per-agent). Deep promotes high-value notes into the shared wiki. Lucid cleans up the wiki on a slower cadence. See dream-phases.md.
- project — an explicit pin of "files / paths / vault notes / research artifacts that belong to this chat session". Auto-injected into the agent's system prompt. See projects.md.
- resource — a configured SSH host.
file_*,exec, andtmuxtools dispatch against it viatarget=<resource>. See resources.md. - skill — a markdown how-to (agentskills.io format) installed at
~/.somora/skills/<slug>/SKILL.md. The agent loads the body on demand via theskilltool when it recognizes the situation. - tools — the things an agent can call. Memory reads, wiki edits, file I/O, shell exec, tmux sessions, web search/fetch, sub-agent spawning, etc. See tools.md.
- voice — two separate features that share only the word. Dictation and spoken replies: a mic button on web/mobile, optional TTS for the answer, plus a generic
/voice/turnaudio-in/audio-out endpoint, all on the OpenAI-compatible audio shape with no extra credentials (voice.md). And realtime voice: a standing, interruptible call with an agent, where a second model does the talking and the agent does the knowing (realtime-voice.md).
Where to next
If you've installed somora and chatted with the default agent, the high-value next steps in rough order:
- Create a second agent with a different persona — copy the default
~/.somora/agents/<name>/AGENTS.md, edit, restart. - Configure a real provider if you started on the subscription defaults — see setup.md → Configuring providers.
- Enable the wiki by pointing somora at your Obsidian vault in
config.yaml. The dream system then has somewhere to consolidate memory into. wiki.md. - Install your phone as a PWA for chatting from anywhere on your tailnet. mobile.md.
- Add an SSH resource so an agent can read files and run commands on a remote machine. resources.md.
- Pin a project to a session so the agent has a persistent working set. projects.md.
Reference docs
| Doc | Reads like | When to open it |
|---|---|---|
| setup.md | Operator runbook | First install; adding providers; HTTPS; mobile; ops |
| agents.md | Persona + agent.yaml reference | Creating an agent; per-agent model, REM, tool/skill visibility |
| builder.md | The builder kind | An agent that is a coding harness: plan → Go → build, task panel, questions, hand-over from an orchestrator |
| lsp.md | Language servers | A builder's writes come back with the compiler's errors; install, config, routes |
| api.md | HTTP/SSE contract | Building a third-party client; debugging streaming |
| tools.md | Tool catalog overview | Wondering what an agent can do |
| mcp.md | External MCP servers | Plugging third-party MCP tools into your agents |
| files.md | File-tool deep dive | Working with file_read/file_write/analyze_file |
| memory.md | Concept + mechanics | Tuning recall, writing notes by hand |
| wiki.md | Concept + mechanics | Shared knowledge base, hand-curation |
| dream-phases.md | Background workers | REM/Deep/Lucid cadence + triggers |
| projects.md | Concept + workflow | Pinning a working set to a session |
| resources.md | SSH-target config | Adding a remote machine |
| skills.md | Markdown skill format | Writing or installing skills |
| tmux.md | Multi-turn shell sessions | Driving long-running CLIs from agents |
| web.md | Browser client | Web-UI specifics + HTTPS notes |
| mobile.md | Mobile PWA | iOS/Android install, scope |
| voice.md | Dictation + spoken replies | Press-to-talk in web/mobile, optional TTS answers |
| realtime-voice.md | Talking to an agent | A standing, interruptible call; the voice talks, the agent knows |
| imagegen.md | Text-to-image | Generating images from the web app or an agent |
| videogen.md | Text-to-video | Job-based renders, and how an agent gets its result without waiting |
| browser.md | Shared browser | A Chromium per agent profile that agents drive and you can take over for logins |
| team.md | Org chart → prompt block | Telling every agent who is who and who to involve |
| sentinel.md | Trigger runtime | Scheduling proactive agent work |
| models.md | Model reference | Models known to run with somora, per engine, with the config values that work and why |
| compaction.md | Context management | When and how a session is summarised, which model does it, what contextWindow controls per engine |
| thinking.md | Reasoning depth | Per-engine thinking levels, session overrides |
| sampling.md | Sampling parameters | temperature, top_p and friends per model, agent and session |
| display.md | TUI toggles | What the terminal client shows, and /queue |
| cache-strategy.md | Prompt-cache mechanics | Why the system prompt is ordered the way it is |
| security.md | Trust model | Network posture, sandbox stance |