Skip to content

Setup ​

End-to-end install for somora as a long-running service on your machine. The short way is the installer plus the setup assistant (next section); the numbered sections after it are the same steps by hand, with the background for each. A short dev-from-checkout section sits at the bottom for contributors.

The short way: installer + assistant ​

bash
curl -fsSL https://somora.ai/install.sh | bash

Run it as the user who will own somora — not as root; the agents get that user's rights. What it does, each step skipped when already in place:

StepWhat happensNeeds admin rights
System packagestmux, ripgrep, git, a C/C++ compiler, python3 — through apt, dnf, pacman, zypper or Homebrewyes, asked first; without them the compiler is the only hard stop
Node.jskept when ≥22.13; otherwise Node 24 system-wide (NodeSource / Homebrew) or, without admin rights, the official build into ~/.local/share/somora/node (checksum-verified)only for the system-wide variant
npm folderwhen npm's global folder is not writable for you, it moves to ~/.npm-global and is added to your PATHno
somoranpm install -g somorano
ServiceLinux: systemd user unit, enabled at boot, lingering on so it survives logout. macOS: LaunchAgent, starts at every loginno (lingering may ask)
Assistantsomora setupno

Options go after bash -s --, or as environment variables:

bash
curl -fsSL https://somora.ai/install.sh | bash -s -- --no-setup     # stop before the assistant
curl -fsSL https://somora.ai/install.sh | bash -s -- --version 2026.930.1
curl -fsSL https://somora.ai/install.sh | bash -s -- --yes --no-sudo  # no questions, no admin rights

--yes (SOMORA_YES=1), --no-setup, --no-service, --no-sudo, --version <v> (SOMORA_VERSION). https://somora.ai/install.sh always serves the script of the latest release; the same file is attached to every GitHub release as install.sh.

somora setup — the assistant ​

bash
somora setup            # all steps
somora setup access     # one step: models | search | agent | memory | team | access | start
StepWhat it does
modelsClaude subscription (installs Claude Code if missing, runs its login), ChatGPT subscription (the bundled Codex login, browser or device code), or your own OpenAI-compatible server (asks the address, lists its models, asks the context window). Writes the providers: block with the tested settings from models.md.
searchWeb search: asks for a Brave Search API key (free plan, 2,000 searches a month), checks it with one real search and stores it under web.brave.apiKey — that switches the agents' web_search tool on.
agentCreates the first agent: name, how it addresses you, answer language, model, backup model. On an existing install it lists the agents and offers to repair one whose model no longer exists.
memoryTurns on REM per agent (with its own model), the duplicate check for new notes, and the shared wiki with Deep and Lucid — in a new folder or an existing Obsidian vault (dream-phases.md, wiki.md).
teamWith two or more agents: writes the first team.yaml (team.md).
accessThis machine only, local network, or Tailscale HTTPS: installs and connects Tailscale if needed, walks you through the one switch in the Tailscale admin page, fetches the certificate and turns on automatic renewal.
startStarts (or, after asking, restarts) the service, waits until it answers, sends the agent a real test message and says which model replied.

The assistant never rewrites a file wholesale: comments and your own settings in config.yaml / agent.yaml stay, the previous version is kept next to the file as <name>.bak-setup-<date>-<time>, and a result the server could not load is not written at all.

1. System prereqs ​

Hard requirements:

ToolWhyInstall
Node ≥22.13runtime + native node:sqlite; every somora command refuses to start on an older Node and prints the upgrade steps, somora update checks the target release before buildingnodejs.org or nvm install 22
tmuxthe tmux tool + the web tmux appsudo apt install tmux (Debian/Ubuntu) · brew install tmux (macOS) · sudo dnf install tmux (Fedora)
ripgrep (rg)needed by file_searchsudo apt install ripgrep · brew install ripgrep · sudo dnf install ripgrep
gitclone the repousually pre-installed; otherwise package-manage

Optional, per feature:

ToolWhyInstall
Chromium / Chromethe shared browser tool (browser.md) — auto-detected, or set browser.executablePathsudo apt install chromium · brew install --cask chromium · sudo dnf install chromium
Xvfbbrowser.headed: true on a host without a display (a virtual screen so Chromium runs with a window and does not announce itself as headless)sudo apt install xvfb · sudo dnf install xorg-x11-server-Xvfb · sudo pacman -S xorg-server-xvfb · macOS: not needed

C/C++ toolchain — npm install builds two native modules (better-sqlite3 for SQLite, @huggingface/transformers's ONNX runtime for local embeddings). Both ship prebuilt binaries for the common platforms; if your build falls through to source you'll need build-essential / Xcode CLT / equivalent.

2. At least one LLM backend ​

Pick one or more. somora doesn't care which you use, but at least one must be installed and authenticated before the first chat (§5).

Claude (recommended) — Anthropic's Claude Code CLI uses your Claude subscription, no API key needed:

bash
curl -fsSL https://claude.ai/install.sh | bash    # installs to ~/.local/bin/claude
claude auth login

ChatGPT — the Codex engine uses your ChatGPT subscription. Codex is bundled with somora (no separate install), so this login step comes after §4 Install somora:

bash
somora codex login   # ChatGPT Plus/Pro/Business; an existing `codex login` is picked up too

Local models — Ollama, LM Studio, vLLM, or oMLX. Install separately per their docs and have a /v1/chat/completions endpoint reachable on some http://host:port. somora talks to them via the openai-compatible engine — see §6 Configuring providers for the config block.

ToolUnlocksInstall
Obsidianthe wiki layer + read-only vault recallobsidian.md
TailscaleHTTPS for the web client (lifts the 6-connection browser limit, unlocks mic/screenshot APIs)tailscale.com

Without Obsidian: somora still works (memory inbox per agent, sessions, all tools), you just lose the long-term shared wiki layer.

Without Tailscale: the web client falls back to plain HTTP/1.1 and is single-window-only; the TUI is unaffected.

4. Install somora ​

bash
npm install -g somora

The package carries the built web clients. One native module (node-pty) is compiled during the install on Linux — that is what the compiler from §1 is for. The install is about 1.4 GB: somora itself is 20 MB, the rest are the bundled Codex and Claude engines (each a complete program for your platform), the embedding runtime for the memory search, the image library and the terminal module. The embedding runtime would also download a 300 MB CUDA library on Linux that somora never uses (embeddings run on the CPU); the installer and somora update skip it with ONNXRUNTIME_NODE_INSTALL=skip — set the same variable when you run npm install -g somora by hand.

If npm answers EACCES, its global folder belongs to root. Give npm a folder of your own instead of reaching for sudo:

bash
npm config set prefix ~/.npm-global
echo 'export PATH="$HOME/.npm-global/bin:$PATH"' >> ~/.profile && source ~/.profile

5. Start the server + chat ​

bash
somora init            # creates ~/.somora/ + writes the systemd user-service unit
somora server start    # starts the unit and enables it at boot
somora tui             # opens the TUI against the running server

somora init is idempotent and one-time: it creates ~/.somora/ (config + lockfile + logs land here) and writes the systemd unit at ~/.config/systemd/user/somora.service with ExecStart baked to your freshly installed global binary. From then on, somora server start just brings the unit up; the server runs in the background and survives logout / reboot. To stop / restart / inspect:

bash
somora server stop
somora server restart
somora server status
journalctl --user -u somora -f      # tail logs

macOS has no systemd; the same three commands work there through a per-user LaunchAgent (~/Library/LaunchAgents/ai.somora.server.plist, written by somora init). somora server start loads it, and from then on macOS starts somora at every login and brings it back if it crashes; somora server stop unloads it until the next login or start. Its output goes to ~/.somora/logs/launchd.log. A LaunchAgent never runs before its user has logged in — on a Mac used as a server, turn on automatic login (System Settings → Users & Groups).

If you don't want a background service (in a container, or while debugging):

bash
somora server start --foreground    # blocks the terminal; Ctrl-C to stop

Updating somora ​

bash
somora update                # the current release
somora update --edge         # the newest build on npm, incl. pre-releases
somora update 2026.930.1     # a specific version

somora update asks npm for the target version, checks that your Node.js is new enough for it, installs it (npm install -g somora@<v>), re-runs somora init so the systemd unit's ExecStart points at the freshly installed binary, then restarts the service. When you are already on that version it does nothing (--force reinstalls). Pass --no-reinit to skip the unit rebake if you've hand-edited somora.service.

Version numbers are dates: 2026.930.1 is the first build of 30 September 2026, 2026.1005.2 the second of 5 October. Versions up to 2026.09.29.12 (four parts) predate the npm package and were installed from a git checkout; their somora update still works and brings you onto the npm package.

Recovering an upgrade that didn't take effect ​

The systemd unit's ExecStart is baked in by whichever copy of somora ran somora init. If you initially ran init from a git checkout (common for early adopters), the unit pins to that checkout path:

ExecStart=/home/<you>/somora/bin/somora.mjs server start --foreground

A subsequent npm install -g <tarball> then has no effect on the running service — systemd keeps launching the checkout binary, which serves its old web/dist. Symptom: new features missing in the web UI even after install + restart.

To spot it:

bash
systemctl --user cat somora.service | grep ExecStart
# good:  ExecStart=/home/<you>/.npm-global/lib/node_modules/somora/bin/somora.mjs ...
# bad:   ExecStart=/home/<you>/somora/bin/somora.mjs ...   (← checkout, not global)

somora update re-bakes the unit automatically. If the unit still points at a checkout, run the manual fix once:

bash
"$(npm root -g)"/somora/bin/somora.mjs init
systemctl --user daemon-reload
systemctl --user restart somora.service
# or, for a config.yaml edit that touches models, providers, caps, tools:
# the gear menu in the web taskbar → Reload config, or /reload in the TUI
curl -ks https://<host>:18737/version

Hard-reload the browser (Ctrl+Shift+R / Cmd+Shift+R) after so it doesn't serve a cached JS bundle from before the upgrade.

On first start somora creates ~/.somora/:

~/.somora/
├── config.yaml                ← server config (created with sane defaults)
├── agents/
│   └── default/               ← seed agent created on first run; rename / customize
│       ├── AGENTS.md
│       ├── SOUL.md
│       ├── USER.md
│       └── agent.yaml
├── index/
│   └── shared.db              ← vault + wiki retrieval index, one per instance (derived, rebuilt if deleted)
└── logs/
    └── server.YYYY-MM-DD.1.log

The log rolls daily. You do not need a shell to read it: the log tile in the web client shows the end of a day's file with filters for level, agent and text (web.md).

Optional: tell the agents who is who ​

With more than one agent, write ~/.somora/team.yaml (start with somora team init --principal "<your name>") so every agent gets the org chart and "who to involve for what" in its prompt. See team.md.

6. Configuring providers ​

Edit ~/.somora/config.yaml. The shipped default has just one provider (Anthropic via Claude Code subscription) and is enough to chat with Claude. Add more providers as needed.

Anthropic via Claude Code subscription (no API key) ​

yaml
providers:
  anthropic:
    engine: claude-cli
    models:
      - id: claude-opus-5-5
        alias: opus
        contextWindow: 1000000
        capabilities: [text, image, pdf, reasoning]

Requires the Claude Code CLI binary at ~/.local/bin/claude (or set SOMORA_CLAUDE_BIN env to its path). Auth is handled by the binary — claude login once and somora rides on the resulting session.

OpenAI via Codex CLI subscription (no API key) ​

yaml
providers:
  openai:
    engine: codex-cli
    models:
      - id: gpt-5.6-terra
        alias: terra
        contextWindow: 258400        # the Codex session window, not the 1.05M API window
        capabilities: [text, image, pdf, reasoning]

contextWindow: 258400 is deliberate: Codex runs a session against a window it delivers itself and reports on every turn (modelContextWindow, 258,400 for every model it offers), and on a CLI engine the value does not trigger somora's compaction anyway — it only decides whether the model is picked as a compaction worker and what the header percentage claims. See compaction.md and models.md.

Codex is bundled — @openai/codex is an exact-version dependency of somora and runs as an app-server per turn; a global codex on the host is ignored (SOMORA_CODEX_BIN remains as a debugging override). Auth: somora codex login (ChatGPT Plus/Pro/Business). somora keeps its own Codex home at ~/.somora/codex-home and mirrors auth.json from ~/.codex on every turn, so a login done with a global Codex CLI works as well. somora codex debug models shows the model catalog the bundled version sees.

somora's tools reach Codex as dynamic tools. codexCli.directTools (config.yaml) names the tools kept in the model's direct tool list every turn; everything else is deferred and found via Codex tool search (or ALL_TOOLS inside Code Mode on the GPT-5.6/GPT-6 models). The default is the everyday core: memory_, file_, exec, tmux, web_, time_now, agent_ask, agent_ask_result, spawn_subagent, subagent_result, somora_docs_.

xAI via Grok Build CLI subscription (no API key) ​

yaml
providers:
  xai:
    engine: grok-cli
    models:
      - id: grok-4.5
        alias: grok
        contextWindow: 500000
        capabilities: [text, reasoning]

Requires the Grok Build CLI (grok) on PATH — installed via curl -fsSL https://x.ai/cli/install.sh | bash, which drops the binary at ~/.local/bin/grok. Override with SOMORA_GROK_BIN if it lives elsewhere.

Auth is handled by the binary: run grok login once, which writes a session to ~/.grok/auth.json. somora's adapter connects over ACP (Agent Client Protocol — JSON-RPC on stdio, grok agent stdio) and the handshake picks up that session as the cached_token auth method automatically. A SuperGrok / Premium subscription authenticates the CLI, not the xAI API — so this path uses your subscription, while pointing an openai-compatible provider at https://api.x.ai/v1 would bill a separate pay-per-token API account instead.

grok-4.5 reports three reasoning efforts (low / medium / high, default high), so somora's /thinking knob maps straight through. Give the model the reasoning capability to activate it. Note there's no "disabled" state — /thinking off maps to low.

Tools. somora's full MCP surface (memory, file_*, exec, wiki, subagents, …) is handed to the ACP session via session/new's mcpServers parameter, scoped to the current agent+session exactly like claude-cli and codex-cli. On top of that Grok Build brings its own file/shell tools, scoped to the working directory ($HOME).

Grok reaches MCP tools through a search_tool / use_tool indirection rather than listing all of them up front, which keeps a large surface cheap context-wise. The adapter unwraps that: a use_tool{tool_name:'somora__memory_list'} call is recorded as mcp__somora__memory_list, so session logs and tool rows match what the other engines emit. The search_tool probes themselves surface as-is.

Budget note: the tool catalogue is not free. A trivial two-tool turn measured ~76k input tokens on a fresh session, ~108k on the follow-up, with most of it served from cache (tokens_in_cached). Against grok-4.5's 500k window that's comfortable, and on a subscription it costs nothing extra — but it's worth knowing before pointing an API-billed provider at the same setup.

Sessions resume across turns via session/load against the grokSessionId stashed in session-meta.

Attachments. Grok Build exposes no attachment channel over ACP — the handshake reports promptCapabilities.image: false. Images and PDFs therefore never reach the engine: the capability gate in run-turn.ts refuses them first, since grok-4.5 declares neither image nor pdf, and the user gets "does not support image inputs" before a process is spawned. Text attachments pass the gate and are inlined into the prompt, the same way codex-cli handles its non-image attachments. Anything else that somehow arrives is named to the model as undeliverable and recorded as an engine_meta item of type attachments_unsupported — never dropped silently.

API failures surface as errors. xAI reports a spent balance or a blocked subscription on the proprietary _x.ai/* channel — a retry_state{type:'failed'} frame plus turn_completed with stop_reason: 'error' — not through the ACP error channel. The adapter reads both and emits a somora error event carrying the message (e.g. "API error (status 402 Payment Required): Grok Build usage balance exhausted"), so a configured fallback: model takes over. Replayed frames from session/load are ignored via _meta.isReplay, so a failure from an earlier turn cannot abort a resumed one.

Local OpenAI-compatible LLM (Ollama, LM Studio, vLLM, oMLX, ...) ​

For this engine contextWindow is the compaction wall — set it to the server's limit (--max-model-len, --context-length), not the model card's. Recommended blocks per model family, with sampling and reasoning vocabularies, are in models.md.

yaml
providers:
  local:
    engine: openai-compatible
    baseUrl: http://localhost:11434/v1   # adjust to your local server
    apiKey: dummy                         # most local servers ignore this
    models:
      - id: llama3.3:70b
        alias: llama
        contextWindow: 131072
        capabilities: [text]
      - id: gemma-3-27b-it
        alias: gemma
        contextWindow: 131072
        capabilities: [text, image]

Multiple OpenAI-compatible providers can coexist — give each one a unique name (local, lmstudio, office, …).

Reliability with smaller / local models ​

Smaller models (deepseek, kimi, and many local ones) drive the OpenAI-compatible tool loop less reliably than the big hosted models. Two well-documented failure modes show up: they re-issue the same tool call over and over without registering the result, and they sometimes echo the provider's internal tool-result template ("Use the results below to formulate an answer…") as their reply instead of answering. somora hardens this path so a weaker model degrades gracefully instead of flooding you:

  • Duplicate tool calls in one round are collapsed to a single execution — the model still gets a result for every distinct call.
  • A per-turn tool-call budget (agentLoop.maxToolCallsPerTurn, default 30) stops a runaway that the round cap can't see, and forces a clean final answer.
  • A budget notice at 75 % of either cap: one user-role message ("N rounds and M calls remain — wrap up") so a "read N things, then summarise" task ends on its own instead of at the hard stop.
  • The forced final answer is a user-role message, valid on every backend (strict chat templates such as vLLM + Qwen reject a trailing system message). If that call fails too, the turn returns a digest of the last tool results instead of only an error line.
  • Output guards detect a leaked template or a repeated-text loop in the stream, cut it before it floods the window, and force one clean no-tools answer.

None of this touches the claude-cli / codex-cli engines — they run their own loop. It also stays out of the way of capable models, which never trip these guards.

yaml
providers:
  local:
    engine: openai-compatible
    models:
      - id: some-strong-local-model
        alias: big
        contextWindow: 131072
        capabilities: [text]
        # Opt a trusted model back into parallel tool calls. Default is
        # sequential (one call per round) because weak models fan out
        # into large duplicate batches; a strong model doing independent
        # reads can safely parallelise.
        parallelToolCalls: true
      - id: some-local-reasoning-model
        alias: thinker
        contextWindow: 262144
        capabilities: [text, reasoning]
        # Output cap sent as `max_tokens`. Unset = not sent, and vLLM then
        # allows the whole remaining context — on a reasoning model, where
        # thinking and answer share that budget, nothing else stops a
        # runaway thinking phase. (Not the memory block's
        # `memory.autoInject.maxTokens`, which caps injected input.)
        maxTokens: 16384
        # Which words this model accepts for somora's thinking levels and
        # where they go in the request — see docs/thinking.md.
        reasoning:
          levels: { high: xhigh }
        # Vendor-recommended sampling for this model; agent.yaml and the
        # session override win per key — see docs/sampling.md.
        sampling:
          temperature: 1.0
          top_p: 0.95

agentLoop:
  maxToolCallsPerTurn: 30   # hard ceiling on tool calls per turn (openai-compatible)

thinkingContent:            # the model's reasoning text in the clients — docs/thinking.md
  capture: true             # false drops it at the server (no SSE, no JSONL)
  maxChars: 65536           # per-turn cap on what is persisted

7. Voice — STT + TTS (optional) ​

Voice is two independent toggles:

  • STT (speech-to-text) — turns the mic button on for the web and mobile-PWA chat. Tap, talk, tap again → transcript drops into the draft.
  • TTS (text-to-speech) — when you submit via mic and the per-chat auto-play toggle is on, somora generates spoken audio for the assistant reply and plays it. A Play-button on the bubble lets you replay.

Both proxy through OpenAI-compatible endpoints on a provider you've already configured. See voice.md for the full picture including the /voice/turn audio-in/audio-out endpoint for integrations.

Speech-to-Text ​

When enabled, the web chat input grows a mic button next to send. Click → record, click again → transcribe → text lands in the textarea ready to send. Audio is forwarded to an OpenAI-compatible STT endpoint (POST /v1/audio/transcriptions) and the result lands locally — never travels via a cloud API unless you configure one.

Supported upstreams: oMLX (mlx-audio, Apple Silicon — full audio API including TTS), faster-whisper-server, OpenAI Whisper API, and anything else exposing the OpenAI audio shape. Configure once, switch backends by editing one block.

yaml
stt:
  enabled: true
  provider: omlx                              # ← name of an entry in `providers`
  model: mlx-community/whisper-large-v3-turbo
  language: de                                # optional default hint (ISO 639-1)

stt.provider must reference an existing openai-compatible provider in your providers block — the STT call reuses its baseUrl and apiKey. The STT model is not listed in providers.<x>.models; keeping it separate keeps it out of chat-model pickers, agent-config validation, and /v1/models.

The web mic button auto-hides when:

  • the server reports stt.enabled: false (or the block is omitted)
  • the browser lacks MediaRecorder / navigator.mediaDevices.getUserMedia
  • the page wasn't loaded over a secure context (HTTPS — required by the mic-permission API in most browsers; via Tailscale, that's already satisfied)

mlx-community/whisper-large-v3-turbo is the practical sweet spot on Apple Silicon: ~800M params, near-real-time on M-series chips, multilingual incl. German, marginal quality difference vs. the full large-v3. If oMLX flags the model as missing the HuggingFace preprocessor/tokenizer files (MLX-converted repos sometimes ship weights only), drop the upstream config JSONs in alongside the weights:

bash
cd ~/.omlx/models/whisper-large-v3-turbo
for f in preprocessor_config.json tokenizer.json special_tokens_map.json \
         tokenizer_config.json generation_config.json; do
  curl -L -O "https://huggingface.co/openai/whisper-large-v3-turbo/resolve/main/$f"
done

Text-to-Speech ​

Mirror config for spoken replies. Same posture — proxies through an OpenAI-compatible TTS endpoint (POST /v1/audio/speech) on a provider you already have.

yaml
tts:
  enabled: true
  provider: omlx                              # ← references providers.omlx
  model: fish-audio-s2-pro-8bit               # whatever your upstream calls it
  language: de
  cache:
    retentionDays: 7                          # 0 disables GC
    maxSizeMB: 500
  reencode:
    enabled: true                             # needs ffmpeg on $PATH
    opusBitrateKbps: 24
  clients:
    web:
      autoPlayVoiceReplies: false             # initial toggle state for new sessions
      allowUserOverride: true
    mobile:
      autoPlayVoiceReplies: false
      allowUserOverride: true

Auto-TTS for the normal chat fires only when all four hold: tts.enabled is true, the user submitted via mic (input_modality=voice), the per-session 🔊 toggle is on, and the assistant text is speakable (no heavy code blocks / large tables). Otherwise the chat stays text-only. The 🔊/🔇 toggle in the chat header is sticky per session (localStorage).

Two flows are wired:

  • Normal chat auto-TTS — voice input → spoken reply on web + mobile.
  • POST /voice/turn — independent audio-in/audio-out endpoint for panel/satellite/bridge integrations; always generates audio regardless of toggles. See voice.md.

System dependency: ffmpeg on $PATH if you want tts.reencode.enabled (opus/m4a output). Without ffmpeg, set reencode.enabled: false and clients receive plain WAV (larger, but works).

Tunables ​

These all live in config.yaml with conservative defaults; uncomment to override. Full schema in src/config/types.ts.

yaml
promptBudgets:                # soft caps for static prompt text — warnings only, nothing is truncated
  teamBlockChars: 3000        # the "# Your team" block per agent (docs/team.md)
  personaFileChars: 8000      # each of AGENTS.md / SOUL.md / USER.md
  personaTotalChars: 14000    # the three together — shown in the web Agent window

realtimeVoice:                # talking to an agent (docs/realtime-voice.md)
  enabled: false              # off until a realtime provider and a key exist
  provider: openai            # openai | google | local
  model: gpt-realtime-2.1-mini
  apiKeyFile: ~/.somora/secrets/openai-realtime.key   # a FILE, chmod 600 — never the key in config
  # url: ws://127.0.0.1:8787/realtime   # where to connect; omitted = OpenAI. Set it (and
  #                             # provider: local) to use a service of your own that speaks
  #                             # the same protocol — same adapter, no key needed on localhost
  defaultVoice: alloy         # ten exist: alloy ash ballad coral echo sage shimmer verse marin cedar
  consultPolicy: always       # auto | substantive | always — when the speaking model must ask the real agent
  consult:
    quickAnswerMs: 8000       # how long a spoken question waits on the agent (1000..120000);
                              # past it the voice says it handed the request over and keeps
                              # talking, the answer is read out when it lands
  maxCallMinutes: 20          # hard stop; a standing call bills while nobody talks
  allowAgentSwitch: false     # move a call to another agent or session mid-conversation
  turnDetection: { threshold: 0.4, prefixPaddingMs: 200, silenceDurationMs: 420 }
  # Separate from `stt`/`tts`: those are dictation and spoken replies
  # (docs/voice.md). A call needs a realtime-capable provider; a normal
  # chat-completions endpoint cannot carry one.

compaction:
  triggerRatio: 0.8           # fraction of the input budget (window minus the answer)
  safetyCushionPairs: 4       # most-recent turns kept uncompacted
  # modelOverride: opus       # force a specific compaction worker model
  # workers: [big-local, small-local]   # who may summarise, in the order they are tried
  # The trigger works off the token count the provider reported for the
  # last request, and off somora's own ESTIMATE only until there is one
  # (new session, model switch, provider without usage). When the backend
  # nevertheless rejects a prompt as too long (400 "Prompt too long",
  # "maximum context length", oMLX's prefill memory guard — typical after
  # switching a long session from a 1M-window model to a 131k one), the
  # openai-compatible engine forces a compaction down to the last
  # exchange and retries the turn once; a second refusal surfaces as a
  # plain-language error (switch model or /reset) instead of the raw 400.
  # Compaction workers are picked from models whose engine has a one-shot
  # path (claude-cli, codex-cli, openai-compatible): the smallest
  # contextWindow that fits the range × 1.3 — which can be a
  # subscription-backed CLI model. `workers:` replaces that pick with an
  # ordered list of your own; a model that is not listed never
  # summarises. Either way up to three workers are asked per compaction,
  # so one refusing (busy host, rate limit, a route being reloaded)
  # costs an attempt, not the compaction. Mechanics, and what
  # contextWindow means on each engine, in docs/compaction.md.

agentLoop:
  maxRounds: 8                # tool-call rounds per turn (openai-compatible)
  toolCallTimeoutMs: 30000    # per-tool-call timeout for fast tools (memory, web, file, time)
  longTaskDefaultTimeoutMs: 300000   # 5 min — slow A2A tools (agent_ask, subagent_result
                              # wait_until_done) when the caller passes no timeout_ms;
                              # also honoured inside the MCP child of the CLI engines
  longTaskMaxTimeoutMs: 1800000      # 30 min — hard ceiling for those, even with an explicit
                              # timeout_ms; past it the tool answers state "pending", the
                              # work keeps running. claudeCli.mcpToolTimeoutMs and
                              # codexCli.toolTimeoutSec must be >= this.
  execMaxConcurrentPerAgent: 8       # background exec jobs one agent may hold
  execMaxConcurrentGlobal: 32        # … across all agents
  wakeGraceMs: 3000           # when work an agent started and walked away from
                              # finishes (a late agent_ask answer, a background
                              # sub-agent, a rendered video), the agent is woken in
                              # the session it asked from — unless it fetched the
                              # result within this grace. 0..60000.
  toolUsageReminder: true     # short "call tools, don't narrate" block in the
                              # system prompt whenever the agent has tools.
                              # Tools reach the model through a separate API
                              # field, never through prompt text; smaller local
                              # models benefit from being told so explicitly.
                              # Constant text — one cache invalidation on
                              # rollout, none afterwards.

# Per-engine idle-event watchdog. If an engine produces no events
# (assistant_delta, tool_call, …) for this duration mid-turn, the
# turn is aborted so the per-session lock releases and the user
# sees a clean error instead of all agents looking dead. Dream
# workers (Deep/Lucid) bypass this — they run on their own path.
#
# While a tool call is in flight the threshold is automatically
# relaxed to the MCP tool timeout (claudeCli.mcpToolTimeoutMs /
# codexCli.toolTimeoutSec, 30 min by default), so a legitimately
# long-blocking tool (agent_ask, subagent_result wait_until_done)
# isn't cut off by the much shorter idle window — a genuinely dead
# child is still caught at the tool-timeout horizon.
# How long a model that was unreachable stays out of every cascade —
# the chat fallback chain, the REM worker chain and the compaction
# workers. The first cascade that hits the outage writes the model
# down; later turns start past it instead of paying the dead hop
# (~30 s each) again. A success clears the note early; a config
# reload or POST /models/availability/reset clears all of them.
fallback:
  retryUnavailableMinutes: 60

engineWatchdog:
  claudeCliIdleMs: 300000        # 5 min — subscription, fast first event
  codexCliIdleMs: 300000         # 5 min — subscription, fast first event
  grokCliIdleMs: 300000          # 5 min — same class as the other CLIs
  openaiCompatibleIdleMs: 1200000 # 20 min — local LLMs can stream slowly;
                                  # raise if your backend regularly silences
                                  # for longer than 20 min mid-turn

# Per-subscriber write budget for SSE broadcasts. A healthy writeSSE
# finishes in microseconds; a wedged subscriber (mobile browser
# backgrounded, TCP receive window stuck at 0, dead-but-not-closed
# stream) can otherwise stall every following turn on that session
# until server restart. This is NOT a per-turn timeout — long-running
# tool calls and slow local LLMs are unaffected, because each
# individual event-write is still microseconds.
sse:
  publishTimeoutMs: 10000         # 10 sec per single event-write; evict the
                                  # subscriber on overrun and continue. Healthy
                                  # writes never hit this; 2 orders of magnitude
                                  # more than a normal write ever takes.
  publishParallel: true           # broadcast in parallel — one slow client
                                  # never blocks the others. Flip to false only
                                  # if you need strict serial delivery order.
  heartbeatMs: 20000              # comment-frame heartbeat on every stream;
                                  # clients treat > ~2 missed as a lost link.
  deadAfterMs: 60000              # a subscriber whose heartbeat write has not
                                  # completed for this long is dead: evicted
                                  # (`sse.publish_evict_dead` in the log), socket
                                  # destroyed. Catches vanished tabs / stuck
                                  # TCP windows that never send FIN.
  h2PingIntervalMs: 30000         # HTTP/2 PING per client session (TLS listener) …
  h2PingTimeoutMs: 30000          # … no ACK within this → session destroyed
  keepAliveDelayMs: 30000         # TCP keepalive on every socket

memory:
  embedding:
    provider: local           # 'local' uses @huggingface/transformers (ONNX)
    model: all-MiniLM-L6-v2   # alias or full HF repo path
  chunking:
    targetTokens: 400
    overlapTokens: 80
  autoInject:
    queryTurns: 3             # current message + last-(N-1) turns as context
    maxResults: 5             # top-N hits injected per turn
    minScore: 0.35            # discard hits below this score (0..1)
    maxTokens: 1500           # hard cap on the injected memory block
    historyWeight: 0.3        # how much the previous turns steer the vector query
    historyWeightShort: 0.55  # … for a message with only 1–2 content words
    historyWeightEmpty: 0.8   # … for a message with none ("das solltest du wissen?")
    historyTurnChars: 800     # head of each previous turn used for the blend
    shortQueryBm25Weight: 0.5 # BM25 share for a 1–2-word question; null = hybrid default
  hybrid:
    vectorWeight: 0.7
    bm25Weight: 0.3
    slugMatchBoost: 1.5       # boost a page whose slug names a query word (1 = off)

# Projects (opt-in, off by default) — pointer-file manifests binding
# a chat session to a real-world thing (Obsidian notes, code dirs,
# URLs, remote-resource paths). When enabled, agents see six tools
# (entity_list, project_list, project_get, project_create,
# project_update, project_focus), the chat-header gets a project chip,
# and slash commands /projekt + /projects activate. See
# docs/projects.md for the full model.
#
# `entities` is a CURATED VOCABULARY — projects belong to one entity
# (e.g. "privat", "acme"), and the agent must pick from this list
# at create time. Prevents STT mishearings from inventing phantom
# entities and gives you a free filter axis ("list all private
# projects"). Agents cannot extend this list via tools.
# projects:
#   enabled: true
#   entities:
#     - slug: privat
#       label: Privat
#     - slug: acme
#       label: acme GmbH
#     # add as many as you need — these are YOURS to curate

Engine-meta — codex todo_list ​

codex-cli (GPT-5.x and codex models) tracks an internal plan/checklist while it works. Codex emits item.completed events with itemType: "todo_list" every time the model marks a task done or adds a new one. somora persists these to the session JSONL as engine_meta records — they're available to:

  • The chat UI when show.tools is enabled (web + TUI). Renders as a dimmer block with a ◌ codex · plan prefix to visually differentiate from real tool calls. Expand to see the task list with ✓ / → / ○ glyphs per status.
  • The REM dream-worker, which scans session history to extract memories. Codex plans appear in that history, so the agent can retain "what was on my list yesterday" implicitly.
  • The session export (?format=markdown), where plans render as GitHub-style task lists.

Mobile PWA hides engine_meta entirely (mobile is intentionally a text-only minimalist surface). The other engines emit their own engine_meta rows through the same mechanism — a forced compaction, a dropped sampling key or an adjusted reasoning effort on openai-compatible, an undeliverable attachment on grok-cli, a restarted session on any CLI engine.

Friendly labels live in src/engine/engine-meta-labels.ts — a tiny engine → itemType → label map. Unknown itemTypes fall back to the raw string so future codex/SDK additions appear immediately, just with a less-pretty label. No configuration needed.

The wiki layer enables long-term shared knowledge across all agents. Requires an Obsidian vault. See wiki.md for the full mental model.

yaml
obsidian:
  vault: ~/Documents/Vault     # required for the wiki to work

wiki:
  enabled: true
  vaultSubfolder: somora       # → <vault>/somora/ becomes the wiki
  language: de                 # de | en — section headings, default folders,
                               # index/log wording, prose language (docs/wiki.md)

  deep:                        # Memory→Wiki consolidation
    enabled: true
    intervalHours: 12
    model: opus                # via claude-cli, subscription

  lucid:                       # Wiki cleanup
    enabled: true
    intervalDays: 7
    model: opus
    requireApproval: true
    maxCallsPerTurn: 3         # cap on wiki_* tool invocations during
                               # an active Lucid review-loop, per user
                               # turn. Forces per-page user confirmation
                               # for bigger plans. Raise (e.g. 5-10) if
                               # you regularly OK multi-page batches
                               # and the default 3 cuts off legitimate
                               # work. Resets on every user message.

  search:
    boostWiki: 1.4             # wiki hits rank above memory in retrieval
    boostMemory: 0.85
    boostVault: 0.65
    overviewMaxChars: 4000     # wiki-overview block in the system prompt;
                               # snapshotted once per session, so this is
                               # paid once inside the cached prefix
    overviewTopNSlugs: 30      # max sections listed when even the bare
                               # page list exceeds the budget

Per-agent REM (session→memory extraction) is configured in each agent.yaml, not here — see agents.md.

Scheduler state files — Deep and Lucid persist their cadence to ~/.somora/dream-state/{deep,lucid}.json so server restarts don't reset the timer. Each file holds the last started / completed / failed timestamps; the worker reads these at boot and schedules the next run at lastCompletedAt + interval, with a 60 s startup grace when the run is already overdue. Fresh installs get an lastCompletedAt = now anchor on first boot so restart-storms before the first auto- fire don't starve out the schedule.

You normally never touch these files. If you want to force the next auto-Deep/Lucid to fire sooner without manually triggering it, edit lastCompletedAt to an older timestamp (or delete the file — the bootstrap anchor is rewritten on the next start). If you want to pause auto-firing without disabling the worker, set lastCompletedAt to a future timestamp.

Sentinel — proactive triggers (optional) ​

The trigger runtime that wakes agents on a schedule (see sentinel.md). One configuration knob:

yaml
sentinel:
  completedRetentionDays: 7   # default; 0 disables auto-cleanup

completedRetentionDays controls how long one-shot at-triggers stay in the registry after they've fired. The scheduler sweeps at boot and on each daily re-arm tick; older completed entries get auto-deleted along with their history file. Set to 0 to disable auto-cleanup entirely (manual sentinel delete only). Recurring triggers don't auto-GC — they end up in paused or error and you choose when to remove them.

The daily update check — what somora.ai sees ​

somora runs on your machine and talks to the model providers you configured. There is exactly one request it makes on its own: once a day the server asks somora.ai whether a newer version exists, and the answer shows up as an update notice in the web client's taskbar, in somora server status and in the server log (update.available). somora update itself keeps using npm.

http
GET https://somora.ai/api/latest-version
User-Agent: somora/2026.1001.2 (linux; node/22.23.2; x64; server)

That is the whole request: no body, no install identifier, no machine id, no random token — the version, operating system, Node.js version, CPU and whether the server or the CLI asked. The answer is {"version": "…", "note": "…"}; the note is an optional sentence from us to everyone ("update Node first"), at most 500 characters.

What we do with it. Like any web server, somora.ai logs the request together with its IP address and the approximate location Cloudflare derives from it (country, region, city). From those logs we count how many installations ask — per day, per version, per operating system — to know whether the project is used and what to test on. The raw logs are deleted after 90 days, daily totals are kept, nothing is published and nothing is passed on. Counting by IP address is approximate by nature: two instances behind one router look alike, and many home connections change their address.

Switching it off. Any of these stops the request entirely:

  • DO_NOT_TRACK=1 in the service's environment (the common convention), or
  • updateCheck.enabled: false in config.yaml, or
  • a CI variable in the environment — a test pipeline is not an installation.

somora telemetry show prints the request as it would be sent, why the check is on or off, and when it last ran. updateCheck.endpoint points the check at a mirror of your own. The last answer lives in ~/.somora/update-check.json.

The first check runs one to six minutes after the server starts (a random delay, so a fleet restarting together does not knock in unison); after a success the next one is a day later, after a failure an hour later. A failed check never affects anything else.

HTTPS (Tailscale) — required for the web client at scale ​

The web client opens one persistent SSE connection per chat window. Browsers cap HTTP/1.1 at 6 concurrent connections per origin, so a multi- agent setup (plus tmux session attaches) hits that ceiling fast — symptom: new chat tabs silently fail to send, agents seem unresponsive.

Solution: serve somora over HTTP/2-over-TLS. HTTP/2 multiplexes every stream over a single TCP connection, lifting the limit entirely. The same upgrade also unlocks secure-context-only browser APIs that the roadmap depends on (mic / screenshare / clipboard write / push notifications / service workers).

somora's blessed path is Tailscale. Tailscale issues publicly-trusted Let's Encrypt certs for your tailnet's *.ts.net hostnames, free, with a single command — Node + every browser accept them with no warnings, no CA installs, and no manual cert pinning. If you're not on Tailscale you'll need to wire up your own cert (mkcert for LAN-only, or a real DNS-validated LE cert) — same config block, different acquisition.

Set up TLS via Tailscale ​

  1. Install Tailscale on the somora host (sudo tailscale up).

  2. In the Tailscale admin DNS panel, enable MagicDNS and HTTPS Certificates (one-time tailnet setting).

  3. Generate certs into ~/.somora/certs/:

    bash
    mkdir -p ~/.somora/certs
    cd ~/.somora/certs
    tailscale cert <your-host>.<your-tailnet>.ts.net

    tailscale status shows your hostname; the FQDN is <host>.<tailnet>.ts.net. The command writes <fqdn>.crt (cert) and <fqdn>.key (private key, mode 0600).

  4. Reference them in ~/.somora/config.yaml:

    yaml
    server:
      host: 0.0.0.0        # REQUIRED for remote clients (default 127.0.0.1 = loopback only)
      port: 18737
      tls:
        cert: ~/.somora/certs/<your-host>.<your-tailnet>.ts.net.crt
        key:  ~/.somora/certs/<your-host>.<your-tailnet>.ts.net.key
        publicHost: <your-host>.<your-tailnet>.ts.net

    host: 0.0.0.0 is what makes the server reachable from other machines (LAN / Tailscale); the default 127.0.0.1 binds loopback-only. Keep this in config.yaml (not as a SOMORA_HOST env in the systemd unit) — config survives somora update, unit env does not (the update rebakes the unit from a template). If you do need custom systemd env, put it in a drop-in (~/.config/systemd/user/somora.service.d/*.conf) — drop-ins survive the rebake; somora init also carries forward existing Environment=/EnvironmentFile= lines and prints what it preserved.

    publicHost MUST match the cert subject — strict TLS verification is on. Internal MCP-child callers (subagent fallback in src/tools/agents/spawn.ts) read this hostname from env at server startup and use it for their own HTTPS callbacks; there is no loopback bypass, everything goes through the one secure listener.

  5. Add renew: tailscale to the block (see Cert renewal) and restart somora. Connect with the full URL: https://<your-host>.<your-tailnet>.ts.net:18737/web/. The :port part is required because somora doesn't run on 443.

Cert renewal ​

Tailscale certs are valid for ~90 days. Add one line and somora takes care of them:

yaml
server:
  tls:
    cert: ~/.somora/certs/<your-host>.<your-tailnet>.ts.net.crt
    key:  ~/.somora/certs/<your-host>.<your-tailnet>.ts.net.key
    publicHost: <your-host>.<your-tailnet>.ts.net
    renew: tailscale

Twice a day the server asks Tailscale for a certificate that is good for at least 30 more days (tailscale cert --min-validity 720h; it only re-issues when the current one is closer to its end) and loads the new pair without a restart — running turns are not interrupted. For that, tailscale cert must work for the user somora runs as:

bash
sudo tailscale set --operator=$USER     # once

somora setup access does both for you. Without renew: the server still watches the two files: renew them any way you like and the new certificate is served at the next check (log line server.tls.reloaded). A failed renewal is logged as server.tls.renew_failed with the reason; from 14 days before the end server.tls.expires_soon warns on every check.

Without Tailscale ​

Drop in any cert + key — the config doesn't care about the issuer, only that the file paths point at valid PEM. For local LAN dev, mkcert is the cleanest non-Tailscale option (mkcert -install once per device, then mkcert <host>.local 192.168.x.y). Self-signed (without an installed CA) will work but every browser will scream — only acceptable for single-developer scratch use.

Falling back to plain HTTP ​

Omit the server.tls block entirely. somora reverts to HTTP/1.1 plain. You'll keep the 6-connection limit and lose secure-context features. Fine for single-window dev, no good for multi-agent daily use.

Mobile PWA — /mobile ​

somora ships a second web client at /mobile aimed at phones. It's a PWA — installable to your home screen on iOS / Android, runs in standalone-app mode with no browser chrome. Minimal-scope by design: chat only, no tmux / no file viewer / no multi-window layout. See mobile.md for the full feature scope, install flow per platform, and configuration knobs.

Installing:

  1. On the phone, with Tailscale connected, visit https://<your-host>.<your-tailnet>.ts.net:18737/mobile/ in Safari (iOS) or Chrome (Android).
  2. iOS: share menu → "Add to Home Screen".
  3. Android: Chrome's install banner, or menu → "Install app".

Build pipeline: build:mobile script alongside build:web, both run by build:all and triggered by the prepack hook on npm pack. The somora update flow picks this up automatically — no manual step. The release tarball always contains both web/dist and web-mobile/dist.

Isolated Claude config dir ​

somora-spawned claude-cli subprocesses run with their own config tree under ~/.somora/claude-home/, not the user's ~/.claude/. On first server start the dir is auto-created and the user's ~/.claude/.credentials.json is copied into it; a continuous sync (see below) keeps both credential stores on the same OAuth session so one claude login covers somora and the user's own Claude Code.

Everything else (project history, sessions, plugin marketplace state, shell snapshots, MCP-needs-auth cache, …) lives separately. somora's agents never see the user's interactive-CLI state, and vice versa.

Why it matters

  • Auto-update insulation. Anthropic's launcher silently rolls forward the user's claude binary. If a release migrates the state schema, somora's spawn — which may run a different binary version — would no longer read the migrated tree cleanly. The isolated dir keeps somora's state out of that path.
  • Privacy + predictability. The user's project conversations, installed plugins, and per-project skill caches never leak into agent context.
  • Reproducible deploys. A fresh somora install on a new machine starts from the same blank slate regardless of how the user's personal Claude Code is set up.

Scope

The isolation applies to the internal engine adapter that somora uses to talk to Claude. Tools the agent invokes for the user — tmux create, exec, the process family — strip somora-internal vars from the spawned shell by default (CLAUDE_CONFIG_DIR, SOMORA_CLAUDE_BIN, the other engine-binary overrides, TSX_TSCONFIG_PATH, NODE_ENV, and Claude Code's MCP-child markers), so a claude or codex you start inside a tmux pane sees your normal ~/.claude login state, and a project's own tsx/next/dotenv see the project's config rather than somora's. The inherit_agent_env: true flag opts back into inheritance when you specifically want it (see docs/tmux.md).

Overriding

Set CLAUDE_CONFIG_DIR in ~/.somora/somora.env (or shell env) to point at any directory you prefer — useful for shared multi-host setups, or when you want somora to read a hand-curated config tree. The auto-create + credential sync still runs on whichever path you supply.

If the user hasn't run claude login

The credentials file at ~/.claude/.credentials.json won't exist yet, and the bootstrap logs a warning instead of failing. Run claude login once interactively (any session) — the running server's credential watcher picks it up within seconds, no restart needed.

Shared-login credential sync

Sharing one login between two config trees has a structural enemy: the claude CLI refreshes OAuth tokens with an atomic write (tmp file + rename). A symlink from the somora-side file to ~/.claude/.credentials.json would not survive that — rename replaces the symlink itself, so after the first token refresh the link would silently materialize into a real file. From then on both trees would rotate the same OAuth session independently, and whichever side refreshes later invalidates the other; the losing side eventually fails with OAuth session expired and could not be refreshed.

somora therefore maintains the sharing as a continuously reconciled content sync (default claudeCli.sharedUserCredentials: true in config.yaml), never as a symlink:

  • Filesystem watcher on both parent directories plus a 60 s fallback poll — a token refresh (or fresh claude login) on either side propagates to the other within seconds, while the server runs.
  • Boot reconcile covers drift that happened while the server was down.
  • Pre-turn reconcile in the claude-cli engine guarantees every turn starts on the newest OAuth chain.
  • Auth-failure reconcile — if a turn still fails auth, somora reconciles inline and tells you whether re-sending the message is enough or a fresh claude login is needed.

On divergence the side with the later OAuth expiry wins (its refresh happened last, so its refresh-token chain is the live one) and is copied over the other — atomic write, mode 0600, previous content kept once at .credentials.json.somora-prev. A corrupt/unparseable file always loses to a healthy one.

Inspect or fix the state manually any time:

somora auth status   # both stores: mtime, OAuth expiry, in sync / diverged
somora auth sync     # one-shot reconcile (what the watcher does continuously)

GET /health also reports the sync state under claudeAuth (existence, expiry, divergence — never token material).

Running somora on a separate Claude account

Set claudeCli.sharedUserCredentials: false in config.yaml — somora then never touches either credentials file, and you manage ~/.somora/claude-home/.credentials.json yourself (e.g. via CLAUDE_CONFIG_DIR=~/.somora/claude-home claude login).

Environment overrides ​

VarDefaultPurpose
SOMORA_HOME~/.somoradata root for config / agents / sessions
SOMORA_PORTconfig.yaml:server.portserver bind port (override)
SOMORA_HOSTconfig.yaml:server.host (127.0.0.1)server bind host override. Prefer server.host in config.yaml — it survives somora update; a SOMORA_HOST env in the systemd unit is dropped by the update rebake. Auto-set to tls.publicHost for MCP-child callers when TLS is on.
SOMORA_TLS0set to 1 by parent when serving HTTPS — MCP-child callers use it to switch to https://
SOMORA_LOG_LEVELinfoPino log level
SOMORA_CLAUDE_BIN~/.local/bin/claudeClaude Code binary path
CLAUDE_CONFIG_DIR~/.somora/claude-homeIsolated state dir for claude-cli subprocesses (auto-created on boot, see "Isolated Claude config dir")
SOMORA_CODEX_BINunset (bundled)Debugging override for the Codex binary; otherwise the bundled @openai/codex is used
SOMORA_COMPACTION_TRIGGER_RATIOfrom configoverride compaction trigger
SOMORA_COMPACTION_SAFETY_PAIRSfrom configoverride compaction cushion
SOMORA_COMPACTION_MODELfrom configoverride compaction worker
SOMORA_COMPACTION_WORKERSfrom configcomma-separated worker cascade

The live values are queryable: GET /env returns the resolved set with isDefault flags, and the same data is logged at server startup as somora.env.

Develop from a checkout (contributors) ​

If you want to hack on somora itself rather than just run it:

bash
git clone https://github.com/thenaxon/somora_agent.git somora
cd somora
npm install                    # local install instead of -g
npm run dev:server             # terminal A — starts via tsx watch
npm run dev:cli                # terminal B — TUI against the dev server

The dev server reads/writes the same ~/.somora/ as the production binary, so anything you configure (providers, agents, persona files) shows up in both. To isolate, point SOMORA_HOME at a scratch dir:

bash
SOMORA_HOME=/tmp/somora-dev npm run dev:server

Other useful scripts:

CommandWhat
npm run typecheckserver-side tsc --noEmit
npm testevery *.test.mts under src/, web/src and web-mobile/src, each against a throwaway SOMORA_HOME; the web clients run from their own folders so their JSX compiles with their own tsconfig
npm test src/browserone subtree, same isolation (npm test web/src/lib for the web client)
npm run verify:fasttypecheck plus the suite
cd web && npm run devVite dev server for the web client (proxies API to :18737)
cd web && npm run buildrebuild web/dist/ (the bundle the production server serves at /web/)

Run tests through npm test, not tsx --test directly. The logger opens its file the moment it is imported, so a test started without an explicit SOMORA_HOME would write into the running installation's log and make its error count meaningless. The launcher sets a temporary home before anything loads and removes it afterwards; SOMORA_TEST_HOME=/some/dir keeps it when you want to read the test's own log. A file started by hand under node:test falls back to a temporary log directory too, but only the launcher isolates sessions, memory and the Claude config as well.

There's no built-in auth on the HTTP API — somora binds to 127.0.0.1 by default and assumes you're its only user. To expose it across a network, use Tailscale (see HTTPS section above) or put it behind a reverse proxy with proper auth.