Skip to content

somora HTTP API reference ​

All clients — TUI, web app, and any third-party tool — talk to the same HTTP+SSE+WebSocket surface. This document is the reference for that surface, so you can build your own client: a Telegram bridge, a status dashboard, a Voice frontend, a Stream-Deck integration, whatever you want to plug into your agents.

The server is a Hono app served by @hono/node-server over HTTP/2 (see setup.md for TLS via Tailscale). The base URL in a typical install is https://<host>.<tailnet>.ts.net:18737. The TUI and web both live at this same origin (TUI is the binary that ships with somora; web is mounted under /web).

Authentication & deployment model ​

There is no API key, no OAuth, no per-route auth. somora's security model is LAN-trust: the server binds to the loopback or the Tailnet, and anyone who can reach the address is authorised.

This is a deliberate choice that follows from somora's positioning:

  • Local-first. somora lives on your machine (or a server in your Tailnet). It holds your memory, talks to your accounts. There is no multi-tenancy concept.
  • Tailscale is the ACL. Tailnet ACLs decide who can reach the port; somora itself trusts whoever's already at the door.
  • Anything an agent can do, you can do. The API surfaces the same capabilities the agent has — memory writes, model switches, dream triggers. If the model is allowed to do it, so is your client.

Do not expose the somora port to the public internet. No auth guard means anyone who finds the URL gets full agent control, including the ability to read all memory + sessions + vault. Keep it on Tailscale or on 127.0.0.1 and tunnel.

Conventions ​

  • Encoding. Request and response bodies are JSON (Content-Type: application/json) unless otherwise noted. Attachments are multipart/form-data.
  • Errors. A non-2xx response carries { "error": "<message>" } in the body. Messages are human-readable strings and may be in German (somora speaks both — error text follows the locale of the originating component).
  • IDs. Session IDs are <YYYYMMDD>-<HHMMSS>_<slug>. The literal string main is the always-on default session per agent and is always addressable. A route that takes <slug|id> resolves in this order: main, an exact id, a session stored under exactly that name, then the newest session with that slug. The third rule matters for sessions written before the id format was enforced: whatever the session list shows can be addressed under the name it was shown with. An unknown reference is refused rather than resolved to a neighbour.
  • Polling cadence. Endpoints that surface live state are cheap by design — poll every 2 s for the dream loop, 30 s for dream phases, 60 s for sessions. The server caches expensive lookups (e.g. session sizes) so polling doesn't burn cycles.
  • SSE. Streaming endpoints use Server-Sent Events with named events (chat, tool, memory_inject, status, heartbeat). See SSE event vocabulary.

Versioning & stability ​

  • The version returned by GET /version is calendar-based and valid semver: 2026.1001.2 is the second build of 1 October 2026. Several builds a day are common during active work.
  • Most endpoints listed here are stable — the TUI and web both use them, and breaking them would break the shipped clients.
  • Endpoints marked ⚠ experimental may change shape in any bump. They are surfaced because the TUI/web already need them; lock yourself to a specific somora version if your client depends on them.

Core ​

GET /version ​

Returns the running somora version and, once the daily update check has answered, what somora.ai says is current (setup.md → The daily update check).

bash
curl https://<host>:18737/version
# { "version": "2026.1001.2", "update": null }
# { "version": "2026.1001.2",
#   "update": { "latestVersion": "2026.1005.1", "available": true, "note": "update Node first" } }

update is null until the first check has run (or when it is switched off); available is true when latestVersion is newer than version; note is optional.

GET /healthz ​

Liveness probe. Returns the plain-text ok with status 200. Use it to wait for server start, smoke-test load balancers, etc.

GET /health ​

Diagnostic snapshot. Use this when something looks stuck — e.g. an agent's session has stopped responding but the server is still accepting HTTP. The response shows which (agent, session) is busy, since when, what its current turn id is, queue depth behind it, and how long ago the last engine event reached SSE subscribers. Read-only, cheap (in-memory only).

bash
curl https://<host>:18737/health

Returns:

json
{
  "ok": true,
  "serverPid": 12345,
  "serverBootedAt": 1778756948428,
  "serverUptimeMs": 5142,
  "lockfilePid": 12345,
  "lockfileStartedAt": "2026-05-14T11:09:08.429Z",
  "activeSessions": 1,
  "totalKnownSessions": 3,
  "claudeAuth": {
    "enabled": true,
    "userExists": true,
    "somoraExists": true,
    "identical": true,
    "userExpiresAt": 1785503549000,
    "somoraExpiresAt": 1785503549000,
    "lastSyncResult": "noop",
    "lastSyncAt": 1785496349000
  },
  "memoryEmbedder": {
    "state": "ok",
    "provider": "local",
    "model": "all-MiniLM-L6-v2",
    "dim": 384,
    "error": null,
    "since": 1785496350120,
    "attempts": 1,
    "loadMs": 2373
  },
  "sharedIndex": {
    "state": "ready",
    "role": "owner",
    "path": "/home/me/.somora/index/shared.db",
    "files": 626,
    "chunks": 1466,
    "built_by": "seed:<agent>"
  },
  "sessions": [
    {
      "agent": "<your-agent>",
      "session": "main",
      "busy": true,
      "activePriority": "user",
      "activeSince": 1778757036447,
      "activeAgeMs": 1024,
      "activeCallId": null,
      "activeTurnId": "b5b7a734-...",
      "queueLength": 1,
      "userWaiting": 0,
      "agentWaiting": 1,
      "queued": [
        { "id": "3f0c2b1e-…", "kind": "agent",
          "preview": "Can you check whether the build log mentions …",
          "enqueuedAt": 1778757040112, "position": 1 }
      ],
      "lastEngineEventAt": 1778757036900,
      "lastEngineEventAgoMs": 571,
      "subscriberCount": 2,
      "lastPublishOkAt": 1778757036900,
      "lastPublishOkAgoMs": 571
    }
  ]
}

activeAgeMs is how long the current turn has been holding the per-session lock. activeTurnId is set for every running turn, whatever started it — a typed message, an agent_ask, a sub-agent brief, a sentinel fire, a tmux or browser wake, a voice consult or a wake-up — so a session that looks stuck can always be matched to a turn and stopped with POST /chat/abort. queued names the waiters behind it in lock order: the entry's id (the turnId of a typed message, the call_id of an agent_ask, the task_id of a sub-agent brief or a sentinel fire), its origin kind, the first 160 characters of its text and its position (1 = next). Any of these ids can go to DELETE /chat/queue/:id; the per-session view GET /agents/:agent/sessions/:session/work adds what is arriving and what the session started elsewhere. lastEngineEventAgoMs ticks up while the engine is silent — if it climbs past a few minutes on a chat turn (vs. a long local-LLM job), the turn is wedged and the engine watchdog will abort it. See setup.md engineWatchdog to tune thresholds per engine.

claudeAuth reports the shared-login credential sync between ~/.claude and somora's isolated claude-home (paths and mtimes elided above; never token material). identical: false with both sides present means the stores have diverged and the watcher hasn't caught up yet — if it persists, claude-cli auth is about to break; run somora auth status on the host. See setup.md.

sharedIndex is the vault/wiki retrieval index shared by all agents (memory.md). state is ready when agents read vault/wiki from it, building while the first build after an update is still running (agents then still answer from their own DB), disabled when no vault is configured, failed with error when the DB could not be opened or built. built_by says where the content came from: seed:<agent> (copied out of that agent's DB on the first boot after the update) or sweep (embedded from disk). null until the server has opened it.

memoryEmbedder is the health of the embedding model behind memory retrieval (see memory.md). The server loads it once at boot; state is ok when the model is loaded, loading while the (first-run) download is in flight, and failed when the last attempt threw — error then carries the reason. failed means every agent's memory search is BM25-only until the next retry succeeds (one retry per agent per minute, on search): the server keeps working, but semantic recall and dream dedup are silently degraded, so treat a persistent failed as an incident. The model cache lives under ~/.somora/models/transformers/ and survives updates.

subscriberCount is the number of currently-connected SSE clients (web / mobile / TUI tail) watching this session. lastPublishOkAt only advances when a broadcast reached at least one subscriber within the sse.publishTimeoutMs budget; if subscriberCount > 0 but lastPublishOkAgoMs grows without bound, at least one client is wedged — the next publish auto-evicts it (see sse.publishTimeoutMs in setup.md).

GET /host-stats ​

Host machine resource snapshot — CPU load and memory usage of the box somora is running on. Surfaced to the web taskbar's cpu / mem widgets and handy when somora lives on a VM you don't otherwise have a metrics view for.

bash
curl https://<host>:18737/host-stats

Returns:

json
{
  "cpu": {
    "loadAvg1": 0.42,
    "cores": 6,
    "percent": 7.0
  },
  "mem": {
    "totalBytes": 25186074624,
    "availableBytes": 23197777920,
    "usedBytes": 1988296704,
    "percent": 7.9
  }
}

cpu.percent is the 1-minute load average divided by core count and expressed as a percentage. Values over 100 mean the box has more runnable processes than CPUs — not capped, an overload signal is more useful than a clamped number.

mem.availableBytes is "memory the kernel can hand back without I/O". The reading is platform-specific so it matches what the OS-native tools report:

  • Linux: /proc/meminfo:MemAvailable (includes reclaimable page cache). Falls back to os.freemem() on kernels that don't expose it.
  • macOS: vm_stat pages free + inactive + speculative, times page size. Mirrors Activity Monitor's "Available" notion. Falls back to os.freemem() if the vm_stat binary is unavailable.
  • Other: os.freemem() straight (degraded but non-erroring).

usedBytes = totalBytes − availableBytes. Read-only, cheap (no disk I/O on Linux, a sub-millisecond vm_stat spawn on macOS).

GET /tui-config · GET /mobile-config ​

Display preferences for thin clients, read from config.yaml by the server so no client parses the file itself. /tui-config returns {show: {memory, tools}, verbose: {tools, memory, system, thinking}} (the tui: block, see display.md); /mobile-config returns {show: {tools, memory}} (the mobile: block). A custom client is free to use either as its own defaults.

GET /env ​

The environment overrides the running server resolved at boot, one entry per variable: {SOMORA_HOME, SOMORA_PORT, SOMORA_LOG_LEVEL, SOMORA_CLAUDE_BIN, SOMORA_CODEX_BIN, SOMORA_CODEX_TOOL_TIMEOUT_SEC, SOMORA_COMPACTION_TRIGGER_RATIO, SOMORA_COMPACTION_SAFETY_PAIRS, SOMORA_COMPACTION_MODEL, SOMORA_COMPACTION_WORKERS}, each {value, isDefault, note?} — value is what is in force, isDefault says the variable was unset or invalid, note explains a fallback (e.g. "unset → uses config.yaml server.port"). Diagnostic; see setup.md.

GET /tools ​

List every tool registered on the server (the same tools agents see). Useful for clients that want to surface a tool catalog.

bash
curl https://<host>:18737/tools

Returns { count, tools: [{ name, toolset, description, inputSchema, maxResultSizeChars, hasAvailabilityCheck }] } — inputSchema is the tool's JSON Schema, maxResultSizeChars is null when the tool uses the default cap, hasAvailabilityCheck says the tool has a runtime probe and may be hidden from some agents.

GET /agents/:agent/skills · PUT /agents/:agent/skills ​

Per-agent skill visibility — the skills half of the web client's Abilities matrix (the tools half is GET/PUT /agents/:agent/tools, documented in mcp.md).

GET returns { agent, gating, hasPatternRules, skills: [{ name, description, available, unavailableReason?, visible }] } — every skill installed on the instance, with visible telling whether this agent sees it. gating is the agent's skills: section ({deny, allow}) or null; hasPatternRules is true when it carries a hand-written allow-list, which the UI shows read-only.

PUT takes { deny: string[], allow: string[] } and rewrites only the skills: block of the agent's agent.yaml (comments and the rest of the file untouched; empty deny+allow removes the block). Names must be skill names ([a-z0-9-]). Takes effect on the agent's next turn. See skills.md.

GET /agents/:agent/tools · PUT /agents/:agent/tools ​

The tools half of the Abilities matrix (see mcp.md).

GET returns { agent, gating, hasPatternRules, tools: [{ name, toolset, mcpServer?, description, visible, availableNow }] } — every tool configured on the instance, built-in and imported from external MCP servers (mcpServer names the origin), with visible telling whether this agent's gating lets it through. gating is the agent's tools: section ({deny, allow}) or null; hasPatternRules is true when it carries an allow-list or a toolset:/glob deny rule, which the UI shows read-only.

PUT takes { deny: string[], allow: string[] } and rewrites only the tools: block of the agent's agent.yaml; 400 when the body has the wrong shape or the write fails. Returns {ok: true}; takes effect on the agent's next turn.

POST /agents/:agent/tools/:name ​

Invoke a tool directly as an agent (without going through a chat turn). The body is the tool's input shape; the response is the tool's output. Same authorisation as everything else (LAN-trust).

bash
curl -X POST https://<host>:18737/agents/<your-agent>/tools/memory_search \
     -H 'Content-Type: application/json' \
     -d '{"query":"voice satellites","limit":3}'

Images ​

Present only when imageGen.enabled is set and at least one model is configured; every route below answers 503 otherwise. See imagegen.md for configuration.

GET /images/status ​

{enabled, outputDir, maxImagesPerTurn, models: [{name, label, model, provider, defaults}]}, or {enabled: false}. Clients use this to decide whether to show an image-generation surface at all.

GET /images ​

Gallery listing, newest first. Query: query (prompt substring, case- insensitive), model, agent, since/until (YYYY-MM-DD), limit (default 60, max 200), offset.

Returns {total, offset, items: [ImageRecord], totalBytes}. total is the unpaged count.

GET /images/:id ​

One ImageRecord: prompt, model, specs, path, mime, bytes, cost, agent, session. A record may carry a linkedTo field; nothing writes it, and the clients show the one canonical path.

GET /images/:id/file ​

The image bytes, with the record's MIME. 410 when the record exists but the file was moved or deleted outside somora.

Files are addressed by record id, never by path — the client cannot name a file, so a user-chosen images directory does not turn this into a way to read arbitrary files.

GET /images/models/:name/capabilities ​

{model, source: 'catalog'|'config'|'unknown', known, values, maxN, maxReferences, sizeAlsoAccepts, defaults}. sizeAlsoAccepts lists named ratios the endpoint takes in size (null when it publishes none) — see the aspect-ratio note in imagegen.md.

values maps a spec field to its allowed values. A field absent from values has no known constraint — clients should offer free input there rather than an empty dropdown.

GET /images/catalog ​

{provider, models: [{id, name?}]} — what the provider currently offers. ?provider=<name> selects the provider; defaults to the one behind the first configured model. Read-only discovery aid; config still decides what somora will call.

POST /images/generate ​

Body: prompt (required), optional model (a configured handle) and any specs — resolution, aspect_ratio, size, quality, output_format, background, output_compression, seed, n — plus optional reference_images. A save_to field is accepted for compatibility and ignored: every image lives in one place, the configured images directory, and the response carries its path.

reference_images is base64 on this route — a browser has the bytes of a file the user picked and no server-side path for it. The image_generate tool takes file paths instead, for the mirror-image reason: an agent works on the same machine, and base64 in a tool argument would mean loading a file into its context just to send it straight back out.

Returns {images: [ImageRecord], costUsd, warnings?, fellBackFrom?}. warnings carries anything the endpoint did differently than asked — a size or aspect ratio it substituted (detected by measuring the returned image, not by trusting the endpoint to report it), an aspect_ratio that had to be sent as the closest size on the OpenAI wire (see imagegen.md), plus any ignored_params or warnings the provider itself sent. fellBackFrom is present only when a fallback: chain had to be walked, and names the models that were unavailable.

503 when the model is configured but not loaded right now (an image backend commonly shares a GPU box that runs one profile at a time and says so) — distinct from 502, which means the endpoint misbehaved.

400 for anything the caller can fix, with a message naming the field and the values that would have worked. 502 when the upstream itself failed.

DELETE /images/:id ​

Forgets the record. The file on disk is kept — returns {ok, path, fileKept: true}.

Media ​

One gallery over everything somora generated — images and videos. The /images/* routes above are the image-only view; these are the medium-agnostic ones the web Media window and the mobile PWA use.

GET /media ​

Query: kind (image | video; omit for both), agent, query (prompt substring), limit (default 60, max 200), offset.

Returns {total, offset, items: [MediaRecord], totalBytes}. A MediaRecord is {id, kind, createdAt, prompt, modelName, modelId, provider, specs, path, filename, mime, bytes, width?, height?, durationSec?, thumbPath?, thumbMime?, linkedTo, costUsd?, agent?, session?, references?, batchId, batchIndex} — kind is absent on records written before video existed and then means image; durationSec and the thumbnail fields are video-only.

GET /media/:id ​

One MediaRecord, 404 when unknown.

GET /media/:id/file ​

The bytes with the record's MIME, Content-Disposition: inline (?download=1 forces attachment), immutable cache headers, and HTTP Range support (206 / 416) so a video player can seek without re-downloading. 410 when the record exists but the file left the disk.

GET /media/:id/thumb ​

A video's still image (image/webp unless the record says otherwise), same headers as /file. 404 when the provider served no thumbnail.

DELETE /media/:id ​

Removes the record. Returns {ok: true} or 404.

Video ​

Video generation runs as jobs: POST starts one and returns at once, the render continues in the main server, the finished file becomes a MediaRecord, and an agent that started the job is woken with a from_system: 'job' turn. Enabled by videoGen.enabled in config.yaml; see videogen.md.

GET /video/status ​

{enabled: false, reason} when video is off or no model is configured. Otherwise {enabled: true, active, limit, models: [{name, label, model, provider, wire}], jobs: [VideoJob]} — active and limit are the concurrent-job slot (videoGen.maxConcurrent, default 4). ?agent=<name> limits jobs to that agent's. A VideoJob is {id, providerJobId, modelName, provider, prompt, specs, status, progress?, queuePosition?, error?, createdAt, updatedAt, mediaId?, path?, agent?, session?, references?} with status one of queued, in_progress, completed, failed; mediaId appears once the file is stored and is what /media/:id takes.

POST /video/generate ​

Body: prompt (required), optional model (a configured handle), the specs seconds, size, aspect_ratio, audio, quality, seed, optional reference_images (base64, as on /images/generate), and optional agent + session naming who should be woken when the job finishes.

Returns {job: VideoJob} immediately. 400 for a bad request, 429 when all job slots are busy, 503 when the model is configured but not available right now, 502 when the provider failed. Poll GET /video/status or watch the session for the completion turn.

Files ​

GET /files/view ​

Read a server-local file by absolute path. Used by the web client's FileView window so users can click absolute-path links emitted in agent messages (e.g. [report.md](/home/user/somoraworkspace/...)) and see the content in-app without SSH-ing into the server.

Policy is reused 1:1 from the file_read tool — the same allowlist (workspace + somora-home roots) and the same blocklist (~/.ssh, credential stores, system dirs, …). Symlink-resolution prevents path escapes via realpath check on the closest existing ancestor.

Read-only, no writes. This route answers what a file is; the bytes of anything it cannot inline come from GET /files/raw.

Query parameters

NameRequiredDescription
pathyesAbsolute filesystem path. ~-prefix is expanded server-side. Relative paths are rejected (no agent-context cwd here).

File kinds. A known text extension decides how the text is highlighted; everything else is classified from the file's magic bytes, because an extension is a claim and not evidence — a .dat holding PNG bytes is reported as an image. Nothing is refused for being the wrong type: an unrecognised file still comes back described, with a download link, which beats an error for a file the user can see referenced in chat.

kindResponse carries
.md, .markdownmarkdowncontent, full Markdown render
.txt, .log, other texttextcontent, monospace
.json, .jsonl, .yaml, .yml, .toml, .svgcodecontent, syntax highlighting
PNG / JPEG / GIF / WebP bytesimageurl, mime
MP4 / MOV / WebM bytesvideourl, mime
WAV / MP3 / OGG / FLAC / M4A bytesaudiourl, mime
PDF bytespdfurl, mime
anything elsebinarymime and a download link only

.svg is listed as code, not as an image: it is markup that can carry script, so it is shown as its own source rather than rendered.

Text responses are capped at 200 000 characters and set truncated: true past that. Every response carries downloadUrl.

Success response (200), text

json
{
  "path": "/home/user/somoraworkspace/somora_feedback/example.md",
  "kind": "markdown",
  "ext": ".md",
  "bytes": 4321,
  "content": "# Report\n…",
  "truncated": false,
  "downloadUrl": "/files/raw?download=1&path=…"
}

Success response (200), media

json
{
  "path": "/home/user/somoraworkspace/shots/run-12.png",
  "kind": "image",
  "ext": ".png",
  "bytes": 184320,
  "mime": "image/png",
  "url": "/files/raw?path=…",
  "downloadUrl": "/files/raw?download=1&path=…"
}

Error responses

StatusWhen
400Missing path query, relative path, path is a directory or non-regular file
403Policy blocked (path resolves under a blacklisted root or a denied somora-internal location)
404File does not exist
bash
curl -G "https://<host>:18737/files/view" \
     --data-urlencode "path=/home/user/somoraworkspace/somora_feedback/example.md"

GET /files/raw ​

The bytes behind a /files/view result: media the viewer displays inline, and a download for every other type. Same path policy, applied in the same two passes — the byte route is not a way around what the metadata route refuses.

Query parameters

NameRequiredDescription
pathyesAbsolute filesystem path, as for /files/view.
downloadno1 forces Content-Disposition: attachment, whatever the type.

Range requests are supported (Accept-Ranges: bytes, 206 with Content-Range, 416 past the end, and bytes=-N for a suffix). A browser seeking inside a video depends on it; without Range, scrubbing re-fetches the whole file. There is therefore no size cap.

Content-Type is taken from the magic bytes, never the extension, and X-Content-Type-Options: nosniff is set so a browser cannot re-guess it. Content-Disposition: inline is limited to image, video, audio and PDF — anything else is served as an attachment, because a file displayed inline runs on somora's own origin.

The path policy is the only boundary: files outside the workspace (/tmp/... screenshots, for instance) are viewable as long as they are not under a blocked root, which matches what file_read allows an agent to see.

bash
curl -G "https://<host>:18737/files/raw" \
     --data-urlencode "path=/home/user/somoraworkspace/shots/run-12.png" \
     -o run-12.png

Agents ​

GET /agents ​

List configured agents.

bash
curl https://<host>:18737/agents

Returns an array of AgentInfo:

json
[
  {
    "name": "<your-agent>",
    "description": "scribe and personal-assistant",
    "icon": "📝",
    "color": "#6366f1",
    "role": "Scribe"
  }
]

Source: ~/.somora/agents/<name>/AGENTS.md frontmatter.

GET /agents/:agent/system-prompt ​

Returns the persona part of the system prompt (SOUL.md, AGENTS.md, USER.md with somora's headings) — {agent, systemPrompt}. For the complete prompt as a turn sends it, use /prompt-preview below.

GET /agents/:agent/prompt-preview ​

The system prompt exactly as the next turn on ?session=<slug|id> (default main) would send it, without running a turn or touching the session:

json
{ "agent": "<your-agent>", "session": "main", "text": "…", "chars": 16210,
  "parts": [{"key": "self", "label": "Self-pointer", "chars": 900},
            {"key": "persona", "label": "Persona (SOUL.md · AGENTS.md · USER.md)", "chars": 9340},
            {"key": "team", "label": "Team block", "chars": 2652}, …],
  "tools": {"count": 41, "schemaChars": 26510, "names": ["exec", …]},
  "budgets": {"teamBlockChars": 3000, "personaFileChars": 8000, "personaTotalChars": 14000},
  "notIncluded": ["tool schemas (…)", "memory recall injected per turn", …] }

parts concatenate to text (separators included), in prompt order: self-pointer, persona, team, tool reminder, wiki overview, skills, session (which session this is, so the agent can name it to tools), project. tools counts the tools this agent can see after gating and the size of their JSON schemas — they travel on the API tool channel, not in text, and engines load them direct or deferred. The list is resolved against the model that session would actually use, including a per-session /model override, because a capability-gated tool differs per model: analyze_file is offered only where the model cannot see images itself. Read-only: a session whose wiki overview was never snapshotted is rendered without persisting the snapshot.

GET /agents/:agent/persona ​

{agent, files: [{name, exists, content, hash, chars, bytes, mtime, readOnly}], budgets, totals: {personaChars}} — AGENTS.md, SOUL.md, USER.md (editable) and agent.yaml (readOnly: true). hash is the optimistic lock for the write below; budgets is config.promptBudgets.

PUT /agents/:agent/persona/:file ​

Body: {content, baseHash}; file is one of AGENTS.md, SOUL.md, USER.md. Writes only when baseHash equals the hash of the file currently on disk — agents self-edit these files, so a stale save must not overwrite theirs: 409 {error, currentHash, currentContent} tells the client to reload. AGENTS.md must keep a parseable frontmatter whose name (when set) matches the agent directory and a non-empty body (400 otherwise). The previous version is kept as <file>.bak-<timestamp> (last five); the write is atomic. Returns {ok: true, hash, backup, chars}. The next turn uses the new text.


Team ​

The org chart from ~/.somora/team.yaml (see team.md). Readable and writable over HTTP; the file can also be edited by hand.

GET /team ​

{enabled, path, exists, valid, issues, file?, principal?, rules?, agents?, order?, unlisted?, missing?, warnings?} — file is the parsed document as written; agents (keyed by name: {name, title, reportsTo, involveFor, notFor, notes?, children, depth}), order (pre-order walk), unlisted (agents on disk missing from the file) and missing (file entries without a directory) are the resolved view. enabled: false with exists: false means no file; with valid: false the issues say what is wrong ({path, message}), and the server keeps the last valid team in force.

GET /team/preview/:agent ​

{agent, enabled, block, chars, softMaxChars} — the exact # Your team text that agent gets in its system prompt. 404 for an unknown agent.

GET /team/check ​

{exists, valid, issues, warnings, unlisted, missing, blocks: [{agent, chars, overSoftMax}], softMaxChars} — what somora team check prints, minus the persona scan.

PUT /team ​

Body: the whole document as JSON ({version: 1, principal, rules?, agents} — the same shape GET /team returns under file). The server validates exactly like the loader; 400 {error, issues: [{path, message}]} writes nothing. On success the previous file is kept as team.yaml.bak-<timestamp> (last five), the new one is written atomically, the read cache is dropped, and the response is {ok: true, backup, …} plus everything GET /team returns. Agents see the change on their next turn.

POST /team/init ​

Body: {principal?: string}. Writes a first document with every agent on disk reporting to the principal (titles from the frontmatter, the default rules spelled out). 409 when a file already exists — this never overwrites. Returns the same shape as GET /team.

POST /team/preview ​

Body: {file, agent} — a draft document and the agent to render for. Returns {agent, valid, issues, warnings, block, chars, softMaxChars}; an invalid draft comes back with valid: false and its issues, and nothing is written. This is what the web Team window's live preview uses.

Sessions ​

A session is a single conversation thread inside an agent. Each agent has a magical main session plus any number of named sessions.

GET /agents/:agent/sessions ​

List sessions for one agent. Archived sessions are filtered out by default — pass ?include_archived=true to surface them.

bash
curl https://<host>:18737/agents/<your-agent>/sessions
curl https://<host>:18737/agents/<your-agent>/sessions?include_archived=true

Returns an array of SessionSummary:

json
[
  {
    "id": "20260511-093251_research-notes",
    "slug": "research-notes",
    "isMain": false,
    "createdAt": "2026-05-11T09:32:51.000Z",
    "lastActivity": "2026-05-12T14:08:22.000Z",
    "messageCount": 47,
    "isArchived": false,
    "byteSize": 18923,
    "engine": "claude-cli",
    "dreamCoverageTs": 1715520120000,
    "dreamLagEvents": 4,
    "unreadAt": "2026-05-12T14:08:22.000Z",
    "seenAt": "2026-05-12T13:55:00.000Z"
  }
]

unreadAt and seenAt drive the unread badge UX: a session is unread when unreadAt > seenAt (or seenAt is null). See /activity/stream for the live feed that keeps these in sync across clients.

GET /sessions ​

Cross-agent session list. Same data as the per-agent endpoint, but flattened across every agent and wrapped as {sessions: [...]}, with each row carrying its agent name. Used by the web Sessions tool and by the session_list agent tool. Each row also says whether a turn is running in that session right now: busy, queueLength (turns waiting behind it) and activeSince (ms timestamp, null when idle).

bash
curl https://<host>:18737/sessions
curl https://<host>:18737/sessions?include_archived=true

POST /agents/:agent/sessions ​

Create a new named session.

bash
curl -X POST https://<host>:18737/agents/<your-agent>/sessions \
     -H 'Content-Type: application/json' \
     -d '{"slug":"research-notes"}'

Returns 201 { id, slug, agent }. The slug must match [A-Za-z0-9_-]+. One live session per slug: when a non-archived session already carries the name, the answer is 409 { error, id, slug, agent, exists: true } with that session's id — the web and TUI simply switch to it. Archiving or resetting the session frees the name. Cannot be the reserved string main (400).

POST /agents/:agent/sessions/:session/archive ​

Hide a session from active views without deleting it. The <id>.jsonl and <id>.meta.json stay on disk; only the archived: true flag in meta is set. Reversible via unarchive.

bash
curl -X POST https://<host>:18737/agents/<your-agent>/sessions/<id>/archive \
     -H 'Content-Type: application/json' \
     -d '{"reason":"old smoke test"}'

The main session cannot be archived directly — use /reset instead.

POST /agents/:agent/sessions/:session/unarchive ​

Clears the archived flag.

GET /agents/:agent/sessions/:session/export ​

Download the session as either raw JSONL (canonical, byte-identical to the source-of-truth file on disk) or a rendered Markdown transcript (human-readable, suitable for Obsidian / GitHub / blog posts).

Query param format:

  • json — Content-Type: application/x-ndjson. The complete JSONL with every event preserved (turn_start, tool_call, engine_meta, …). Use this for backups and cross-host transfer.
  • markdown (default) — Content-Type: text/markdown. Renders user/assistant turns as ## sections, tool calls as collapsible <details> blocks with their JSON args/results, and engine_meta items (e.g. codex's plan/todo lists) as task-style bullet lists with status glyphs. Skips bookkeeping events (turn_start, turn_end, assistant_audio) — those don't add value in a transcript.

Both responses set Content-Disposition: attachment so browsers trigger a file save.

bash
# Markdown transcript
curl https://<host>:18737/agents/<your-agent>/sessions/main/export?format=markdown \
     -o <your-agent>-main.md

# Raw JSONL (full fidelity)
curl https://<host>:18737/agents/<your-agent>/sessions/main/export?format=json \
     -o <your-agent>-main.jsonl

The web client surfaces this via per-row download icons in the Sessions tool (file-text icon = markdown, file-json icon = JSONL). The TUI has /export [json|markdown] [path] — see display.md for the slash-command reference.

POST /agents/:agent/sessions/:session/reset ​

Archive the current session content and start fresh. Triggers an asynchronous REM extraction over the archived content if REM is enabled for the agent.

bash
curl -X POST https://<host>:18737/agents/<your-agent>/sessions/main/reset
# { "agent": "<your-agent>", "session": "main",
#   "archivedId": "20260512-140822_main-archive",
#   "dreamSpawned": true }

The reset returns immediately. The REM run continues in the background; check its progress via GET /dream-states or via the per-agent REM badge in the web AgentDock.

Builder sessions ​

Agents of kind builder (see builder.md) carry mode, phase, plan file and task list per session, plus at most one open question. The web task panel reads and writes these; the builder's own tools (todo_write, ask_user, plan_write) call the same routes.

  • GET /agents/:agent/sessions/:session/builder → {agent, session, kind, state, question, turn} — question also carries the write-scope question a builder's file_write/file_patch raises for a path outside the project folder (attended mode; see builder.md) — turn is {turnId, startedAt, toolCalls, lastTool?, lastToolAt?} while a turn runs on the session, else null; state is {mode: "attended"|"unattended", phase: "plan"|"build", planPath, todos: [{content, status, priority?}]} or null before the session's first turn; question is {questionId, question, header?, options: [{label, description?}], multiple, askedAt, expiresAt} or null.
  • PATCH /agents/:agent/sessions/:session/builder {mode?, phase?, planPath?, orderer?} → {agent, session, state}. Publishes builder_state. orderer {agent, session?} names the agent that handed the order over (builder_dispatch sets it).
  • POST /lsp/diagnostics {agent, session, path, touch?} → {diagnostics: {server, errors, errors_in_other_files} | null} — the language server's verdict on a file a builder just wrote (what file_write/file_patch call; see lsp.md); touch: true only starts the server. GET /lsp/status → the servers found, the config and the running instances.
  • GET /builders → {builders: [{name, role, description, busy: [{session, workdir, since}]}]} — the agents of kind builder and the folders each is working in right now.
  • GET /builders/busy?workdir=<path> → {workdir, busy: {agent, session, turnId, workdir, since, reason} | null} — whether a builder's running turn claims that folder (or a parent/child of it). Without workdir: {claims: [...]}.
  • POST /agents/:agent/sessions/:session/builder/go {note?} → 202 {agent, session, state, turnId, callId?, wakes?} — sets the phase to build and starts a turn telling the builder the plan is approved (note is appended to that message). With an orderer in the state the turn runs as that agent's detached ask (callId = turnId, wakes = the orderer): it is woken with the report like after agent_ask with wait:false. 409 {error, busy} when another builder's turn is working in the session's folder right now.
  • PUT /agents/:agent/sessions/:session/todos {todos: [{content, status, priority?}], by_agent?} → {agent, session, todos} — replaces the whole list (max 100 items). Publishes todo_updated.
  • PUT /agents/:agent/sessions/:session/plan {content} → {path, bytes, archived?, note?} (what plan_write calls); archived names where a plan file that was already there and not written by this session was moved: PLAN-<date>-<session>.md beside it. The path is the session's plan path (the pinned project's PLAN.md, else <workspace>/PLAN.md), written atomically under the file write policy.
  • POST /agents/:agent/sessions/:session/ask {question, header?, options: [{label, description?}] (2-6), multiple?, timeout_ms?} — blocks until the person answers or the wait runs out (default 30 min, max 4 h) and returns {answered, answers: string[], text?}. Publishes question_asked when the question opens. A second question on the same session supersedes the first (which returns unanswered).
  • POST /agents/:agent/sessions/:session/answer {questionId, answers?: string[], text?} → {ok: true}, or 404 when no such question is open. Publishes question_answered.

GET /agents/:agent/sessions/:session/work ​

What this session is doing right now, what waits behind it, which answers are about to arrive, and what it started elsewhere. This is the data behind the queue badge in the web client, the sheet on mobile and /queue in the TUI. Previews only: the first 160 characters of each message, never a prompt, a tool body or a secret.

bash
curl https://<host>:18737/agents/<your-agent>/sessions/main/work
json
{
  "asOf": 1789300000000,
  "agent": "<your-agent>",
  "session": "20260913-101500_main",
  "busy": true,
  "active": {
    "id": "1f3a…", "kind": "agent", "state": "running",
    "preview": "Can you check whether the release notes mention …",
    "target": { "agent": "<your-agent>", "session": "20260913-101500_main" },
    "requester": { "agent": "<other-agent>", "session": "20260910-083000_main" },
    "enqueuedAt": 1789299990000, "startedAt": 1789299991000, "turnId": "1f3a…"
  },
  "queued": [
    { "id": "b5b7…", "kind": "human", "state": "queued",
      "preview": "and the changelog too",
      "target": { "agent": "<your-agent>", "session": "20260913-101500_main" },
      "requester": { "human": true },
      "enqueuedAt": 1789299995000, "position": 1 },
    { "id": "task-…", "kind": "sentinel", "state": "queued",
      "preview": "[sentinel] inbox digest …",
      "target": { "agent": "<your-agent>", "session": "20260913-101500_main" },
      "enqueuedAt": 1789299998000, "position": 2 }
  ],
  "pendingWakes": [
    { "id": "9d2c…", "kind": "agent", "about": "a2a", "state": "done",
      "preview": "Which of the three drafts …",
      "target": { "agent": "<other-agent>", "session": "20260910-083000_main" },
      "requester": { "agent": "<your-agent>", "session": "20260913-101500_main" },
      "enqueuedAt": 1789299900000, "startedAt": 1789299901000, "finishedAt": 1789299999500 }
  ],
  "children": [
    { "id": "task-…", "kind": "subagent", "state": "running",
      "preview": "Summarise the three …",
      "target": { "agent": "<your-agent>", "session": "sub-<your-agent>-20260913-101700" },
      "requester": { "agent": "<your-agent>", "session": "20260913-101500_main" },
      "enqueuedAt": 1789299970000, "startedAt": 1789299971000 }
  ]
}

Every entry has the same shape:

  • id — the work id: the turnId of a typed message, the call_id of an agent_ask, the task_id of a sub-agent brief or a sentinel fire, the consult id of a question from a call, the job id of a video render.
  • kind — where it came from: human, agent, subagent, sentinel, tmux, browser, voice or wake; a wake also says about (a2a, subagent or job).
  • state — queued, running, done, failed, cancelled or dequeued (taken out of the queue before it started).
  • preview — the first 160 characters of the text, line breaks folded to spaces. The only text exposed.
  • target — the {agent, session} the turn runs on.
  • requester — who asked for it: {agent, session} for an agent_ask, a sub-agent brief or a wake; {human: true} for a typed message; {voiceCall} for a question from a call. Absent for sentinel, tmux and browser turns, which nobody waits on.
  • enqueuedAt, startedAt?, finishedAt? — epoch ms. turnId? once the turn runs; error? for failed, cancelled and dequeued.

The four lists:

  • active — the turn holding the session lock, or null. busy is the lock itself; when a turn holds it without a ledger entry the response carries activeTurnId instead.
  • queued — the waiters in lock order, each with position (1 = next). Any of them can be removed with DELETE /chat/queue/:id.
  • pendingWakes — work this session asked for that has finished and whose wake-up turn is scheduled (the grace is agentLoop.wakeGraceMs, see setup.md); about says what kind of answer is on its way.
  • children — the sub-agents and agent_ask calls this session started that are still queued or running, with their target so a client can jump there.

Everything here lives in memory and is gone after a restart, like the sub-agent registry. 404 for an unknown agent or session.


Models & thinking ​

Per-session overrides for the active model and the active thinking level. Both fall back to the persona default when unset.

Model ​

bash
# Read current
curl https://<host>:18737/agents/<your-agent>/sessions/main/model

# Set per-session override
curl -X PUT https://<host>:18737/agents/<your-agent>/sessions/main/model \
     -H 'Content-Type: application/json' \
     -d '{"model":"claude-opus-4-7"}'

# Clear override (back to persona default)
curl -X DELETE https://<host>:18737/agents/<your-agent>/sessions/main/model

The model field accepts an alias (claude-opus-4-7), a <provider>/<id> tuple (anthropic/claude-opus-4-20250514), or anything else resolvable by GET /models.

A switch takes effect at the next turn: a turn that is already running keeps the model it started with. Every open client is told through the session_model SSE event.

Both routes take optional by_agent / by_session in the body — sent by the session_model agent tool. A switch made by an agent is written into the affected conversation as an engine_meta row (engine: "somora", itemType: "session_model", label model switched) naming who switched to what, so a person reading that session sees it. A switch without by_agent — a person using their own client — leaves no row.

Switching models mid-session is safe on every engine. On codex-cli the thread simply continues under the new model (thread/resume takes the model; Codex may compact the thread context once) and somora drops a model switch marker into the conversation. The somora session (id, history, meta) is untouched; alias changes that resolve to the same underlying model don't trigger a re-thread.

Thinking ​

bash
# Read
curl https://<host>:18737/agents/<your-agent>/sessions/main/thinking

# Set
curl -X PUT https://<host>:18737/agents/<your-agent>/sessions/main/thinking \
     -H 'Content-Type: application/json' \
     -d '{"level":"medium"}'

# Clear override
curl -X DELETE https://<host>:18737/agents/<your-agent>/sessions/main/thinking

Levels: off, low, medium, high — anything else is a 400. Per-model reasoning.levels in config.yaml decides which wire word (minimal, xhigh, max, …) each level becomes; see thinking.md.

GET returns {agent, session, effective, override, personaDefault, source, modelSupportsReasoning, wire}: effective is the level in force (session override > persona default > null = engine default), source says which of those won (session-override / persona-default / engine-default), modelSupportsReasoning tells a client whether the setting is live or dormant on the current model, and wire is the value actually sent when it differs from the level (e.g. high → xhigh). PUT answers {agent, session, level}, DELETE answers {agent, session, cleared: true}.

The GET …/model payload is {agent, session, provider, modelId, alias, engine, contextWindow, source, override, personaDefault} with source session-override or persona-default; PUT answers {agent, session, model, resolved: "<provider>/<modelId>"} and 400 names an unknown model; DELETE answers {agent, session, cleared: true}.

GET /models ​

List models the server knows about. Sources: config-level model definitions plus engine-discovered models.

bash
curl https://<host>:18737/models

Each entry: { provider, id, alias, engine, contextWindow, capabilities, ref }. ref is the canonical handle to pass to the model-set endpoints.


External MCP servers ​

Status and control of the MCP hub (mcp.servers in config.yaml, see mcp.md). All three answer 503 when no external server is configured.

Each entry carries unavailable: {since, until, reason} (epoch ms) while the model is marked unreachable — a cascade (chat fallback, REM worker chain, compaction) hit a host error on it within the last fallback.retryUnavailableMinutes; cascades start past such models.

POST /models/availability/reset ​

Forget every "unavailable" mark, so the next cascade tries the primaries again. 200 {ok: true, cleared, retryUnavailableMinutes}. A config reload does the same as a side effect.

GET /mcp/status ​

{enabled: true, servers: {<name>: {state, toolCount, transport?, lastError?, lastConnectedAt?, consecutiveFailures}}} — state is pending, connected, failed, needs-auth or disabled.

POST /mcp/servers/:name/reconnect ​

Tears the connection down and reconnects immediately, resetting the backoff. {ok: true, status} with the server's new status entry; 400 when the name is unknown or the server is disabled in config.

POST /mcp/call ​

Body: server, tool (the upstream tool name), args (object), optional timeoutMs. Calls the tool through the hub and returns {isError, text, images: [{data, mimeType}]}. 502 when the upstream call failed. This is the loopback path somora's own MCP children use; it bypasses per-agent tool gating, so treat it as an operator surface.

Config reload + restart ​

bash
curl -sk https://<host>:18737/config/status          # loadedAt, changedOnDisk, restartRequiredSections, restartAvailable
curl -sk -X POST https://<host>:18737/config/reload  # → { ok, changed: [...], restartRequired: [...] } or 400 with the schema issues
curl -sk -X POST https://<host>:18737/server/restart # → { ok, via: "systemd", expectedDowntimeSeconds } or 409 without a systemd unit

Reload validates the file first and keeps the running config on any error. Sections listed in restartRequiredSections (server, memory, obsidian, wiki, mcp, claudeCli, codexCli, stt, tts, sentinel, tmux, web, mobile) are consumed at boot and only change after a restart; the rest applies to the next request. See web.md for the taskbar surface and the TUI's /reload / /restart.

Sampling ​

Per-session sampling override (temperature, top_p, …), merged over the agent's and the model's defaults. Only the openai-compatible engine applies it — engineSupportsSampling tells clients whether the setting is live or dormant. Full description in sampling.md.

bash
curl https://<host>:18737/agents/<your-agent>/sessions/main/sampling

curl -X PUT https://<host>:18737/agents/<your-agent>/sessions/main/sampling \
  -H 'Content-Type: application/json' -d '{"temperature":0.7}'      # merges; null drops a key

curl -X DELETE https://<host>:18737/agents/<your-agent>/sessions/main/sampling

Keys: temperature, top_p, top_k, min_p, frequency_penalty, presence_penalty, repetition_penalty, seed, stop; an unknown key or an out-of-range value is a 400 naming the field. GET returns {agent, session, effective, override, personaDefault, modelDefault, source, engineSupportsSampling} (source is session-override, persona-default, model-default or engine-default). PUT merges the body into the override and returns {agent, session, override} (null once the last key is dropped); DELETE returns {agent, session, cleared: true}.

Chat ​

POST /chat/send ​

Fire-and-forget. The server returns 202 immediately; the actual turn runs in the background and emits SSE events to /chat/stream subscribers on the same (agent, session).

bash
curl -X POST https://<host>:18737/chat/send \
     -H 'Content-Type: application/json' \
     -d '{"agent":"<your-agent>","session":"main","text":"Was steht heute an?"}'

Body fields:

  • agent — agent name; when omitted the server falls back to the alphabetically first configured agent, so pass it
  • session (optional) — defaults to "main"
  • text (required) — user message
  • attachments (optional) — array of {hash, name, mime, size} — refs from prior POST /attachments calls
  • input_modality (optional) — "voice" when the client filled the text through its microphone (STT); recorded as user_message.input.modality and a precondition for a spoken reply
  • stt_provider (optional) — free-form tag of the STT path used
  • auto_play_requested (optional) — the client wants a spoken reply for this turn (assistant_audio); honoured only together with input_modality: "voice", see voice.md
  • subagent_depth (optional) — nesting depth when the turn is a sub-agent brief; the turn's origin becomes {kind: "subagent"}
  • from_agent (optional, A2A) — when set, the turn is attributed to another agent (used by agent_ask tool)
  • from_session (optional, A2A) — the session the asking agent wrote from; ignored without from_agent. Same meaning as on /chat/send-sync.
  • agent_ask_call_id (optional, A2A) — correlation UUID. A call id posted here is registered like one from /chat/send-sync, so GET /a2a/ask-result finds it while the turn is queued or running, not only in the target's history afterwards.
  • steer (optional, default false) — hand the text to the turn that is running on this session right now instead of queuing a turn of its own; see Steering below. Ignored together with agent_ask_call_id.

Response: { ok: true, turnId }. The turnId is the server-issued identifier for the queued/running turn — clients echo it through to match later SSE events (turn_queued, user_message) back to the optimistic bubble they rendered locally. With steer: true the response is { ok: true, steered: true, steerId, turnId } when the message went into the running turn (turnId is that turn's), or { ok: true, turnId, steered: false } when nothing steerable was running and it queued as usual.

Streaming responses arrive via /chat/stream; this endpoint just acknowledges receipt.

Steering ​

A message for a session whose turn is still running can be handed into that turn instead of waiting behind it: steer: true. The server keeps it in a per-session letterbox; the engine reads the letterbox at its next step boundary — after the tools of the current round returned, before the next model call — and gives the model the text as a user message framed as "sent while you were working, delivered before your next step". The model then changes course, stops, or carries on as told. A running tool call is never interrupted; the message lands after it returns.

What a client sees: the 202 response carries steerId; a steer_queued SSE event tells the other windows on the session; and when the engine has handed the text to the model, a user_message event (and a session record) with steer: true and the same steer_id follows — that is the moment the bubble stops being "pending". The record sits in the session file exactly where the model read it, between the tool results of one round and the next assistant step.

Engines: openai-compatible (somora's own loop), claude-cli (streamed input) and codex-cli (turn/steer) take steer messages; any other engine does not, and the request queues instead (steered: false). A message that arrives when the turn is already finishing, too late for the engine to read it, becomes an ordinary queued turn of its own — nothing is dropped. Sub-agent and voice turns are steerable like any other; agent_ask calls are not steered (their answer must come from a turn of their own).

The web composer shows a steer / queue toggle next to Send while a turn runs: "steer" = the next message steers, "queue" = it queues. The agent's steering: setting in agent.yaml is the default position (see agents.md).

Queuing ​

Sends on a (agent, session) that already has a turn running are enqueued, not rejected. The server holds a per-session lock with one FIFO queue, in arrival order, and every turn takes the same path into it: typed and dictated messages, agent_ask calls, sub-agent briefs, sentinel fires, tmux and browser wakes, voice consults and the wake-ups that bring a late answer, a finished sub-agent or a rendered video back. No origin jumps ahead of another. Turns are labelled by where they came from, for diagnostics only (activePriority in /health):

  • user — a person typing or dictating (no from_agent, no system origin)
  • agent — everything else

The currently-running turn always finishes — preempting would corrupt JSONL — so a queued turn starts only after the lock holder releases. Where a turn came from is recorded on its user_message as origin (see GET /chat/stream).

Every turn is an entry in one work ledger, whoever started it. GET /agents/:agent/sessions/:session/work shows the running entry, the waiters in order, the answers about to arrive and the sub-agents and agent_ask calls the session started; /health lists the waiters of every session next to its counters; and DELETE /chat/queue/:id takes any waiting entry out again.

Clients can opt into rendering a queue indicator by listening for the turn_queued SSE event (see below). UIs without it still work; the turn runs eventually, just without a visible "waiting" hint.

DELETE /chat/queue/:id ​

Take a waiting turn back before it starts, whoever queued it. :id is the work id the queue views show: the turnId that POST /chat/send returned for a typed message, the call_id of an agent_ask, the task_id of a sub-agent brief or a sentinel fire. Nothing is written to the session — the message never became a turn.

Without a body the request acts as the person and may remove anything. With a body {"requesting_agent": "<name>"} it acts as that agent and may remove only entries that agent asked for itself; anything else is 403 {ok: false, reason: "forbidden"}. This is what agent_ask_cancel sends.

  • 200 {ok: true, id, kind, agent, session, text?, attachments?} — the waiter is gone. For a typed message (kind: "human") the payload comes back with text and attachments: [{hash, name, mime, size}], so the client can put it into its composer (attachment refs are still valid, no re-upload needed). turnId carries the same id for older clients.
  • 409 {ok: false, reason: "already_started"} — the lock went to this turn meanwhile; it is running. POST /chat/abort is the tool now. With a body {"requesting_agent": "<name>", "withdraw_running": true} (what agent_ask_cancel sends) a running call of that agent's own is instead withdrawn: 200 {ok: true, state: "withdrawn", id, turnId, agent, session, steered, steerId?, ranMs} — its outcome wakes the requester no more, and when the target's turn reads its steer inbox (steered: true) a stop message is put there with a steer_queued SSE event like any steered message; the turn is not aborted. 403 when the call is not that agent's.
  • 404 {ok: false, reason: "unknown"} — nothing waits under that id (started, finished, or never queued here).

Whoever asked for a removed entry is told:

  • an agent_ask, on the line or already pending, reads state: "failed" with error: "removed from the queue by the user before it started" — as its inline result, through agent_ask_result and through GET /a2a/ask-result; an asker that had already stopped waiting is woken with an [agent answer] turn saying so, the same way it would have been woken with the answer;
  • a sub-agent brief reads cancelled with the same error in subagent_status, subagent_result and /spawn-status, and its parent is woken with a [subagent attention] turn saying there is no result;
  • a sentinel fire is recorded in the trigger's history as skipped with skipReason: "removed from the queue by the user";
  • a wake-up turn is dropped quietly; the result it was bringing stays readable.

Side effects on success: a turn_dequeued SSE event for every open client on the session, followed by fresh turn_queued events for the typed messages that moved up (their ahead shrank).

POST /chat/send-sync ​

Synchronous variant. Waits for the turn to finish and returns the full result inline. Slower (you block on it) but simpler for clients that don't want to manage SSE.

bash
curl -X POST https://<host>:18737/chat/send-sync \
     -H 'Content-Type: application/json' \
     -d '{"agent":"<your-agent>","session":"main","text":"…"}'

Body fields: same as /chat/send (agent, session, text, attachments, from_agent, agent_ask_call_id, subagent_depth), plus:

  • model (optional) — an alias or provider/id from config.yaml that answers this one turn instead of the session's model.

  • max_rounds (optional) — per-turn override of agentLoop.maxRounds.

  • attachments (optional) — refs from POST /attachments, exactly as on /chat/send. This is the route the A2A tools use, so it is also how one agent hands another a picture the receiving model has to actually see. A target whose model has no vision gets the vision worker's description instead, the same as for a chat attachment. Agents normally do not build this by hand: agent_ask and spawn_subagent take images: ["/absolute/path.png"] and upload for them (see agents.md).

  • from_session (optional, A2A) — the session the asking agent wrote from (id or main). Persisted as user_message.from_session and shown to the target in the attribution header ([Message from agent <other-agent>, session <slug>]) so it can address a follow-up. The server composes that header into the turn's frame (the stored row's ephemeral, see user_message under GET /chat/stream); the stored text is the message itself. Ignored without from_agent.

  • waiter_agent / waiter_session (optional, A2A) — identify the caller turn that blocks on this request. Used by agent_ask and spawn_subagent internally to register the wait in the server's deadlock guard; set both or neither.

  • create_session (optional, default false) — when session is a named slug that does not exist on the target yet, create it (with the standard timestamped id) and deliver the message into it. Only slugs: main always exists, exact ids and sub-* names answer 400. The target is told beside its first message — in the turn's frame, not in the stored text — that the session was just created ([Your session '<slug>' was just created by <agent> for this conversation; it runs on model <model>.]).

  • create_model (optional, with create_session) — alias or provider/id pinned on the session if this call creates it (same effect as PUT …/sessions/:session/model). An unknown model is 400 {error, known_models} and nothing is created. When the session already exists the model is ignored and the response carries session_note saying so.

  • detach (optional, with from_agent and agent_ask_call_id) — hand the message over and return at once instead of waiting for the reply: 202 {call_id, state: "pending", session_id, session_created?, session_model?, session_note?}. The turn queues and runs as usual; the asker is woken with an [agent answer] turn when the reply lands (the record line names the target, the session and the first words of the answer; the instruction to read the whole answer with agent_ask_result accompanies it as the turn's frame), or reads it with GET /a2a/ask-result. This is agent_ask with wait: false. Ignored without the two A2A fields.

The success response is the turn result plus session_id (the resolved id), session_created: true and session_model when this call created the session, or session_note when create_model was ignored.

An unknown session answers 404 with the target's existing, non-archived session slugs, so a caller that guessed wrong can correct itself instead of retreating to main:

json
{ "error": "session 'reserch' not found for agent '<other-agent>'",
  "known_sessions": ["main", "research", "somora-dev"] }

When waiter_* are present and the request would close a wait cycle (the target is already — directly or through a chain of waits — blocked on the caller), the server responds 409 instead of deadlocking:

json
{ "error": "circular A2A wait: …", "circular_wait": true,
  "chain": ["scribe/main", "coach/main", "scribe/main"] }

Response on success: the full turn result (finalText, usage, model, ms, …). A call that a person removes from the target's queue before it starts (DELETE /chat/queue/:id) answers 200 with outcome: "failed", dequeued: true and error: "removed from the queue by the user before it started".

Sub-agent tasks — /spawn-* ​

The HTTP twins of the spawn_subagent / subagent_* tools. Every spawn is one of these tasks, whether the tool was called with wait: false or wait: true — a synchronous spawn registers the task, runs it in the background and waits for it through the result route, which is why it is listed, stoppable and cancellable like any other task while the parent waits. Agents running in an MCP child (claude-cli, codex-cli) reach the task store this way; a custom client can use them to run a sealed background task in a fresh session and collect the result.

POST /spawn-async ​

Body: agent and session (required — a slug that does not exist is created with the standard timestamped id; an exact id that is gone is a 404), text (the task), optional from_agent, parent_agent + parent_session (who to report back to; default the caller), subagent_depth, model (override), max_rounds, attention (false suppresses the [subagent attention] wake of the parent), and attachments (refs from POST /attachments — pictures that belong to the brief).

Returns 202 {task_id} at once; the turn runs in the background under the target session's lock. 429 when the per-agent concurrent spawn cap is full.

GET /spawn-status?task_id=… ​

{task_id, state, parent_agent, parent_session, target_agent, target_session, started_at, finished_at?, error?} — state is running, done, failed or cancelled. 404 for an unknown id.

GET /spawn-result?task_id=… ​

Same fields plus result (the full turn result: finalText, outcome, tool_calls, files_written, media, usage, …) once terminal. result.follow_ups: string[] (oldest first) holds texts that reached the parent after the sub's report: the outcome of work the sub started and did not wait for — its own subs, agents it asked. Each one is announced to the parent with a [subagent attention] wake whose text says Task '<task_id>' … has a follow-up: the work it started has finished. It begins: "…" and whose frame reads Fetch it with subagent_result({ task_id: "<task_id>" }) — the follow-up is in its follow_ups field — then continue whatever depended on it. If nothing does, a short acknowledgement to the user is enough. (see agents.md). A task still running answers 409 {task_id, state: "running", error}. wait_until_done=1 + timeout_ms block server-side (capped at agentLoop.longTaskMaxTimeoutMs); with waiter_agent / waiter_session the wait joins the deadlock guard and a cycle answers 409 {circular_wait: true, chain}. Reading a terminal result here cancels the parent's pending [subagent attention] wake.

GET /spawn-list?parent_agent=… ​

{tasks: [entry]} — every task this agent spawned since server start (the store is in-memory).

POST /spawn-cancel ​

Body: task_id, optional requesting_agent (an agent must be the spawning agent — otherwise 403; without it the request acts as a person and the parent reads stopped by the user), optional reason. Cancels the task — a running turn is aborted, a task still waiting in its session's queue is taken out of it — and cascades to child spawns; returns {cancelled: [task_ids], skipped: [{task_id, state}]} (tasks that were already terminal are skipped). Disk artifacts stay.

GET /a2a/ask-result ​

Outcome of an agent_ask call by call_id — backs the agent_ask_result tool.

A call whose asker stopped waiting (agent_ask returned pending) does not depend on anyone remembering to poll: when the answer lands, the asker is woken in the session it asked from, with the first lines and the call_id. Reading the result — here or through the tool — cancels that wake, and a caller still on the line never gets one, because the answer reaches it as its tool result.

A done call may be followed later by a follow-up: when the target answered "I am working on it and will report back" and the work it started while answering — sub-agents, calls of its own, and whatever those start — finishes after its reply, the asker receives the outcome once, as an ordinary message from the target into the session it asked from. That message is a normal turn with origin: { kind: "agent", from: { agent, session }, callId } where callId is this call_id and text is the follow-up; its frame is the [Follow-up on the question you sent earlier (call_id "…"): …] note quoted under user_message in GET /chat/stream. Nobody waits for it, so DELETE /chat/queue/:id on it wakes no one. It is not sent when the asker was a person, when the call did not finish done, when the target already wrote to the asker itself in the meantime, when the asker read this route during the wake grace, or after a server restart (a restart instead wakes the asker once with a failure note, see agents.md); a follow-up whose reporting turn failed says so, with the error, in place of the text.

GET /a2a/ask-result?call_id=<uuid>
GET /a2a/ask-result?call_id=<uuid>&wait_until_done=1&timeout_ms=300000
GET /a2a/ask-result?call_id=<uuid>&agent=<target>&session=<slug>    # after a restart
json
{ "call_id": "…", "state": "done", "target_agent": "<other-agent>",
  "target_session": "20260906-172957_research", "started_at": 1788…,
  "finished_at": 1788…, "response": "…", "outcome": "completed", "source": "registry" }

state is queued (behind another turn on the target session), running, done or failed. A failed call carries error; the value stopped by the user means a person pressed Stop on the target's turn, and removed from the queue by the user before it started that a person took the waiting call out of the target's queue — the target never saw it. Neither is something to retry on the agent's own. The live registry is fed by /chat/send-sync and by /chat/send when it carries an agent_ask_call_id; when it has no record (server restarted since the call) pass agent + session and the route reads the target's JSONL (user_message.agent_ask_call_id) — source: "history", and state becomes unknown when the turn never reached turn_end. With wait_until_done the request blocks until the call finishes or timeout_ms passes; waiter_agent / waiter_session register the wait in the deadlock guard, and a cycle answers 409 with circular_wait: true like /spawn-result.

GET /a2a/turn-origin/:agent/:session ​

Who started the turn currently running on agent/session: the A2A asker (from_agent/from_session of the live turn, kind: "a2a"), for a sub-agent session the spawning parent from its spawn meta (kind: "subagent"), or — when the turn is an [agent answer] wake — the agent and session whose answer woke it (kind: "wake"), so a reply written from the wake turn goes back to the conversation it belongs to.

json
{ "origin": { "agent": "<other-agent>", "session": "20260906-172957_research", "kind": "a2a" } }
{ "origin": null }

agent_ask calls this when its session argument is omitted: if the target is the origin agent, the message goes back to the origin session instead of main (logged as agent_ask.session_inferred, reported as session_inferred: true in the tool result). An explicit session always wins. Every agent_ask result names the rule that applied as routing_reason: explicit, reply_to_origin, or default_main — the last with a routing_note when the caller sits in a non-main session, since a turn started by tmux, sentinel, a wake or a person has no origin and its report would land in the target's main unannounced.

GET /chat/stream ​

Server-Sent Events stream for a single (agent, session). Subscribe once per chat window you want to display.

bash
curl -N "https://<host>:18737/chat/stream?agent=<your-agent>&session=main"

Event types:

  • chat — {state: 'delta'|'final', text} — streaming assistant output

  • agent — {phase: 'start'|'end', usage?, contextWindow?, provider?, model?, thinking?, fallback?} — turn lifecycle around the model call. usage carries tokens_in, tokens_out, optional tokens_in_cached, tokens_out_reasoning (+_estimated) and context_tokens. Read them as two different things: tokens_in/tokens_out are what the turn SPENT, summed over every request it made, so a tool-using turn exceeds the window several times over. context_tokens is OCCUPANCY, the prompt size of the turn's last request, and is the only one to compare against contextWindow. All three engines report it; a client should still treat it as optional.

  • user_message — {text, ts, turnId?, origin?, input?, from_agent?, from_session?, from_system?, agent_ask_call_id?, steer?, steer_id?} — broadcast when a turn's user_message is written to JSONL. steer: true marks a message that was handed into the running turn turnId (see Steering under POST /chat/send); steer_id pairs it with the steer_queued event and the send response. input is {modality?: 'text'|'voice', source?: 'stt'|'realtime'} when the turn was not typed — stt is the microphone button, realtime a sentence from a standing call — so the live bubble and the one rebuilt from history render the same way. Self-typed sends, A2A inbounds, and system wakes all flow through here; from_system is one of sentinel, tmux, subagent, job, browser, voice, a2a, and web, mobile and TUI draw each of them as a divider rather than a bubble. turnId lets a sending client match this event to the optimistic bubble it rendered after POST /chat/send (which echoes the same id).

    text is the record of the turn: what a person typed, what an agent asked, what a trigger's prompt says, or the one-line statement of a wake-up. It is what the dream phase learns from, what recall is built from and what a reader sees. Everything the model is told about the turn — who wrote it, why it arrives now, what to do with it — is the turn's frame, and the frame travels in the stored row's ephemeral field beside the memory-recall block (frame first, then the recall block), never in text. Every engine puts ephemeral in front of the user message, and it replays byte-identically on later turns, so caching is unaffected. Per origin the frame is: the `[Message from agent <name>, session

    <slug>]` header and, for a session `agent_ask` created, the `[Your session '<slug>' was just created by …]` note, or for a follow-up on an earlier call the `[Follow-up on the question you sent earlier (call_id "…"): …]` note (`agent`, see below); the evidence block of a sentinel fire (`sentinel`); the inspect-now instructions of a tmux wake (`tmux`); the take-a-fresh-snapshot advice of a browser hand-back (`browser`); the call framing of a voice consult (`voice`); and the "read the whole answer with `agent_ask_result(…)`" / "fetch the full answer with `subagent_result(…)`" instruction of a wake-up (`wake`). `human` and `subagent` turns carry no frame. The SSE event carries `text` and `origin`; `ephemeral` is on the stored row only.

    origin says where the turn came from as one value. It is present on the SSE event and on the stored user_message row of every turn; rows recorded by earlier releases have none, so a client keeps reading from_* as the fallback.

    ts
    origin?:
      | { kind: 'human';    via: 'chat' | 'voice-stt' }
      | { kind: 'agent';    from: { agent: string; session?: string }; callId?: string }
      | { kind: 'subagent'; parent?: { agent: string; session: string }; taskId?: string; depth: number }
      | { kind: 'sentinel'; triggerId: string; taskId: string; triggerName?: string }
      | { kind: 'tmux';     tmuxSession: string; tmuxKind?: string }
      | { kind: 'browser';  viewId: string; cause: 'handoff' | 'activity'; handoffId?: string }
      | { kind: 'voice';    callId?: string; consultId: string }
      | { kind: 'wake';     about: 'a2a' | 'subagent' | 'job'; ref: string; depth?: number };

    human is a person typing (chat) or dictating (voice-stt). agent is an agent_ask — or a follow-up on one: when the target answered and work it started while answering (sub-agents, calls of its own) finishes later, the target's session sends the asker one more turn with origin.kind: "agent" and callId = the original call. Its text is the follow-up itself; its frame, beside the [Message from agent …] header, is [Follow-up on the question you sent earlier (call_id "<id>"): the work it started has finished. Below is its result — treat it as the answer to that question. Continue whatever depended on it; if nothing does, tell your human in one line.]. It is broadcast as a user_message like any agent turn; there is no separate event. subagent a sealed brief running in its own sub-… session; sentinel, tmux, browser and voice the four system triggers; wake brings something the agent started earlier back to it — a late agent_ask answer (ref = call id), a finished async sub-agent (ref = task id) or a rendered video (ref = job id) — and, with about: "subagent", a follow-up on a finished sub whose own work finished later (the text reads Task '<id>' … has a follow-up: the work it started has finished. It begins: "…", and the frame points at the follow_ups field of subagent_result). The legacy fields stay and are derived from it: agent fills from_agent, from_session and agent_ask_call_id (= callId); sentinel, tmux, browser and voice set from_system to the same word; wake sets from_system to its about; human and subagent set none of them.

  • steer_queued — {steerId, text, ts, turnId, origin} — fired when POST /chat/send with steer: true accepted a message for the turn turnId that is running. Not yet in the session file: the matching user_message with steer: true and the same steer_id follows once the engine has handed the text to the model. Other windows on the session render a pending bubble from this.

  • builder_state — {mode, phase, planPath} — a builder session's mode or phase changed, or its plan file was set (see Builder sessions).

  • todo_updated — {todos: [{content, status, priority?}], by?} — the builder replaced its task list (todo_write).

  • question_asked — {questionId, question, header?, options, multiple, expiresAt} — the builder is waiting for the person (ask_user); answer with POST …/answer.

  • question_answered — {questionId, answered} — the open question was answered.

  • turn_queued — {turnId, ahead, workId?, kind?} — fired when POST /chat/send hit a busy lock and the turn had to wait. ahead is the number of turns this one must wait for (≥1, includes the currently- running one). Static snapshot at enqueue time, not updated as the queue drains — except after a DELETE /chat/queue/:id, which re-emits it for the waiters that moved up. Clients render "queued · N ahead" until the matching user_message event arrives (= the turn is now actually running, lock acquired).

  • turn_dequeued — {turnId, workId} — a waiting entry was taken back via DELETE /chat/queue/:id, whoever queued it; both fields carry the removed entry's id. Clients drop the optimistic bubble for a typed message with that id and refresh their queue view.

  • turn_started — {turnId} — the engine opened the turn; this is the engine's own id (t-…), the one assistant_media, assistant_audio, turn_error and the session file's turn_end carry. Clients stamp the assistant bubble they are about to build with it, so artifacts that arrive after the turn closed pair to THIS turn instead of "the most recent bubble".

  • turn_error — {turnId?, message, engine} — the turn ended with an error instead of (or after) an assistant message. The status event still carries the same text (error: … / turn failed: …) for older clients; this one adds the turn id so the failure can be rendered as a block inside the right turn.

  • session_model — {model, resolved?, source} — the session's model override was set (PUT …/model) or cleared (DELETE …/model, model: null, source: "persona-default"). Sent to every subscriber of the session, because the switch is often made from outside the window showing it (an orchestrator agent, another client). Clients re-read GET …/model and update their header.

  • tool — {phase: 'call'|'result'|'error', tool, summary?, details?, error?} — tool-call events

  • thinking — {state: 'delta'|'final', text, truncated?} — the model's reasoning text, cumulative deltas like chat; the final precedes the chat final of the same turn. Only engines that surface thinking send it; thinkingContent.capture: false in config.yaml suppresses it entirely. See thinking.md.

  • engine_meta — {engine, itemType, label, summary?, payload} — engine-internal side-channel state. The canonical case is codex's todo_list (an internal plan/checklist the model updates mid-turn) — somora persists these so memory / REM-dream can read them later and clients can optionally render them. label is server-resolved (e.g. todo_list → "plan"); unknown item-types fall back to the raw itemType. payload is the engine's original event, opaque. Codex's own error notifications arrive as error — except those codex marks willRetry: a dropped stream it reconnects by itself is one reconnecting row per turn, and its switch to HTTPS after the retries is a transport_fallback row; neither is a failure. Besides engine-native items, somora emits its own: model_switch (codex thread continued under a new model), thread_recreated (the Codex thread no longer existed; a fresh one was started with the session history replayed), tools_changed (the agent's tool set changed — Abilities toggle, hub server, review loop, update — and Codex threads keep their tools, so a fresh thread was started with the history replayed), mcp_server_renamed (engine session rebuilt after somora's MCP server rename, label "session restarted"), context_compacted (history compacted after a context overflow — on codex-cli it means codex compacted its own thread and said so, which is otherwise invisible), context_trimmed (the oldest tool results in a running turn were shortened to keep the request inside the window — the turn continues, the model keeps the record that those tools already ran), attachments_unsupported (the engine cannot forward attachments, grok-cli), voice_handover (a call was handed to or from another agent), voice_spoken (what a voice call said out loud — found in sessions recorded by earlier releases), reasoning_effort_adjusted and sampling_dropped (backend rejected the parameter, turn retried without it). Each carries a human-readable payload.text.

  • model_fallback — {requested, actual, reason, hops?} (refs are provider/modelId) — the persona's primary model failed before producing anything and a configured fallback: model is answering this turn. requested is always the primary, actual the model now answering, reason the failure that triggered this hop. hops lists every model that failed so far, in order ([{model, reason}, …], primary first) — with a fallback chain (fallback: [a, b]) one event is sent per hop and the last one carries the whole chain. Sent before the fallback's first delta; the following agent phase:'end' also carries fallback (same shape) and reports the ACTUAL provider/model. Persisted to history as the same kind, so a reload keeps the marker on that turn.

  • memory — {count, topScore?, refs, fullText} — the memory recall injected into this turn: number of hits, best fused score, the source/slug refs, and the full <memory-context> block text. Sent after agent phase:'start'; count is 0 with empty refs when recall found nothing.

  • status — {msg} — connection events, error notices

  • heartbeat — current ms timestamp, every sse.heartbeatMs (20 s). The server watches these writes: one that fails or stays pending for sse.deadAfterMs (60 s) marks the stream dead — it is closed and its socket destroyed, logged as sse.disconnect {reason: 'dead'}. On the TLS listener every HTTP/2 session is also PINGed (sse.h2PingIntervalMs) and destroyed when it stops answering (sse.h2PingTimeoutMs), and every socket carries TCP keepalive — a tab that vanished without closing is gone server-side in about a minute instead of never.

Tool names are normalised through the same path the wire serializer uses — clients receive memory_search, not mcp__somora__memory_search. Tool input/output payloads ride in the details field as pretty-printed JSON.

GET /activity/stream ​

App-wide activity feed. One subscription per client gives streaming markers for every busy (agent, session) plus unread state for sessions the user hasn't viewed since new movement arrived. Distinct from /chat/stream, which is per-session and per-window.

bash
curl -N https://<host>:18737/activity/stream

Event types:

  • streaming — {agent, session, phase: 'start'|'end'} — emitted when any turn begins or ends on any session. Drives multi-agent streaming-dots in clients that aren't subscribed to every per- session /chat/stream.
  • turn — {agent, session, unreadAt} — a new unread-candidate event landed in the session's JSONL. Unread candidates are:
    • chat:final (assistant answer)
    • user_message with from_agent set (A2A peer wrote to us)
    • user_message with from_system set (sentinel, tmux, subagent, job, browser, voice or a2a woke us) Plain self-typed user messages, tool / memory / engine_meta events, and lifecycle (agent:start, agent:end) are excluded.
  • seen — {agent, session, seenAt} — broadcast when any client POSTs /sessions/:agent/:session/seen. Sibling clients clear their unread badge for that session.
  • status — {msg} — connection lifecycle.
  • heartbeat — current ms timestamp, every sse.heartbeatMs; same liveness rules as /chat/stream.

Unread state is persisted in the session's meta as unreadAt and seenAt (both ISO timestamps). A session is unread when unreadAt > seenAt (or seenAt is null). The persistence path is authoritative, so a server restart preserves the badge state.

POST /sessions/:agent/:session/seen ​

Tell the server "I am looking at this session now". Updates seenAt in the session's meta and broadcasts a seen event on /activity/stream so other open clients clear their badge live.

bash
curl -X POST https://<host>:18737/sessions/<your-agent>/main/seen

Optional body:

json
{ "ts": "2026-05-27T14:00:00.000Z" }

ts lets a client claim an older "I last looked at this at …" time — useful for scrolling-into-view triggers where the wall-clock isn't exactly "now". The server clamps to max(currentSeenAt, ts), so a later arrival can never regress the seen marker.

Response:

json
{ "ok": true, "agent": "<your-agent>", "session": "main", "seenAt": "2026-…" }

GET /chat/history ​

Snapshot of past events for a session. The TUI and web both hydrate this on open.

bash
curl "https://<host>:18737/chat/history?agent=<your-agent>&session=main"

Rows are the session's JSONL events in order. A turn with thinking carries one thinking_message row ({kind, ts, engine, text, truncated?}) directly before its assistant_message; clients fold it onto that bubble.

bash

Pagination: pass ?limit=200 to get the last 200 events plus a hasMore + oldestTs cursor; subsequent calls supply ?before=<oldestTs>&limit=200 for older windows.

json
{
  "agent": "<your-agent>",
  "session": "20260511-093251_research-notes",
  "events": [...],
  "hasMore": true,
  "oldestTs": 1715512000000
}

Event kinds: user_message, assistant_message, thinking_message, tool_call, tool_result, engine_meta, model_fallback (precedes the assistant message the fallback model produced), assistant_audio, assistant_media, project_switched, error, turn_start and turn_end. Each carries kind, ts, and kind-specific fields. Tool names are normalised here too. engine_meta rows preserve the raw itemType + opaque payload; clients resolve the friendly label on render (see setup.md).

POST /chat/abort ​

Cancel an in-flight turn on a (agent, session). The TUI fires it on ESC; web and mobile fire it from the Stop button overlaid on the streaming assistant bubble. Idempotent — returns aborted: false when no turn is running. Cancels the currently-running turn only; queued waiters keep their slots and still execute — a waiting entry is removed with DELETE /chat/queue/:id instead.

It stops whatever is running on the session, regardless of what started it — a command the turn is running through exec is killed with it (process group, SIGTERM then SIGKILL), on the in-process engine and in the MCP child alike, and its result says killed: the turn was stopped: a typed message, an agent_ask from another agent, a sub-agent brief, a sentinel fire, a tmux or browser wake, a voice consult or a wake-up. Whoever asked for that turn learns why it ended: an agent_ask still on the line, agent_ask_result and subagent_result report state: "failed" with error: "stopped by the user", and a sentinel fire is recorded with outcome error and the same text. The tool descriptions tell the asking agent not to retry that on its own.

bash
curl -X POST "https://<host>:18737/chat/abort?agent=<your-agent>&session=main"

Memory ​

Each agent has its own memory layer — markdown notes under ~/.somora/agents/<agent>/memory/ plus, optionally, a shared Obsidian vault and wiki layer. See memory.md for the storage architecture.

GET /agents/:agent/memory/notes ​

List indexed notes for an agent. Optional ?source=memory|vault|wiki to filter by source layer.

POST /agents/:agent/memory/recall-preview ​

What auto-inject would recall for a message in a given conversation state — the turn's own recall path (query construction, history blend, search, block budget) without running a turn. Body:

json
{
  "text": "und wer gehört sonst noch zur familie?",
  "history": [
    { "kind": "user_message", "text": "was kannst du mir über karl erzählen?" },
    { "kind": "assistant_message", "text": "Karl ist …" }
  ],
  "autoInject": { "historyWeight": 0.4 }
}

history is optional (oldest first; only user_message and assistant_message entries count). autoInject optionally overrides any memory.autoInject knob for this call only — for measuring a setting before changing config.yaml. Response: hits with slug, source, score, vecScore, bm25Score, line range, plus injectedCount and ephemeralContextChars. Used by the recall replay harness; loopback-only like every debug route.

Hybrid (BM25 + vector) search over the agent's own memory notes plus the shared vault/wiki index — the same MemoryManager.search the memory_search tool and auto-inject use: filler words are dropped from the keyword side, a page whose slug names a query word is boosted (memory.hybrid.slugMatchBoost), see memory.md. No history blend here — that is auto-inject's, use recall-preview above to see a turn's actual recall.

bash
curl "https://<host>:18737/agents/<your-agent>/memory/search?q=voice+satellites&limit=5&minScore=0.3"

Query params: q (required), limit (1–50, default 5), minScore (0..1, default 0 — the route shows everything; auto-inject applies autoInject.minScore).

Returns { agent, query, limit, minScore, count, hits: [...] }; each hit carries slug, source (memory | wiki | vault), score (fused, min-max normalised within this query's candidates — a rank, not a similarity), vecScore, bm25Score, startLine, endLine, filePath and the chunk text.

For full content of a hit, agents call memory_get via the tool endpoint (POST /agents/:a/tools/memory_get). Same path is available to your client.


Wiki explorer ​

Read-only browse surface over the shared wiki, backing the web client's wiki window. All routes return 503 unless wiki.enabled and obsidian.vault are both configured.

Pages are addressed by slug, never by path — a request can only name pages the index already found under the wiki root, so traversal attempts come back as a plain 404 rather than needing a filter to catch them.

GET /wiki/status ​

{ enabled: boolean, root?: string }. Cheap enough to call on UI mount; clients use it to decide whether to show the wiki entry point at all.

GET /wiki/tree ​

json
{
  "root": "/path/to/vault/somora",
  "pages": 262,
  "builtAt": 1784750000000,
  "nodes": [
    { "type": "dir", "name": "personen", "path": "personen",
      "children": [
        { "type": "page", "slug": "personen/familie-klein",
          "name": "familie-klein.md", "title": "Familie Klein",
          "description": "…", "mtimeMs": 1784700000000 }
      ] }
  ]
}

Titles come from the page's first # H1, falling back to frontmatter title, then the filename.

GET /wiki/page?slug=<slug> ​

Returns markdown plus resolved relationships:

json
{
  "slug": "projekte/somora", "title": "somora", "folder": "projekte",
  "mtimeMs": 1784700000000,
  "markdown": "## Aktueller Stand\n…",
  "frontmatter": { "type": "project", "created": "2026-05-08" },
  "links":       [{ "slug": "personen/jane-doe", "title": "Jane" }],
  "backlinks":   [{ "slug": "agenten/<your-agent>", "title": "Your Agent" }],
  "unresolved":  ["personen/familie-doe"],
  "linkTargets": { "personen/jane-doe": "personen/jane-doe",
                   "familie-doe": null }
}

linkTargets maps every raw [[target]] in the body to a slug, or null when nothing matches. Resolution — exact slug, case-insensitive slug, unique basename — lives here so clients don't reimplement Obsidian's matching. An ambiguous basename resolves to null on purpose: guessing one of several same-named pages fabricates a relationship.

404 when the slug names no page.

GET /wiki/graph?scope=local&slug=<slug> · ?scope=global ​

json
{
  "scope": "local",
  "nodes": [{ "id": "projekte/somora", "label": "somora",
              "folder": "projekte", "degree": 41 }],
  "edges": [{ "from": "agenten/<your-agent>", "to": "projekte/somora",
              "type": "wikilink" }],
  "truncated": false
}

local returns the page, everything it points at, everything pointing at it, and the edges among those neighbours. global returns the whole wiki, capped at the 400 most-connected pages — truncated says whether the cap bit. degree always counts the full wiki, so a node stays recognisable as a hub inside a local view.

index.md is excluded from both scopes: it links to every page by construction, which makes it a table of contents rather than a relationship.

Edge type is wikilink for inline [[links]] and related for frontmatter related: entries.

POST /wiki/refresh ​

Drops the cache and re-scans. The index otherwise caches for 10 seconds and then re-parses only files whose mtime or size changed, so ordinary edits appear without this call.


Dream system ​

The dream system runs in three phases (REM, Deep, Lucid) — see dream-phases.md for the model. Some of these endpoints surface state for monitoring; others trigger phases manually.

GET /dream/loop-state ​

Read-only snapshot of the active Lucid review loop, if any.

json
{ "active": true, "agent": "<your-agent>", "dreamId": "lucid-...",
  "startedAt": "…", "lastActivityAt": "…" }

Returns { active: false } when no loop is held.

GET /dream-states ⚠ experimental ​

Per-agent REM state + server-global Deep / Lucid state. Drives the web AgentDock pulse indicators and REM badges.

json
{
  "rem": {
    "<your-agent>":  { "active": false, "pendingCount": 3 },
    "<agent-b>":  { "active": true,  "pendingCount": 0 }
  },
  "deep":  { "active": false },
  "lucid": { "active": false, "pendingRuns": 0, "pendingFindings": 0 }
}

rem[<agent>].active is filesystem-driven (presence of <agent>/memory/.dreams/<id>.dream.running.md). rem[<agent>].pendingCount is the number of completed REM dreams waiting for review (<id>.dream.md files in .dreams/, not in processed/).

POST /dream/run-deep ​

Trigger a Deep run manually. Default is fire-and-forget (returns immediately); set {"wait": true} to wait for the run to finish and get the result inline.

bash
curl -X POST https://<host>:18737/dream/run-deep \
     -H 'Content-Type: application/json' \
     -d '{"wait":true, "force":false}'

force: true bypasses the per-agent skip-cache so every memory file gets re-evaluated.

POST /dream/run-lucid ​

Same shape as run-deep, for the Lucid (wiki review) phase. While a previous run still has findings waiting for review, no new run starts (the response names that run); {"force": true} runs anyway.

POST /wiki/migration/plan ​

Step one of moving a grown wiki onto the folder template (wiki.md): read the whole wiki and write down what a migration would do — nothing in the wiki is touched. The plan lands under ~/.somora/wiki-migration/<id>/plan.md (and plan.json). 400 when the wiki layer is off.

json
{ "id": "20260929-104047", "plan": "/home/you/.somora/wiki-migration/20260929-104047/plan.md",
  "pagesTotal": 991, "foldersTotal": 71,
  "summary": { "move_folder": { "items": 19, "pages": 40 }, "unite_twins": { "items": 19, "pages": 38 },
               "fold_report": { "items": 190, "pages": 190 }, "review_pages": { "items": 37, "pages": 447 },
               "describe_folder": { "items": 0, "pages": 0 }, "unclear": { "items": 0, "pages": 0 } },
  "durationMs": 1830 }

POST /wiki/migration/refine ​

Step two: the Lucid model (falling back to the Deep model) judges every page the plan is unsure about, and every page a rule would move — keep, move to a folder, fold into an existing page, or unclear — in batches of 25 pages. Body { "id": "<plan id>", "wait": false, "batchSize": 25 }. Runs in the background by default and writes refined.md / refined.json next to the plan; wait: true returns the result inline. Still no write into the wiki. 404 for an unknown plan, 409 while a refine for that plan is running.

json
{ "id": "20260929-104047", "started": true, "message": "Refine started in background. GET /wiki/migration/plans/:id for progress." }

GET /wiki/migration/plans/:id ​

The plan's summary, the progress of a running refine or execute (done / total), the approvals, and the refined result's groups when it exists.

json
{ "id": "20260929-104047", "dir": "…/wiki-migration/20260929-104047",
  "plan": { "pagesTotal": 991, "foldersTotal": 71, "summary": { … }, "createdAt": "…" },
  "refine": { "started": "…", "done": 12, "total": 27 },
  "execute": null,
  "approvals": { "planId": "…", "groups": { "move:regeln": { "status": "approved", "at": "…" } }, "twins": {} },
  "refined": { "model": "…", "pagesJudged": 943, "batchesTotal": 38, "batchesFailed": 0,
               "groups": [ { "action": "move", "target": "regeln", "pages": 50 }, … ] } }

POST /wiki/migration/reindex ​

One full sweep of the shared search index now (somora wiki migrate undo calls it after putting a backup back). Returns { indexed, skipped }; 503 while the index is still building.

POST /wiki/migration/plans/:id/relink ​

A second pass over every page with the renames a finished run of this plan recorded: [[links]] and frontmatter related: entries that still name a moved, folded or united page are pointed at its new place. Idempotent. {"dryRun": true} only counts. Runs before v2026.09.29.07 left related: untouched — this is the repair.

json
{ "id": "…", "dryRun": false, "renames": 819, "refsRewritten": 393, "pagesTouched": 224 }

POST /wiki/migration/plans/:id/approve ​

Mark groups of the refined plan. Body: groups (group keys such as move:regeln or fold:projekte/somora), or action (move / fold / unclear — every group of that action), or twins (names, or "all"); status is approved (default), dismissed or pending. 404 until the plan has a refined result.

json
{ "id": "…", "status": "approved", "touched": 3, "approvedGroups": 12, "approvedPages": 310, "approvedTwins": 19 }

POST /wiki/migration/plans/:id/execute ​

Run the approved part. Default {"dryRun": true} writes dry-run.md next to the plan and touches nothing. A real run needs {"dryRun": false, "confirm": "move my wiki"}, copies the whole wiki into backup-<time>/ under the plan first, and refuses to start when the copy is incomplete. Background by default (wait: true for the result inline); 409 while a run for that plan is going. See wiki.md for what a run does.

json
{ "id": "…", "dryRun": false, "report": "…/execution-20260929-131500.md",
  "counts": { "move": 207, "fold": 540, "unite": 19, "failed": 2, "skipped": 0 },
  "linksRewritten": 812, "foldersRemoved": 51, "backupDir": "…/backup-20260929-131412",
  "reindex": { "indexed": 610, "skipped": 380 } }

POST /agents/:agent/dream/run-rem ​

Catch up one agent's unread conversations now: starts the REM cycle the idle timer would start — resume a paused dream, else read the sessions with unread events one after another — without waiting for silence. Always in the background; chat activity pauses it as usual.

json
{ "agent": "<your-agent>", "outcome": "started", "started": true, "message": "…" }

outcome is started, busy (a cycle is already running) or nothing_to_do. 400 when REM is not enabled for the agent, 404 for an unknown agent, 409 when REM was enabled after the server started.


Projects ​

Curated pointer-file manifests linking a session to a real-world project. Opt-in feature — every route below returns 503 when projects.enabled is false in config.yaml. Clients should probe GET /projects/feature once at boot to decide whether to surface the feature at all. See projects.md for the user-level model.

GET /projects/feature ​

Feature-flag probe. Always returns 200, regardless of the configured state — clients use this to detect availability without ambiguity (empty entities array vs. feature off).

json
{ "enabled": true, "entityCount": 2 }

GET /projects/entities ​

The controlled entity vocabulary from config.projects.entities.

json
{
  "entities": [
    { "slug": "privat", "label": "Privat" },
    { "slug": "acme", "label": "acme GmbH" }
  ]
}

Agents call this before project_create when they're uncertain about an entity name they heard via STT — the response is the canonical match list. Direct clients fetch it once to populate filter dropdowns.

GET /projects ​

List configured projects.

Query params (all optional):

  • entity=<slug> — filter to one entity
  • tag=<string> — filter to projects whose tags[] contains this
  • includeArchived=true — include soft-deleted projects (hidden by default)
json
{
  "total": 3,
  "projects": [
    {
      "slug": "heimkino",
      "name": "Heimkino",
      "entity": "privat",
      "description": "Receiver, beamer, …",
      "color": "#4f46e5",
      "tags": ["hardware", "wip"],
      "created": "2026-04-15T10:23:00Z",
      "updated": "2026-05-13T09:42:00Z",
      "archived": false,
      "paths": [
        { "ref": "~/code/heimkino", "label": "Sourcecode" },
        { "ref": "https://drive.google.com/..." }
      ]
    }
  ]
}

GET /projects/:slug ​

Full project file content. 404 if the slug doesn't exist.

bash
curl https://<host>:18737/projects/heimkino

POST /projects ​

Create a new project. Returns 201 with the full project on success.

Body:

json
{
  "slug": "heimkino",
  "name": "Heimkino",
  "entity": "privat",
  "description": "Receiver, beamer, acoustic treatment",
  "color": "#4f46e5",
  "tags": ["hardware", "wip"],
  "expires": null,
  "paths": [
    { "ref": "~/code/heimkino", "label": "Sourcecode" },
    { "ref": "https://drive.google.com/..." }
  ],
  "workdir": "~/code/heimkino"
}

workdir (optional) is the project's working directory — pinning the project to a session makes it that session's working directory (see projects.md).

Validation:

  • slug must match [a-z0-9_-]+ and be unique (409 on collision)
  • entity must match one of config.projects.entities[].slug (400 with the available list when unknown)
  • each paths[].ref must be scheme-recognised — https://..., ~/abs, /abs, or <resource-slug>:/path where the slug exists in config.resources (400 with the available list when the resource is unknown)

PATCH /projects/:slug ​

Transactional multi-op update. All ops validate first; if any one fails, nothing is written. Returns 200 with the updated project on success.

Body:

json
{
  "ops": [
    { "op": "add_path", "ref": "~/research/atmos.md", "label": "Atmos notes" },
    { "op": "set_field", "field": "description", "value": "Updated" },
    { "op": "set_tags", "tags": ["hardware", "wip", "avr"] }
  ]
}

Supported op shapes:

opRequired fieldsEffect
set_fieldfield ∈ {name,description,color,expires}, value (string or null)Update top-level field; null clears optional fields (cannot clear name).
add_pathref, label?Append a pointer. Same scheme validation as POST /projects.
remove_pathrefRemove by exact ref match. 400 if ref isn't in the list.
set_tagstags: string[]Replace the full tag array.
archivereason?Soft-delete.
unarchive—Restore.

Slug and entity are intentionally not mutable in v1 — would break session pins. Workaround: delete the file and recreate.

GET /agents/:agent/sessions/:session/project ​

Current pinned project for a session. Always returns 200 (or 404 if the agent/session itself doesn't exist).

json
{
  "agent": "<your-agent>",
  "session": "main",
  "slug": "heimkino",
  "project": { … full ProjectInfo … }
}

When no project is pinned: { "agent": …, "session": …, "slug": null, "project": null }. When the slug is set but the file is missing on disk: { …, "slug": "ghost", "project": null, "missing": true }.

POST /agents/:agent/sessions/:session/project ​

Pin a project to a session.

bash
curl -X POST https://<host>:18737/agents/<your-agent>/sessions/main/project \
     -H 'Content-Type: application/json' \
     -d '{"slug":"heimkino"}'

Returns { agent, session, previousSlug, currentSlug }. Re-pinning the same project is a noop — no project_switched event is written to the JSONL.

Emits an SSE project event to every subscriber of (agent, session).

DELETE /agents/:agent/sessions/:session/project ​

Clear the pin. Returns { agent, session, cleared: true, previousSlug }. Also emits an SSE project event.


Realtime voice ​

Talking to an agent: a standing, interruptible call where a realtime model speaks and the agent knows. Off unless realtimeVoice.enabled. Concept and configuration in realtime-voice.md; this is not the dictation/TTS path (voice.md).

GET /voice/status ​

json
{ "enabled": true, "provider": "openai", "model": "gpt-realtime-2.1-mini",
  "maxCallMinutes": 20, "agents": ["<your-agent>", "<other-agent>"], "calls": [] }

agents lists who may be called — realtime voice on, and the agent's agent.yaml carries voice.enabled: true. With the feature off the answer is { "enabled": false, "agents": [] } (200, not an error: a client asks this to decide whether to show the tile at all).

GET /voice/instructions?agent=<name>&session=<slug> ​

Exactly what the speaking model would be told, without starting a call:

json
{ "agent": "<your-agent>", "session": "main", "text": "You are <your-agent>, speaking out loud …",
  "chars": 1528, "source": "derived", "voice": "ash", "language": "de",
  "consultPolicy": "always" }

source is derived (built from the persona plus agent.yaml voice:) or VOICE.md when the operator wrote one. The derived text exists only in memory — this is the only way to read it. 503 when voice is off, 404 when that agent has no voice.

WS /voice/attach?agent=<name>&session=<slug> ​

The call itself. One JSON frame format in both directions; the browser holds no provider knowledge, no key, and never sees a tool call.

Client → server:

json
{ "type": "audio", "base64": "<PCM16 24 kHz mono>" }
{ "type": "interrupt" }
{ "type": "hangup" }

Keep sending audio while nobody speaks. The provider ends a turn on SILENCE, not on missing packets — a client that stops sending gets one "speech started" and then nothing at all.

Server → client:

json
{ "type": "ready", "call": { … }, "rateHz": 24000 }
{ "type": "audio", "base64": "…", "rateHz": 24000 }
{ "type": "state", "call": { "state": "consulting", "consults": 2, "target": { … } } }
{ "type": "event", "event": { "kind": "user_transcript", "text": "…", "final": true }, "call": { … } }

state is connecting | listening | consulting | speaking | closed. consulting is somora running a real turn in the bound session — no provider event announces it. target changes when a call is handed to another agent, and the client follows it.

Closing the socket ends the call: a standing connection bills by the minute. Work the agent already accepted keeps running — hanging up and cancelling are two different things.

Attachments ​

Files attached to chat turns (images, PDFs, plain-text snippets) travel via this two-step flow: upload first, then ref the hash on /chat/send.

POST /attachments ​

Upload one file as the raw request body, with its name in the X-Somora-Filename header (URL-encoded, so spaces and non-ASCII survive). Content is stored once on disk, deduped by hash.

Raw bytes rather than a form upload is deliberate: multipart parsing pulls the whole file into memory, while a raw body streams and keeps the size cap meaningful. A multipart/form-data request is refused with 415 rather than stored as an opaque text attachment.

bash
curl -X POST https://<host>:18737/attachments \
     --data-binary @./screenshot.png \
     -H "Content-Type: image/png" \
     -H "X-Somora-Filename: screenshot.png"

Returns one object (the type is sniffed from the bytes, not the name):

json
{ "hash": "sha256-…",
  "name": "screenshot.png",
  "mime": "image/png",
  "kind": "image",
  "size": 184320 }

Collect the objects for as many files as you need and pass them as attachments[] on the next /chat/send or /chat/send-sync. The per-turn count and per-file size come from config.attachments.

GET /attachments/:hash ​

Serve the bytes for a previously-uploaded attachment. Useful for clients that want to preview the same image the agent saw.


Tmux integration ​

somora knows about tmux sessions on the host and lets clients attach to them through a WebSocket bridge. See tmux.md for the full model.

GET /tmux/sessions ​

List live tmux sessions, joined with somora's origin store so each session carries the agent/session that created it (if known).

bash
curl https://<host>:18737/tmux/sessions

WS /tmux/attach?session=<name> ​

WebSocket bridge to tmux attach-session -d -t <name>. Binary frames carry the terminal stream; text frames carry control messages ({type:'resize',cols,rows}).

WS /terminal/attach ​

Fresh shell (no tmux session). Same binary/text frame protocol as /tmux/attach.


Browser — shared managed Chromium ​

Served by the somora web server. See browser.md. Mutating HTTP routes answer 503 while browser.enabled is false; status and the change stream return {enabled:false,browsers:[],warnings:[]}. The viewer WebSocket refuses attachment when disabled.

GET /logs ​

The server's own log, for the log window in the web client. Query: day (YYYY-MM-DD, default the newest file), minLevel (pino numbers: 20 debug, 30 info, 40 warn, 50 error), q (case-insensitive substring over the raw line), agent, limit (default 300, max 2000).

Returns { day, days, lines: [{ ts, level, msg, agent?, session?, fields }], offset, truncated }. days lists every day that has a file, newest first. offset is a byte position to continue from. truncated says older lines were outside the read window.

Only the tail of one day's file is read (512 KiB), so the cost does not grow with the log — the directory routinely holds hundreds of megabytes. The caller names a day, never a path.

GET /logs/since ​

?offset=<n> plus the same filters. Returns { lines, offset, day } with only what was appended after offset — the follow path for the log window. A file that shrank (rotation, truncation) snaps the offset back to its real size instead of reading backwards.

POST /browser/op ​

Run one browser tool operation for an agent, in that agent's window — the path the MCP tool child takes for claude-cli/codex-cli turns, because the Chromium lives in the server process. Body { agent, session?, input } where input is the tool's argument object ({ op: "open", url }, { op: "snapshot", tab }, …). Returns the tool's result object (ok, error, hint, tab, snapshot, …), never a non-2xx for a tool-level refusal.

GET /browser/status ​

{ enabled, headed, browsers: [{ view_id, agent, browser_id, profile, ephemeral, state, control, human_by?, handoff?, tabs: [{ tab_id, url, title, agent, session?, generation, emulation? }], last_used, headed? }], warnings } — one entry per open window, not per Chromium process: view_id is <browser_id>@<agent> and agent owns it. Agents sharing a profile run in one process (browser_id) with one window each, own tabs, own control state, own handoff. Stopped browsers are not listed, except one still holding a pending handoff. control is agent_control, handoff_requested, human_control or paused. headed is the host plan for browser.headed (headless, display, xvfb, unavailable); warnings lists what the operator must fix (no Chromium found, headed configured but no display and no Xvfb). emulation shows a tab's device / locale from open.

GET /browser/stream ​

SSE events browsers contain {enabled,browsers,warnings} with the same browser entries as /browser/status. Every connection starts with a complete snapshot; later events replace it. heartbeat follows the normal SSE liveness settings. Slow readers receive coalesced snapshots; stalled writes terminate the connection.

POST /browser/:id/restart ​

Explicitly reopen a known stopped browser with a blank page and the same profile. :id may be a window id or a browser id — a restart is per process; every window that existed comes back with its control state. No old navigation or input is replayed. Pending handoffs remain pending. Returns {ok:true} or 409 with {error} when recovery is refused (unknown browser or removed shared-profile configuration).

POST /browser/:id/control ​

:id is a window id (profile:team@<your-agent>); a bare browser id works while only one agent has a window on that process, and is otherwise refused as ambiguous. Body { mode: "human" | "agent", by?, handoffId? }. human takes control of that window: its agent's ops are refused until handed back, while other agents on the same process keep working. agent hands it back; with a pending handoff the requesting agent is woken once in its session (pass the handoffId from the status so a stale button press after a newer handoff does not wake twice). Without a pending handoff, a hand-back after real activity wakes the window's own agent, in the session of the last tab it used there. 404 when the browser is not running, 409 on a state conflict. A mismatched handoff ID is refused; a duplicate completed ID does not release a newer manual takeover. With by, a different current controller cannot be handed back. The web client sends control over its viewer WebSocket so the identity matches subsequent input. Wake dispatch is not a durable queue; see the recovery limitations in browser.md.

GET /browser/attach (WebSocket) ​

?view=<viewId>&tab=<tabId>&viewer=<id> — the live view behind the web client's browser window. viewId is <browser_id>@<agent>; only that window's tabs are reachable, and browser= is still accepted as the parameter name. Control, input and hand-back over this socket act on that window. Binary frames carry one JPEG each: [u32 BE header length][JSON header][JPEG], header { tabId, generation, seq, url, cssWidth, cssHeight, scrollX, scrollY, ts }. Frames are ack-paced (browser.stream.maxFps, default 15, at most 20 fps) and skipped for a viewer whose socket has more than 2 MB pending. Text frames (JSON):

  • server → viewer: ping (answer {"type":"pong"}; 80 s of silence drops the socket), ready {browser, tabId, viewerId} after attach or a tab switch, tabs {browser} whenever tabs or control changed, control {control} after a control request, notice/error{text}.
  • viewer → server: control {mode:"human"|"agent", handoffId?}, tab {tabId} (switch the streamed tab), and — only while this viewer holds human control — navigate {url} (same policy as the tool), newtab, closetab {tabId}, resize {width,height}, mousemove/click/mousedown/mouseup {x,y,button?,clickCount?} in CSS pixels of the streamed viewport, wheel {x,y,deltaX,deltaY}, text {text} (composed text incl. paste; CDP Input.insertText), key {key, ctrl?, alt?, shift?, meta?, action?} (Playwright key names, e.g. Enter, Control+a), back, forward, reload. Input from a viewer without control gets a notice, nothing is applied. Input other than navigate, newtab, closetab, tab and control must also carry frameTab and generation from the displayed frame; obsolete metadata is refused. Messages above 64 KiB and queues above 64 commands close the connection.

Close codes: 1008 bad request (unknown browser, disabled), 4000 heartbeat timeout, 4001 tab closed, 1009 oversized input, 1012 server shutdown. 1008 also covers excessive pending input.

Sentinel — proactive triggers ​

Sentinel installs time-based triggers that wake agents on a schedule. The agent does its work into its own chat session — same surface as when you interact with it directly. See sentinel.md for the conceptual overview.

The same operations are also exposed as the sentinel tool agents can call (POST /agents/:agent/tools/sentinel); the HTTP routes below are for the web-UI sentinel tab and for external clients.

GET /sentinel/triggers ​

List all triggers. Optional query filters:

  • ?owner=<agent> — only triggers whose ownerAgent matches
  • ?status=active|paused|error|completed
bash
curl https://<host>:18737/sentinel/triggers?status=active
json
{
  "count": 2,
  "triggers": [
    {
      "id": "morning-mail-summary-a7c3",
      "name": "morning-mail-summary",
      "ownerAgent": "<other-agent>",
      "source": { "type": "time", "spec": { "type": "daily", "time": "08:00" } },
      "evaluator": { "type": "none" },
      "dispatch": { "agent": "<other-agent>", "session": "morning-routine",
                    "prompt": "Check inbox via gog skill…" },
      "createdAt": "2026-05-17T11:00:00.000Z",
      "status": "active",
      "fireCount": 5,
      "lastSuccessAt": "2026-05-17T08:00:01.234Z",
      "errorStreak": 0,
      "nextFireAt": "2026-05-18T08:00:00.000Z"
    }
  ]
}

GET /sentinel/triggers/:id ​

{trigger} — the full trigger document for a single id. 404 if missing.

GET /sentinel/triggers/:id/history ​

Newest-first fire log. ?limit=N capped at 200, default 50.

json
{
  "count": 5,
  "entries": [
    {
      "firedAt": "2026-05-17T08:00:01.234Z",
      "scheduledFor": "2026-05-17T08:00:00.000Z",
      "outcome": "success",
      "taskId": "task-..."
    },
    {
      "firedAt": "2026-05-16T08:00:00.500Z",
      "scheduledFor": "2026-05-16T08:00:00.000Z",
      "outcome": "skipped",
      "skipReason": "cooldown (1320s remaining)"
    }
  ]
}

Outcomes: success / error (with error: string) / skipped (with skipReason: string). Plus optional catchUp: true (boot recovery fire) and testMode: true (fired via /test). A fire that a person took out of the session's queue before it started is skipped with skipReason: "removed from the queue by the user"; one stopped while running is error with stopped by the user.

404 once the trigger is deleted (its history file goes with it).

POST /sentinel/triggers/:id/pause ​

Set status to paused. Trigger stops firing until explicitly resumed. Returns {"ok": true}. 404 if missing.

POST /sentinel/triggers/:id/resume ​

Set status back to active, recompute nextFireAt from the spec. Idempotent for already-active triggers. 404 if missing.

POST /sentinel/triggers/:id/test ​

Fire NOW, bypassing cooldown and daily-cap. The fire is recorded with testMode: true in history. The dispatched agent receives the same evidence-prefixed prompt as a real fire.

bash
curl -X POST https://<host>:18737/sentinel/triggers/morning-mail-summary-a7c3/test

DELETE /sentinel/triggers/:id ​

Remove the trigger and its history file. Idempotent — already-deleted ids return 404.

GET /sentinel/status ​

Scheduler diagnostic snapshot. Useful when sanity-checking that the scheduler is armed for the next due trigger.

json
{ "started": true, "nextFireAt": 1779013800000 }

Creating triggers ​

Triggers are created through the agent-facing tool, not a dedicated HTTP route, so the safeguards (min-interval, per-agent cap, limit enforcement) all run through the same validation path:

bash
curl -X POST https://<host>:18737/agents/<your-agent>/tools/sentinel \
  -H 'Content-Type: application/json' \
  -d '{
    "action": "create",
    "name": "morning-mail-summary",
    "intent": "Daily 8am inbox digest",
    "source": { "type": "time", "spec": { "type": "daily", "time": "08:00" } },
    "dispatch": {
      "agent": "<other-agent>",
      "session": "morning-routine",
      "prompt": "Check inbox via the gog skill, group by topic, tell me what is important."
    }
  }'

The full sentinel tool surface (create / list / get / pause / resume / delete / test / history / purge_completed) is described in sentinel.md.


Voice ​

Two flows: STT for filling chat drafts, TTS for spoken replies. Plus /voice/turn as the audio-in/audio-out endpoint for integrations. All routes return 503 when the matching block is missing or disabled in config.yaml. See voice.md for the full picture.

GET /stt/config ​

Reports STT availability + the default language hint.

json
{ "enabled": true, "language": "de" }

Returns { "enabled": false } when STT is off in config — clients auto-hide their mic button.

POST /stt/transcribe ​

Forwards a multipart audio recording to the configured upstream's /v1/audio/transcriptions and returns the transcript.

http
POST /stt/transcribe
Content-Type: multipart/form-data

[email protected]
language=de              # optional, overrides config default

Response: { "text": "<transcript>" }. 503 when disabled.

GET /tts/config ​

Reports TTS availability + supported wire formats + per-client auto-play defaults.

json
{
  "enabled": true,
  "formats": ["audio/wav", "audio/opus", "audio/mp4"],
  "language": "de",
  "voice": null,
  "clients": {
    "web": { "autoPlayVoiceReplies": false, "allowUserOverride": true },
    "mobile": { "autoPlayVoiceReplies": false, "allowUserOverride": true }
  }
}

POST /tts/synthesize ​

Generate (or fetch from cache) spoken audio for a piece of text. Content-negotiates the wire format from Accept.

http
POST /tts/synthesize
Content-Type: application/json
Accept: audio/opus, audio/wav;q=0.5

{ "text": "Es ist 10:29 Uhr.", "voice": null, "language": "de" }

Response body is audio bytes. Useful response headers:

  • Content-Type — audio/wav, audio/opus, or audio/mp4.
  • X-Tts-Cache — hit or miss.
  • X-Tts-Cache-Key — sha256 hex used for caching.
  • X-Tts-Duration-Ms — set on WAV cache misses; omitted otherwise (clients can compute on-play).

400 on missing text or text > 4000 chars. 502 on upstream failure. 503 when TTS disabled.

GET /tts/cache/:filename ​

Stream a previously-generated audio file by its cache key. Filenames are <64-hex>.<wav|opus|m4a>; anything else returns 400. Supports single-range requests (Range: bytes=N-) so mobile <audio> can seek.

This is the URL emitted as assistant_audio.url in SSE and JSONL — clients render it directly into <audio src=…> without ever calling /tts/synthesize for cached turns.

POST /voice/turn ​

Independent audio-in → audio-out endpoint. STT-transcribes the recording, runs a normal agent turn (with input_modality=voice), sanitizes the assistant reply for speech, generates TTS, and returns JSON with the artifact URL. The session JSONL + SSE stream see the turn live, same as a /chat/send turn.

http
POST /voice/turn
Content-Type: multipart/form-data
Accept: audio/opus, audio/wav;q=0.5

agent=<name>             # required
session=<name>           # required: "main" / exact id / new slug (creates)
[email protected]    # required
voice=<voice-id>         # optional, falls back to tts.voice
language=<lang>          # optional, falls back to tts.language

Response:

json
{
  "ok": true,
  "agent": "<your-agent>",
  "session": "main",
  "transcript": "Wie spät ist es?",
  "text": "Es ist 10:29 Uhr.",
  "audio": {
    "url": "/tts/cache/abc123….opus",
    "mime": "audio/opus",
    "durationMs": 1800,
    "cacheKey": "abc123…"
  }
}
  • Session lock priority: user (treated as human input).
  • Always generates audio, regardless of any per-chat auto-play toggle (those toggles only affect /chat/send).
  • 404 when session=<exact-id> doesn't exist; auto-creates for free- form slug names. 503 if either stt or tts is disabled.

assistant_audio SSE event ​

After a turn whose reply got TTS (auto or via /voice/turn), the session's SSE stream emits:

event: assistant_audio
data: {"turnId":"…","url":"/tts/cache/…","mime":"audio/opus","durationMs":1800,"cacheKey":"…"}

Clients pair on turnId and render a Play-button on the matching assistant bubble. The event is also appended to the session JSONL, so /chat/history returns it on reload and Play-buttons survive.


Web bundle ​

GET /web/ ​

Serves the bundled web UI from web/dist/. Same-origin as the API, so the web app's fetch('/agents') works without CORS. GET /web (no slash) redirects here.

GET /mobile/ ​

Serves the mobile PWA from web-mobile/dist/ the same way; GET /mobile redirects to it. Both bundles are static files — a custom client does not need them, every function they use is in the routes above.


Building a custom client — typical flow ​

A minimal client that wants to send a message and stream the response back works like this:

bash
# 1. Discover agents
curl https://<host>:18737/agents

# 2. Optionally set the model + thinking for this conversation
curl -X PUT https://<host>:18737/agents/<your-agent>/sessions/main/model \
     -H 'Content-Type: application/json' \
     -d '{"model":"claude-opus-4-7"}'

# 3. Subscribe to the stream (background)
curl -N "https://<host>:18737/chat/stream?agent=<your-agent>&session=main" &

# 4. Send a message — server fires the turn, events arrive on the stream
curl -X POST https://<host>:18737/chat/send \
     -H 'Content-Type: application/json' \
     -d '{"agent":"<your-agent>","session":"main","text":"Was steht heute an?"}'

For richer clients (a dashboard, a desktop app, a phone bridge), the typical loop is:

  1. On boot: GET /agents + GET /sessions + GET /version to build the navigation.
  2. For each open chat window: open one SSE subscription (/chat/stream) and one history hydration (/chat/history?limit=200).
  3. On user input: POST /attachments for any files, then POST /chat/send with attachments[] set.
  4. On user reset: POST /agents/:a/sessions/:s/reset. The server triggers REM in the background; your UI can show the resulting pending count via /dream-states.
  5. For monitoring: poll /dream-states every 30 s for dream-phase indicators, /dream/loop-state every 2 s if you want to surface the Lucid review loop.

The TUI's API client lives at src/cli/tui/api.ts; the web's at web/src/lib/api.ts. Both are short, focused, typed wrappers over the surface above and make good starting points for your own.


Files of interest in the somora source ​

  • src/server/index.ts — every route definition lives here
  • src/server/sse-serializer.ts — wire format for SSE events
  • src/server/tool-format.ts — tool name + arg + result pre-formatting (normalises mcp__… prefixes)
  • src/cli/tui/api.ts — reference client (TypeScript)
  • web/src/lib/api.ts — second reference client (TypeScript, browser-targeted)