Web client
A browser-based desktop for somora — multi-window chat with every agent on your LAN. Same backend as the TUI, just a different head.
Mental model
┌─────────────────────────────── browser ─────────────────────────┐
│ ┌─ agent dock ─┐ ┌── chat: scribe ───┐ ┌── chat: coach ────┐ │
│ │ scribe ● │ │ history + stream │ │ history + stream │ │
│ │ coach ● │ │ tool blocks │ │ tool blocks │ │
│ │ archi. ◯ │ │ [paperclip] ▢▷▶ │ │ [paperclip] ▢▷▶ │ │
│ └──────────────┘ └───────────────────┘ └───────────────────┘ │
│ │
│ taskbar: [scribe] [coach] ▢ auto-arrange 💾 save layout │
└─────────────────────────────────────────────────────────────────┘
│ HTTP + SSE
▼
somora server :18737Each chat window is a ChatProvider-managed React subtree subscribing to one SSE stream keyed by agent::session. Multiple windows for the same agent share state — open an agent's main session twice and both windows render the same live transcript.
Access
The somora server mounts the production bundle at /web/*. HTTPS is a hard requirement once you open more than one chat window — see "Why HTTPS is required" below. The blessed path is Tailscale, which hands out free Let's Encrypt certs for your tailnet's hostnames.
https://<your-host>.<your-tailnet>.ts.net:18737/web/For full Tailscale + cert setup steps see setup.md. Short version:
# in the Tailscale admin: enable MagicDNS + HTTPS Certificates (one-time)
mkdir -p ~/.somora/certs && cd ~/.somora/certs
tailscale cert <your-host>.<your-tailnet>.ts.net # use your own FQDNAdd to ~/.somora/config.yaml:
server:
port: 18737
tls:
cert: ~/.somora/certs/<your-host>.<your-tailnet>.ts.net.crt
key: ~/.somora/certs/<your-host>.<your-tailnet>.ts.net.key
publicHost: <your-host>.<your-tailnet>.ts.netThen systemctl --user restart somora (or however you launch it).
By default the server binds 127.0.0.1. To reach it across the tailnet/LAN, set SOMORA_HOST=0.0.0.0:
# ~/.config/systemd/user/somora.env
SOMORA_HOST=0.0.0.0There is no auth — same trust model as the API server. Tailnet-only by design (everyone with a Tailscale node on your tailnet can reach it, no public exposure).
For development:
cd web
npm install
npm run dev
# vite reads ~/.somora/certs automatically and serves
# https://<host>.<tailnet>.ts.net:5173/web/ (HTTP/2)
# proxies /agents /attachments /browser /chat /dream /files /health
# /models /terminal /tmux /tools /tui-config /version to the https
# somora serverIf ~/.somora/certs/<host>.{crt,key} aren't present (SOMORA_TLS_HOST overrides the host name), Vite falls back to plain HTTP/1.1 — works for one-window debugging, hits the 6-connection wall fast otherwise.
Why HTTPS is required
Browsers cap HTTP/1.1 at 6 concurrent connections per origin. Each chat window holds one persistent SSE stream — so 6 windows max before new tabs silently fail to send and agents look unresponsive. Tmux session attaches and any xterm.js/WebSocket panels eat connections from the same pool.
HTTP/2-over-TLS multiplexes every stream over one TCP connection. The 6-limit becomes effectively unlimited, end of problem.
Plus the secure-context features the client relies on — getUserMedia (mic), getDisplayMedia (screenshare), Clipboard async, Service Workers (offline + push notifications), Web Push: all require HTTPS. You can serve a chat app over plain HTTP, but you can't have voice or push without it.
Window manager
- Desktop icons: one tile per agent from
/agentsplus the app tiles, laid out on a grid that spans the whole desktop. Fresh install = the classic left column; from there drag any icon to any cell (dropping on an occupied cell swaps the two), or move the focused icon withAlt+Arrow. The arrangement is stored per browser inlocalStorage(somora-desktop-icons); a narrower window relocates icons that no longer fit, widening restores them. Icons sit below windows like on a real desktop. Click an agent tile to open its chat window; clicking again focuses the existing window. Right-click an agent tile for its menu: Open main, the three most recently active other sessions (click one to open it in its own window), New session… (type a name — letters, digits,-,_; the field tells you what's wrong while you type — Enter creates it and opens a new window, Esc cancels), All sessions… (the Sessions tool) and Configure… (the Agent window, see "Features" below). Agent-wide actions live here; per-chat settings stay in the chat's•••menu. Each agent tile carries up to three live signals:- Status dot (bottom-right of the icon): green = idle, amber = streaming, violet = holds the dream review loop, grey = offline.
- Pulse glow around the icon when a dream phase is running for this agent: green = REM (per-agent session→memory extraction), indigo = Deep (server-wide consolidation, every participating agent pulses), violet = Lucid (review-loop holder). Polled every 30 s from
GET /dream-states. See dream-phases.md for what each phase does. - REM badge (top-right of the icon) when the agent has REM extractions waiting for review: a small green counter (
1,2,9+). The badge clears as you work through findings viadream_apply/dream_dismiss.
- App tiles: non-agent surfaces —
tmux(attach to an existing tmux session),terminal(fresh shell in the somora workspace),sessions(cross-agent session browser — see next section),sentinel(the proactive triggers: list, detail with history, test / pause / resume / delete — see sentinel.md),abilities(per-agent visibility matrix for tools and skills plus external MCP server health — see mcp.md and skills.md),team(the org chart editor forteam.yaml— see team.md), andlog(the server's own log, see below). Opt-in tiles appear with their feature:voice(a call with an agent, see voice.md),browser(see below),media(the image-generation form and gallery, see imagegen.md) andwiki. Thewikitile carries a violet Lucid badge when completed lucid runs are waiting for review — lucid is platform-wide wiki cleanup, so its review backlog lives here rather than on any single agent. The badge counts pending findings; the tooltip names the oldest waiting run. Review with any agent viadream_review. - Window: drag the title bar to move, drag the bottom-right corner to resize. Close button removes the window without unsubscribing other clients.
- Windows never leave the desktop. Dragging and resizing stop at the edges (a window may hang off the right edge while you drag it, but its title bar and body never go under the taskbar). When the browser gets smaller — you move it from an external display to the laptop screen, or a saved layout comes back on a smaller screen — every window that no longer fits is pushed back inside and, only if it is bigger than the desktop itself, shrunk to fit. Nothing is rearranged and nothing grows back when the browser gets bigger again: use Save/Restore layout in the taskbar for that.
One chat window per conversation: opening a session that is already on the desktop — from the dock, the sessions tool, a link in another chat, or /session typed into a window — brings that window forward instead of adding a second one. The rule sees through the two spellings of a session (its slug and its dated id), so /session main and the sessions list never end up on separate windows of the same stream.
Taskbar gear: reload config, restart
The gear left of Arrange opens a small server menu:
- Reload config re-reads
~/.somora/config.yaml, validates it and swaps it in without a restart. A typo leaves the running config untouched; the toast shows the schema error with field and message. On success the toast lists the changed sections and, when one of them only applies at boot (server, memory, mcp, voice, wiki schedulers, …), says so. The menu marks changed on disk when the file is newer than what the server loaded. - Restart somora asks systemd to restart the user unit. Every open stream drops for a few seconds; the page polls
/healthand reloads itself once the new process answers. Greyed out when somora is not running assomora.service.
agent.yaml needs neither: it is read on every turn. The TUI has the same two actions as /reload and /restart YES.
- Taskbar (bottom): lists open windows and always stays on top — no window can cover it, so Arrange (tile all windows across the desktop) is always reachable. Save/restore persists positions in
localStorage. - Arrange tiles every non-minimized window over the full desktop. Counts that don't fill a grid get a full-height master on the left with the rest stacked beside it (3 → one left, two right; likewise 5 and 7) instead of a grid with a hole in it; 1, 2, 4, 6 … tile as an even grid. Windows keep their left-to-right, top-to-bottom order, so the leftmost window becomes the master and arranging twice changes nothing. Icons are not worked around — they sit below windows, so Arrange uses the width right up to the left edge and a covered icon is back when you close or minimize the window. Chat windows whose agent no longer exists (a renamed or deleted agent in a restored layout) are dropped once the server has answered, rather than sitting there invisible and taking up a tile.
Layout state is per-browser-profile. There's no server-side window manager — each device remembers its own arrangement.
Sessions tool
A cross-agent session browser launched from the sessions tile in the app dock. The single window where you keep order across the hundreds of sessions somora accumulates over weeks.
┌───────────────────────────── Sessions ──────────────────────────────┐
│ 172 total · 1 live · 28 archived · 94 dreamed · 50 partial [⟳] │
│ [Active] [Archived] [All] │
│ search: ____ agent: nova luna engine: codex-cli REM: partial │
│ ─────────────────────────────────────────────────────────────────── │
│ ☐ Agent Slug Engine Status Last act. Msgs Size │
│ ☐ nova main ★ codex-cli ●★ 5 min ago 46 280k │
│ ☐ luna debug-auth-x claude-cli 3h ago 12 22k │
│ ☑ nova sub-self-477… openai-c. yesterday 2 35k │
│ ... │
│ [2 selected] [Archive selected] [Clear] │
└─────────────────────────────────────────────────────────────────────┘What it shows per row:
| Column | Meaning |
|---|---|
| Agent | Persona name, coloured with the agent's color from AGENTS.md frontmatter |
| Slug | Session slug — ★ marks the magic main session |
| Engine | claude-cli / codex-cli / openai-compatible (last engine that touched it) |
| Status | ● (green) = at least one SSE subscriber is live on this session right now · 📦 = archived · ★ = main |
| Last activity | Human-readable relative time (5 min ago, yesterday, 3w ago) |
| Msgs | user_message + assistant_message count |
| Size | Bytes on disk for the JSONL |
| REM | 🧠 ✓ (green) = REM dreamed up to the latest event · 🧠 ⚠N (orange) = REM ran once but N new events have arrived since · 🧠 ○ (grey) = never dreamed |
About the REM column. REM is the per-agent background phase that turns finished session content into memory candidates (see dream-phases.md). The column tells you, per session, how much of its history REM has already processed:
🧠 ✓(dreamed). REM has worked through every event in this session. Nothing waiting.🧠 ⚠N(partial). REM ran at some point — but you've chatted further since. The number is how many user/assistant events have piled up beyond REM's last read-through marker. They'll be picked up on the next idle- triggered REM run for this agent.🧠 ○(never). No REM run has touched this session yet. Common for brand-new sessions, sub-agent spawns that finished quickly, or sessions on agents where REM is disabled inagent.yaml.
The lag count is informational — you don't have to do anything about it. If you want to nudge a session, end it (/reset) and REM will fire on the archived copy at the next idle window.
Tabs:
- Active — non-archived sessions (default). Same set the
/sessionslash-popup sees. - Archived — sessions explicitly archived OR legacy
/resetoutputs (ids ending in-archive). - All — both.
Filters + sort: agent (multi-select chip group), engine, REM-state (dreamed / partial / never), plus a free-text search across slug / agent / id. Column headers sort by last-activity / messages / size / agent (click to toggle direction).
Click a row (outside the checkbox / action buttons) → opens a chat window for that (agent, session). Same wm.openChat() path the agent dock uses.
Archive / Unarchive: the right-edge action button on each row archives (or unarchives in the Archived tab). Bulk-select with the checkboxes and the bulk-action bar archives multiple at once. The magic main session can't be archived directly — use /reset to spawn an archived copy.
Export: two download icons sit next to the archive button on every row — a file-text icon downloads a Markdown transcript (readable, with ## user/assistant headers, fenced code blocks for tool calls, and engine plan items as task lists), a file-json icon downloads the raw JSONL (byte-identical to the on-disk source, full fidelity). Markdown is great for sharing or saving into Obsidian; JSONL is the canonical backup you'd drop on another somora host. Both work for archived sessions too. Backend route: GET /agents/:agent/sessions/:session/export?format=… (see api.md).
Reload: manual reload icon top-right, plus a 60-second auto-refresh toggle (default on). Stats are cached in each session's <id>.meta.json and invalidated by JSONL mtime, so reloads stay cheap.
Archive semantics: archive is meta-flag based, no file movement. meta.archived = true (with archivedAt + optional archiveReason) is the source of truth. The <id>.jsonl and <id>.meta.json files stay where they are. Default-filtering at listSessions() keeps archives out of the slash-popup, chat-window session picker, and the Active tab — they only surface in the Sessions tool. No hard-delete option, on purpose: archive is fully reversible, and you can always clean up ~/.somora/agents/<agent>/sessions/ by hand if you really want bytes gone.
Why this exists: sessions accumulate fast (REM idle-trigger, sub-agent spawns, /reset archives, debugging sessions). Without a place to see everything at once, slash-popups grow until they're useless and you can't tell which old sessions are still worth keeping. The Sessions tool is the housekeeping surface — search, filter, archive, see at a glance which sessions have unconsolidated memory waiting for REM.
Chat window anatomy
┌──────────────────────────────────────────────────────┐
│ 🧠 scribe · assistant ● streaming │ ← header
│ main · opus · think:medium · 🔧 on · ▣21% Σ↑12k ↓4k · ●│ ← live meta
├──────────────────────────────────────────────────────┤
│ [user] summarize today's notes │
│ │
│ [tool] memory_search "today notes" │ ← above
│ [tool] → notes/2026-05-10 · 0.71 │
│ [scribe] You spent the morning on the cache fix... │ ← below
│ ▮ │
├──────────────────────────────────────────────────────┤
│ 📎 Type a message… ▷ │ ← input
└──────────────────────────────────────────────────────┘- Header: agent name, role badge from
AGENTS.md, streaming pill. - Meta line (10px mono): session id, model, thinking level, tools toggle, context fill, token counts, connection dot. The two readings mean different things.
▣is how full the window was on the turn's last request, amber past 75 % and red past 90 %.Σ↑and↓are what the turn spent; the Σ says it is a sum over every request the turn made. A turn with 21 tool rounds sends its context 21 times, so the sum runs far past the window and says nothing about how full it is — a real one read▣ 62% Σ↑ 6.0Magainst a 524k window. The TUI header shows the same pair. On a CLI engine the window is the one the engine reports for itself where it reports one (codex sendsmodelContextWindowper thread), because the configuredcontextWindowis somora's guess at a cap that engine enforces on its own. If the last request was still bigger than the window somora knows, the badge reads▣ >100%with the raw number rather than a percentage that keeps climbing — a plain "300 %" said nothing a reader could act on, and the tooltip names the cause (configured window too small). When the last turn was answered by the persona'sfallback:model, a warn-coloured⇄ <backup-model>marker sits next to the model (tooltip: why the primary failed). - Chat text zoom (⊖ / ⊕ in the header, left of the project chip): 75–200 % in discrete steps, per agent — one conversation can be enlarged without touching other agents, the window chrome or the desktop. The percentage readout appears only off-default and doubles as the reset. Persisted per browser (
somora-chat-zoom). - Fallback marker on bubbles: an assistant turn produced by a fallback model carries a
⇄ fallback · <model>chip. Hover shows why the primary failed — and with a fallback chain (fallback: [a, b]in agent.yaml, see agents.md) every model that failed before this one, in order. It survives reloads — the server persists amodel_fallbackevent in the session history. A short notice appears once at the start of a fallback streak and once more when the primary model answers again; turns in between get the chip only. When every model in the chain fails before producing anything, the turn's error row names all of them with their reasons (All 3 models failed. …) instead of only the last one's raw error. The TUI prints the same as a warn line in the scrollback, the mobile client shows the chip on the bubble. - Server log (the log tile): the end of somora's own log in a window, so a look at what the server did no longer needs an ssh session. Pick the day, a minimum level (debug, info, warn, error) and a text filter; new lines follow every two seconds while the window is open, and following pauses when you scroll up to read. Only the tail of one day's file is ever read.
- Browser window (when
browser.enabled): the browser tile lists browser sessions, one row per agent window — agents sharing a profile run in one Chromium but get a window each, with their own tabs, control state and handoff. A row opens the live view with tab bar, URL bar and take-over buttons. Stopped browsers are not listed, except one still holding an unanswered handoff. A pending handoff adds an Open browser notice in its source chat and a taskbar marker. These and the list update from one change stream. Broken viewer connections reconnect; stopped browsers can be explicitly reopened without replaying old actions. Closing the viewer keeps human control in place. See browser.md. - Peer origin caption: a message another agent sent via
agent_askrenders with that agent's icon and colour; when it was sent from one of the sender's non-main sessions,<agent> · <session>sits left of the timestamp. Main-session traffic shows no caption. The TUI shows the same as↬ <agent>/<session>. - Body: pinned-to-bottom auto-scroll. Manually scroll up to read history; new messages won't yank you down. Scroll back to the bottom to re-pin.
- Tool-rendering toggle (wrench in the meta line): show or hide tool-call / tool-result blocks AND
engine_metarows (codex's internal plan/todo state — see setup.md). Persists peragent::sessioninlocalStorage— a reload restores your choice. Tools render above the agent's answer for a given turn (TUI-style ordering) — the assistant bubble is a single message that re-anchors to the bottom of the transcript whenever a tool event lands, so cumulative text never duplicates around tool boundaries. Engine-meta blocks render dimmer than tool calls and prefix with◌ codex · planso it's visually clear they came from the engine, not from somora's tool layer. - Input: auto-grow textarea up to 120px, then internal scroll. Enter sends, Shift+Enter newline. The textarea stays editable while a turn is streaming — pressing Send during a running turn enqueues the message rather than blocking it (see "Queueing & Stop" below).
- Steer / queue toggle (next to Send, while a turn runs): a small labelled pill reading queue (hourglass, off) or steer (bolt, on). With steer on, what you type goes into the running turn instead of behind it — the model reads it at its next step and changes course; the bubble shows
steering…until the engine has taken it andsteeredafter. With queue the message waits as before. The agent'ssteering:in agent.yaml sets which way the toggle starts; a steer the turn ends before reading becomes an ordinary queued turn. Works on every engine; a Claude turn keeps its channel to somora open until every steered message has been answered, so tools keep working in the steered part of the turn; see api.md → Steering and agents.md. - Stop buttons (two, same abort): while a turn is in flight a red Stop appears in the composer next to Send and on the streaming assistant bubble (in the slot where copy/pin sit on finished bubbles). Send stays live the whole time, so queued sends keep working from the button — Stop is additive, it never replaces Send. One click aborts the current turn server-side via
POST /chat/abort; if there was nothing to stop (turn just finished) or the abort fails, a notice says so instead of silently doing nothing. - Mic button (next to send, when STT is configured): click-to-toggle voice input. Click once to start recording — the icon switches to a red stop-square and
MediaRecordercaptures from the default mic. Click again to stop; the audio blob is POSTed to/stt/transcribe, somora forwards it to the configured upstream (e.g. oMLX Whisper), and the returned transcript is appended to the textarea — your existing draft is preserved, voice gets added with a separating space. Hidden when the browser lacksMediaRecorder/getUserMedia(capability gate identical to the screenshot button), or when somora's/stt/configreports STT disabled. Configuration in setup.md §7. - Voice replies (when TTS is configured): the chat header shows a
🔊/🔇toggle. When on, mic-submitted turns trigger automatic TTS playback of the assistant reply (text replies are unaffected). Toggle is sticky per<agent>:<session>. A Play-button appears on bubbles that have generated audio so you can replay any time. Full details in voice.md.
Markdown is rendered via react-markdown + remark-gfm + rehype-highlight. Code blocks scroll horizontally inside the 75%-width bubble; tables overflow-scroll independently.
Session actions menu (•••)
The three-dots button in the chat header opens a popover anchored to that button — same screen position every time, escapes the chat window via a portal so it can't be clipped. Three sections:
- MODEL — current model + engine + context window. "Switch model…" expands an inline picker with a free-text filter; clicking a model commits a per-session override (
PUT /agents/<a>/sessions/ <s>/model). Same flow as the/modelslash command. - THINKING — current effective level and its source (session override / persona default / engine default). A segmented control (off / low / medium / high) sets a session override; "Reset to default" removes the override and falls back to persona / engine default.
- DANGER ZONE — Reset session — archives the current chat (the jsonl is kept as
<id>-archive) and starts a fresh session. If the agent has REM enabled inagent.yaml, REM fires asynchronously over the archived range to extract memory candidates. Two-click confirm so it can't trigger by accident.
Close the menu by clicking the ••• button again, clicking anywhere outside the popover, or pressing Escape.
Bubble actions: copy and pin
Hover any finished assistant bubble and two small icons appear in the top-right corner. They show up only once the message has finished streaming — partial content can't be copied or pinned. The buttons sit inside the bubble so they travel with their message as you scroll and don't clutter the chat header.
- Copy writes the bubble's raw markdown to the clipboard. The icon swaps to a check glyph for ~1.5 s as a confirmation cue, then reverts. Pasting elsewhere preserves headings, lists, code blocks, tables — anything markdown carries.
- Pin opens a free-floating pin-note window with a snapshot of that message (see below). The pin icon turns yellow + filled and stays visible even without hover, so pinned messages stand out while you scroll the history.
Clicking an active pin again closes its pin-note window. Closing the pin-note window via its × does the same thing — the two affordances stay in sync.
Pin-note windows
A pin-note is a small free-floating window that captures one assistant message for working memory. Use it when an answer is important to refer back to while the conversation continues — a recipe, a config snippet, a decision summary.
Layout:
┌──────────────────────────────────────┐
│ 📌 nova note × ─ ⤢ │ ← yellow titlebar tint
├──────────────────────────────────────┤
│ 🤖 nova · main 14:08 │ ← agent + source + when said
├──────────────────────────────────────┤
│ │
│ rendered markdown body, scrollable, │ ← same renderer as the chat
│ selectable… │
│ │
├──────────────────────────────────────┤
│ 📌 pinned 3 min ago │ ← when the pin itself was made
└──────────────────────────────────────┘Behaviour:
- Free-floating. Drag the titlebar to move, drag the bottom-right corner to resize, same as any other window. Pin-notes have their own z-stack so you can leave one over a chat while continuing the conversation.
- Yellow titlebar tint distinguishes them at a glance from chat, tmux, and sessions windows.
- Snapshot semantics. The content is frozen at pin time. If the agent continues streaming additions to the same turn, the pin doesn't update — re-pin to capture the new state.
- Multiple in parallel. Pin as many messages as you like; each gets its own window. Re-pinning the same message focuses the existing note instead of duplicating.
- Survives reloads. Pin-notes live in the window-manager's
localStoragelayout, so a browser refresh restores them. - Survives the source. Even if the original session is archived or the chat window is closed, the pin-note keeps its content. The header still shows the source agent and session label so you can retrace where the message came from.
Closing a pin-note (× button on the window, or click the active pin button on the source bubble) just removes the note window — the original message stays in chat history.
Queueing & Stop
You don't have to wait for a turn to finish before typing the next one. Submits during a running turn flow into the per-session queue on the server (see api.md) and execute in order once the lock frees. The alternative is the steer toggle in the composer: a message steered into the running turn is read by the model mid-work instead of after it (see "Chat window anatomy").
The optimistic user-bubble shows up immediately with a small hourglass marker next to its timestamp:
┌──────────────────────────────────────┐
│ next thing I want to ask │
│ ⌛ queued · 14:08 │
└──────────────────────────────────────┘When other waiters sit in front, the marker reads ⌛ queued · N ahead. The marker disappears as soon as the server starts the turn — at that point the bubble looks like any other finished user message, and the assistant's reply streams in below it.
While the marker is showing, the message is still yours: the small ↩ edit next to it takes the message back out of the queue and into the composer (DELETE /chat/queue/:id), attachments included, so you can add the thing you forgot and send again. The re-sent message joins the end of the queue. If the turn started in the meantime the bubble just loses its marker and a notice says so — Stop is the handle from then on.
The queue badge
The window header says what the session is doing as a whole: waiting 3 · running · 1 arriving, and 2 sub-agents / 1 ask when the session has work out elsewhere. It counts every kind of turn, not only yours — a question from another agent, a sub-agent brief, a sentinel fire, a question from a call. Click it for the list:
- Running — the turn holding the lock: where it came from, its first line, since when, and Stop.
- Waiting — the queue in order, each entry with its origin, a preview and how long it has waited, and × to remove it, whoever queued it. Your own message comes back into the composer as with ↩ edit; another agent's question is reported to that agent as failed with the reason, a sub-agent brief as cancelled, a sentinel fire as skipped.
- Arriving — answers on their way back into this session: a late reply to an
agent_ask, a finished background sub-agent, a rendered video. - From here — the sub-agents and
agent_askcalls this session started that are still queued or running; a row opens the target session. × on one that has not started yet removes it from the target's queue; ■ on a running one stops it — a sub-agent together with everything it started, a question as the turn it runs on the target. This session hears about it the way it would have heard the result.
Badge and list both read GET /agents/:agent/sessions/:session/work (api.md): refetched on every queue event, and every 3 s while the list is open so the wait times keep counting.
Aborting (Stop button on the streaming bubble) cancels the currently-running turn only. Queued waiters keep their slots and run in order. It does not matter what started that turn: a message from another agent, a sub-agent brief, a sentinel fire, a tmux or browser wake, a voice consult or a wake-up stops the same way as your own message, and whoever asked for it is told it was stopped by the user.
Failed turns
A turn that ends in an error (engine 5xx, watchdog, abort) renders as a red Turn failed block inside the turn, right where it happened. Media the turn produced before failing (a generated picture, say) hangs under that block — not under the previous answer. The block comes from the turn_error SSE event live and from the session file's error rows on reload.
Activity feed (multi-agent dots + unread)
The agent dock and the Sessions tool both show two passive indicators that come from a single app-wide SSE on /activity/stream:
- Streaming dot — fills on every agent whose any session is mid- turn, not just agents whose chat window you currently have open. When a sentinel job wakes a different agent in the background or a peer-agent message kicks off an A2A reply, that agent's dock tile goes busy even if its window is closed.
- Unread dot — appears (different colour from streaming) when a session has activity since you last viewed it. Counts: A2A inbounds, sentinel-triggered messages, and assistant final replies. Self-typed messages and tool/memory side-effects do not count.
A session is "viewed" when its chat window is open and focused. The client POSTs /sessions/:agent/:session/seen on focus and the server broadcasts the cleared state — open the chat on the TUI or mobile and the badge here disappears too. Per-session badges also show in the Sessions tool's session list.
State persists across server restarts: unreadAt and seenAt live in each session's meta. Endpoint docs: see api.md.
Cross-client echo
When you type a message in a Web window, the somora server echoes a user_message SSE event to every subscriber of that agent::session. The TUI tail will show what you typed in the web, and a second web tab on the same chat will show it too. Self-echoes are deduped against the optimistic local-user message via a pending list, so you never see the same text twice.
SSE event vocabulary
The web client listens for these named events on /chat/stream:
| Event | Payload | Meaning |
|---|---|---|
status | {msg} | Connection state — initial connected, periodic keepalive |
user_message | {text, ts, turnId?, origin?, from_agent?, from_session?, from_system?, agent_ask_call_id?} | A user-typed message landed in the session (any client). turnId pairs the event with an optimistic bubble made by POST /chat/send. from_system marks a system-trigger inbound: sentinel, tmux (attention watcher), subagent (finished task), job (video job), browser (hand-back), voice (question from a realtime call), a2a (late answer to an agent_ask). Each is drawn as a centered divider, not a bubble — in the web client, on mobile and in the TUI alike. origin carries the same information as one structured value (shape in api.md); turns recorded before it existed have only the from_* fields. |
turn_queued | {turnId, ahead, workId?, kind?} | Fired when a send hit a busy lock. ahead ≥ 1 includes the currently-running turn. workId is the entry's id in the session's work queue (equals turnId for a typed turn) and kind its origin. Drives the ⌛ queued marker. Re-emitted for the waiters that move up after a dequeue. |
turn_dequeued | {turnId, workId} | A waiting entry was taken back (↩ edit or × in the queue list, from any client). The bubble for a typed message is dropped; the queue list refreshes. |
turn_started | {turnId} | The engine's own turn id — stamped on the assistant bubble so assistant_media / turn_error pair to this turn. |
turn_error | {turnId?, message, engine} | The turn failed. Rendered as the Turn failed block inside the turn. |
agent | {phase: 'start'|'end', usage?, provider?, model?, fallback?, ...} | Turn boundary. On end, provider/model are the model that ACTUALLY answered; fallback {requested, actual, reason, hops?} is set when that was a fallback model (same shape as model_fallback). |
model_fallback | {requested, actual, reason, hops?} | The primary model failed before producing anything; a fallback model is answering this turn. requested is always the primary, actual the model now answering; hops lists every model that failed so far (chain). One event per hop, precedes the first chat delta of the model that answers. |
chat | {state: 'delta'|'final', text} | Cumulative assistant text (each delta carries the full running text, not just the new chunk) |
thinking | {state: 'delta'|'final', text, truncated?} | The model's reasoning text, cumulative like chat; truncated when the server cut it at its cap. Renders as the 🧠 thinking block above the reply (the "Show thinking in replies" switch in the ••• menu, same as /verbose thinking). See thinking.md. |
tool | {phase: 'call'|'result'|'error', tool, summary?, details?, error?} | Tool invocation lifecycle |
engine_meta | {engine, itemType, label, summary?, payload} | Engine-internal side-channel (e.g. codex todo_list). Renders under the tools toggle. |
memory | {count, topScore, refs, fullText?} | Memory auto-inject for this turn |
project | {from, to, via} | Project focus change — fired by /projekt slash + HTTP routes. MCP-routed agent tool calls don't emit this; clients re-GET /…/project on agent:end instead. Only fires when the projects feature is enabled. |
assistant_audio | {turnId, url, mime, durationMs?, cacheKey} | Server-generated TTS artifact for the matching turn. Pairs by turnId; drives the Play-button on the bubble. |
assistant_media | {turnId, media: [{type, id, prompt, mime, filename, url}]} | Media produced while the turn ran. Pairs by turnId and renders under the bubble. Published by the server after the turn finalizes — the agent doesn't have to attach anything. Each entry's type decides the renderer; an unknown type is skipped rather than guessed at. See imagegen.md. |
The server pubsub key is ${agent}::${session} — multiple agents can share a session id like main without leaking events across windows.
FileView
Agents reference files by absolute path in chat ([report.md](/home/…)), and the web client's Markdown renderer turns those into links that open a FileView window. Nothing is refused for being the wrong type: what the viewer cannot render, it describes.
| The file is | You get |
|---|---|
| Markdown | full render, same plugins as chat |
| Text / code | monospace, syntax highlighting for the known extensions |
| Image | inline, click for full size |
| Video | <video controls>, seeking served by Range requests |
| Audio | <audio controls> |
| the browser's own viewer | |
| anything else | name, type, size — and a download |
A download button is present for every file type, including the ones that render inline.
The classification comes from the file's magic bytes, not its extension: a .dat holding PNG bytes is shown as the image it is, and a .png that is really a JPEG is served with the type it really has. .svg is the deliberate exception — it is markup that can carry script, so it is shown as its own source rather than rendered.
Bytes come from GET /files/raw, which streams and honours Range; without that, scrubbing a video would re-fetch the whole file. Inline display is limited to image/video/audio/PDF, because anything else shown inline would run on somora's own origin.
The boundary is the read policy, not the workspace: a file under /tmp/ is viewable, a file under a blocked root is not — the same answer file_read gives an agent. See api.md.
Features
Chat windows — one per
(agent, session), live SSE streaming, full history hydration on open, per-window tool/memory toggles, cross-client echo (multiple windows on the same session stay in sync), window manager with persistence (positions/sizes survive reload).Windows never leave the desktop — shrinking the browser pushes every window back inside and shrinks only what no longer fits; the taskbar (with Arrange, Save/Restore layout) always stays on top. See the window-manager section above.
Agent context menu — right-click an agent tile: open main, jump to one of its recent sessions, start a named new session in its own window, open the Sessions tool, or Configure… (the Agent window below).
Agent window — what a persona is made of and what it costs. A budget strip on top: the three persona files together, the team block, the full assembled prompt and the tool schemas, each in characters and estimated tokens against
promptBudgetsin config.yaml (soft caps: over budget turns the counter yellow, nothing is truncated). Tabs forAGENTS.md,SOUL.mdandUSER.mdas plain editors with a per-file counter and Save/Discard,agent.yamlread-only (model, fallback, REM — edit the file or use the session controls), and Full prompt: the system prompt exactly as the next turn on a chosen session would send it, split into its parts with sizes, plus the tool-schema total and a note on what is not in that text (memory recall, history, the engine's own instructions). Agents that can be called by voice get one more tab, Voice prompt: the whole instruction the talking model is given, whether it came from a hand-writtenVOICE.mdor was derived from the persona, with voice, language and consult policy beside it. Saving is guarded: the agents edit these files themselves, so a save is refused when the file changed on disk since you loaded it, and you are asked to reload; the previous version is kept as a backup.AGENTS.mdmust keep a loadable frontmatter whosenamematches the agent directory.Queued messages are editable — a message waiting behind a running turn shows ⌛ and an edit link that takes it back into the composer; a turn that fails renders as a Turn failed block inside the turn, with any media it produced under it.
Abilities window — per-agent matrix for tools and skills, plus external MCP server health. See mcp.md and skills.md. The matrix follows the agent's kind: a chat agent sees the whole programme minus the three builder-only tools, a builder sees its coding set on top and every other tool under more, off unless switched on. A toggle takes effect on the agent's next turn on every engine; on codex-cli the Codex thread is restarted with the session history carried over (a
tools changedmarker appears in the chat), because Codex keeps a thread's tool set for its lifetime.Builder windows — an agent of kind
builder(builder.md) has a grey outline and a builder label on its tile, and its chat window opens wider than a chat window (960 × 620, so the panel is not folded from the start) and carries the task panel docked on the right: mode (attended / unattended) and phase (plan / build) switches, the Go button that approves the plan and starts the build, the plan file's path, the task list the builder keeps withtodo_write, and any question it asks withask_user, answered right there. The panel folds to a strip when the window is narrower than about 640 px and unfolds again with the window. The agent window shows onlyAGENTS.mdfor a builder (noSOUL.md/USER.md: the harness rules replace the persona), and the Full prompt tab shows the harness prompt with its parts.Team window — the org chart editor for
~/.somora/team.yaml(team.md): drag an agent card onto its new superior (or onto you), edit title, "involve for" / "not for" chips, notes and the active switch per agent, your own name/title/about and the global rules; the preview pane shows the exact# Your teamblock any agent would get — rendered from the unsaved draft, so you see the effect before saving. Save validates on the server, writes atomically and keeps the last five versions as backups; agents pick the change up on their next turn. With no file yet the window offers to create one from the agents on disk.Drag & drop / paste / paperclip attachments. Per-turn user-attachments end-to-end through all three engines: claude-cli inlines as native ImageBlock / DocumentBlock; codex-cli sends images as native turn inputs and rasterises PDFs to per-page PNGs; openai-compatible builds an array-content user message (
image_urlfor images,fileor rasterised pages for PDFs depending onpdfMode). Bytes live content-addressed at~/.somora/attachments/<sha256>.<ext>; JSONL refs only, never inline. Seedocs/files.md(and config block below) for the Server side; the client surface is paperclip + Cmd/Ctrl+V paste- drag&drop overlay + a per-bubble thumbnail row. Capability gate refuses uploads for models without
image/pdfcapability with a clear nudge to/model. Switching an existing session to a text-only model afterwards keeps working: the history is packed for the model that will read it, so past attachments it cannot process are replayed as a text marker naming the file instead of being re-sent and rejected. Seefiles.md.
- drag&drop overlay + a per-bubble thumbnail row. Capability gate refuses uploads for models without
Older-messages lazy-load. Initial history load is paginated to the last 100 events; scrolling near the top auto-fetches the next 100, anchoring the visible position so you stay where you were reading. There's also an explicit "↑ load older" button for clarity. A window left open on a busy session does not grow without bound either: past 400 rows it keeps the newest 300 once the running turn has ended, and the older rows come back through the same "load older" path. Each window redraws only for events of its own session, so several streaming windows side by side stay responsive.
Memory-inject banner — per-turn
🧠 memory · N hits · refs…line in the chat flow, mirrors the TUI's◇ memory · …row. Brain icon in the meta-line toggles visibility; chevron expands the full injected text. Persists peragent::sessionlike the tools toggle.Slash-command popup — type
/in the input to bring up a command picker:/model <ref>— switch model for this session (autocompletes from/models)./session <slug>— switch the window to another session of this agent (in-place, the SSE re-subscribes)./new <slug>— create a new session and switch this window to it./thinking <off|low|medium|high|default>— set / clear the thinking-effort override./sampling [key=value …|default]and/temp <n>|default— sampling parameters for this session (openai-compatible engine), see sampling.md./verbose thinking on|off— show or hide the 🧠 thinking block above replies in this session (display only; same switch as the checkbox in the ••• session menu)./projekt <slug>(alias/project) — pin a project to this session. Autocompletes from/projectswith entity + path-count detail per row./projekt unlinkis always the first row so clearing the pin doesn't require waiting for the list to load. Only present when the projects feature is enabled inconfig.yaml./reset YES— archive this session and start fresh (same as the Danger-zone action in the•••menu);YESis the confirmation.
Arrow-keys navigate, Enter or Tab accepts, Escape dismisses. Note:
/agentis intentionally not a slash-command — the agent dock on the left is the agent switcher, switching agent inside an existing window would mean discarding the window's identity./skillis also out, because skills are activated by the agent itself via theskilltool, not by user-facing CLI shortcuts.Tmux session windows — attach to a running tmux session in its own window via xterm.js + a WebSocket bridge. Sessions you started in the terminal show up in the AppDock; output streams live and keystrokes go straight through.
Sessions tool — cross-agent session browser in the AppDock with filter (agent, engine, REM state), sort, search, click-to-chat, bulk archive/unarchive, and 60s auto-refresh. Archive is a meta-flag (no file movement, no hard delete) so archived sessions stay inspectable and can be restored. When the projects feature is on, the table also has a Project column with color-coded chips per row.
Wiki explorer (opt-in, read-only) — a dock tile that opens the shared wiki in three columns: folder tree, rendered page, link graph plus backlinks. Obsidian
[[wikilinks]]are clickable and navigate in place; targets that match no page render as broken rather than silently disappearing, so gaps in the wiki stay visible. The graph toggles between the current page's neighbourhood and the whole wiki, and clicking a node opens that page. The tile only appears whenwiki.enabledandobsidian.vaultare both set — the same gate the server applies to the routes. Nothing here writes: the wiki is owned by Deep/Lucid, and a viewer that could also edit would race them. See wiki.md.Project chip + switcher (opt-in) — when
projects.enabledis on, the chat-window header gains a chip next to the ••• action button. Color-pill with project name when pinned; ghost folder button when unpinned. Click opens a switcher popover with all configured projects grouped by entity + a search field. Cross- client live update via SSE: pin via/projektin the TUI, the open web window updates within 100 ms. See projects.md for the feature model.
Files of interest
web/
├── src/
│ ├── components/
│ │ ├── ChatProvider.tsx ← global per-session state + SSE multiplexer
│ │ ├── Desktop.tsx ← window manager host
│ │ ├── Window.tsx ← drag/resize chrome
│ │ ├── ChatWindow.tsx ← header + body + input
│ │ ├── MessageItem.tsx ← per-message renderer (user/agent/tool)
│ │ ├── DesktopIcons.tsx ← icon grid: drag/swap/keyboard placement
│ │ ├── AgentTile.tsx ← one agent tile (status dot, pulse, badge)
│ │ ├── AppTile.tsx ← one app tile
│ │ ├── WikiWindow.tsx ← wiki explorer: tree + reader + links
│ │ ├── WikiGraph.tsx ← d3-force layout, plain-SVG rendering
│ │ └── Taskbar.tsx ← bottom bar + layout actions
│ ├── hooks/
│ │ ├── useAgents.ts ← /agents poll
│ │ ├── useSessionInfo.ts ← /tui-config + per-session model/thinking
│ │ ├── useLoopState.ts ← /dream/review-state poll
│ │ └── useWindowManager.ts ← layout state + localStorage persistence
│ ├── lib/
│ │ ├── api.ts ← typed wrappers around /agents, /chat/*, etc.
│ │ └── colors.ts ← per-agent gradient + role-tint resolver
│ └── styles/
│ ├── desktop.css ← desktop chrome (windows, taskbar, tiles)
│ ├── globals.css ← Tailwind + targeted overrides
│ └── tokens.css ← CSS variables (theme)
└── vite.config.ts ← base: '/web/', proxy in dev mode