Skip to content

Tools ​

somora exposes tools to the agent through a single registry. Every tool is engine-agnostic: claude-cli (and grok-cli) see them via an MCP server somora spawns per turn, codex-cli receives them as Codex dynamic tools on the app-server and somora answers the calls, openai-compatible sees them in-process. The model never knows which path it's on.

Tool families ​

ToolsetToolsPurpose
memorymemory_search, memory_get, memory_list, memory_write, memory_edit, memory_deleteRead/write across the three layers — agent memory inbox, shared wiki, read-only vault. Hybrid retrieval (vector + BM25) with per-source boost; filler words are dropped from the keyword side and a page whose slug names a query word is lifted, so memory_search "karl" finds the page about Karl, not the pages that mention him most (see memory.md).
dreamdream_list, dream_get, dream_apply, dream_dismiss, dream_run, dream_reviewInspect findings from REM (per-agent) and Lucid (platform-wide); trigger Deep/Lucid via dream_run({phase: 'deep'|'lucid'}). dream_review({dream_id, action:'start'|'end'}) opens/closes a conversational wiki-edit loop for a Lucid run. See dream-phases.md.
wikiwiki_edit, wiki_create, wiki_delete, wiki_moveLoop-scoped wiki write tools. Only exposed to the agent currently holding the active dream_review loop; otherwise hidden. wiki_edit mutates body and/or related:/sources: frontmatter; wiki_move moves or renames a page into a folder the wiki map lists and rewrites every [[link]] to it. A per-turn cap on wiki_* calls (wiki.lucid.maxCallsPerTurn, default 3) keeps the model from batch-editing without user check-in. See dream-phases.md.
timetime_nowCurrent date/time/timezone — model never hallucinates "today".
webweb_search, web_fetchSearch via Brave + fetch web pages as Markdown (Mozilla Readability + SSRF guards).
filefile_read, file_write, file_patch, file_search, file_list, analyze_fileGeneric filesystem I/O — local or against any configured SSH resource via target=.... analyze_file is the multimodal companion for models that cannot see: it dispatches images/PDFs to the configured vision.worker — a single model or an ordered chain tried until one answers — and returns a text description. It is hidden from any agent whose active model has the image capability (see files.md). file_read returns numbered lines (N: text, 2000 per call, with the offset to continue), file_patch tolerates whitespace and indentation drift between the model's old_string and the file and shows a diff of what changed, file_search takes include globs, context lines and files_only, file_list walks recursively with a glob and honours .gitignore; output that had to be shortened is kept in full under ~/.somora/agents/<agent>/tool-output/ and the result names the file. Limits and the matching chain in files.md. For a builder agent, a write or patch result also carries the language server's errors for the file (lsp.md).
execexec, processOne-shot shell commands (sync + background) on local or any SSH resource. Hard-blacklisted destructive patterns; per-resource allowBlocked opt-in lets dedicated agent-workstations whitelist admin commands like sudo … or systemctl reboot (see resources.md). Background jobs are fully detached (survive somora/MCP restarts until they finish or are killed) and disk-tracked with poll/log/kill via process — poll verifies real process liveness (exit-code file + PID probe), never a stale in-memory state.
exec (tool tmux)tmuxPersistent multi-turn terminal sessions (claude/codex/vim/REPLs). One typed tool with `action: create
agentsspawn_subagent, spawn_subagents, subagent_status, subagent_result, subagent_list, subagent_cancel, agent_ask, agent_ask_result, agent_ask_cancel, session_list, session_model, builder_dispatchAgent-to-agent orchestration. spawn_* create sealed sub-sessions for delegated work; subagent_* poll/collect results — a done result carries the sub's final text plus runtime-derived facts the model cannot fake: outcome (completed / partial — the engine had to force a finish at the round cap or tool budget / degraded — the text is a somora marker, the sub looped or never answered / failed) with outcome_reason, tool_calls and rounds, files_written (files created or edited via file_write/file_patch), and media (absolute paths of every image or video the sub generated, straight from the media store). All of these are present even when the final text degraded, and the [subagent attention] wake names outcome, files and media too; subagent_cancel cancels a spawn tree of your own, running or still waiting in its queue (cascades to child-spawns, disk artifacts untouched). Finished async subs wake their parent with a [subagent attention] turn unless spawned with attention:false or the result was already fetched. Every spawn is such a background task: wait:true waits for it and returns the result inline plus its task_id, so a synchronous sub is listed by subagent_list, stoppable, and reached by subagent_cancel together with its children while the parent waits; when the wait runs out (agentLoop.longTaskMaxTimeoutMs) the tool returns state: "pending" with the task_id, the sub keeps running and the parent is woken when it finishes. Child spawns count against the parent persona's concurrent cap (4), with the last slot reserved for top-level spawns so an orchestrator sub can never lock its own agent out. agent_ask posts a live question into another agent's existing session (request-response: the target's normal reply flows back as the tool result). The target sees [Message from agent <name>, session <slug>], so it can address a follow-up; an agent_ask without session that answers the agent who wrote to it (or a sub's spawning parent, or — from an [agent answer] wake turn — the agent whose answer woke it) goes back to the session that message came from, otherwise to main — an explicit session always wins, and an unknown slug returns the target's existing sessions instead of a bare 404. With create_session: true (and optionally model: "<alias>") a missing project session is created on the target, the model pinned on it and the target told — so "work with <agent-a> and <agent-b> in their <project> sessions, create them if needed, <agent-a> on <alias-1> and <agent-b> on <alias-2>" is two calls; the model is only applied to a session the call creates, never to an existing one. When the target does not answer within timeout_ms, agent_ask returns state: "pending" with a call_id — and when that answer finally lands, the asker is woken with the first lines of it and the call_id to read the rest (the same attention wake a finished sub-agent sends). wait: false hands the message over and returns pending at once, for jobs that may take their time; the wake is the same. agent_ask_result fetches or waits for that call's outcome (done / failed / pending with phase queued or running) — the message is never re-sent, so the target's work cannot run twice. agent_ask_cancel takes a call of your own back: out of the target's queue while it waits (removed), or, while it runs, by telling the target to stop through its running turn and letting the outcome wake nobody (withdrawn, steered says whether the stop message reached the turn) — it never aborts the target's turn. A target that answered "I am working on it with sub-agents and will report back" does not have to remember to: when the work it started while answering finishes later — its subs, calls of its own, and whatever those start — the asker hears the outcome once, as a follow-up: an ordinary message from the target under the original call_id for an agent, one more [subagent attention] wake with the text in follow_ups of subagent_result for a parent, a spoken follow-up for a voice call; so an asker never re-asks. Both agent_ask and the spawn_* tools take images: ["/absolute/path.png"] — pictures the receiving agent SEES (somora uploads them and puts them on that turn), because a path in the message text is just text to a model; a target whose model has no vision gets the vision worker's description instead. Blocking A2A waits are deadlock-guarded — the server tracks who waits on whom and rejects any call that would close a wait cycle (even across chains of 3+ agents) with an instructive error instead of letting both sessions hang. session_model switches the model of an EXISTING session — the agent's own (default: the one it is in) or another agent's, named by agent + session — or clears the override with clear:true; the running turn keeps its model, the next one uses the new one, and the switch is written into that conversation with the agent's name and shown in every open window. (agent_ask with create_session + model remains the way to pin a model on a session that does not exist yet.) session_list is the read-only view of somora's chat sessions — the agent's own by default, another agent's with agent, everyone's with "*": session name (usable as session in agent_ask and sentinel dispatch), last activity, message count, pinned project, whether a turn is running there right now and how many wait behind it, and which row is the session the agent is in. See agents.md. builder_dispatch hands a build order to a builder agent (kind: builder, see builder.md) in one call — fresh session, project pinned, unattended + build — and returns a call_id; the caller is woken with the builder's report like with any agent_ask (wait:false). builder may be omitted when exactly one builder exists; a project folder another builder is working in right now is refused (one builder per folder, see builder.md).
mediamedia_listFind images and video generated earlier, newest first, with an optional type filter. Its own toolset because the gallery is shared: a video-only install still needs to list what it made.
videovideo_generate, video_status, video_modelsJob-based text-to-video. video_generate returns immediately and the agent is woken when the render lands — it never waits. Global concurrency cap; the caller that finds it full is refused with the numbers rather than queued. reference_images takes paths and the order means something (one = opening frame, two = opening and closing). Hidden unless videoGen is configured. See videogen.md.
imageimage_generate, image_modelsText-to-image against a configured image model, plus an index over everything generated. Specs (aspect_ratio, resolution, quality, n, …) are real request fields, validated against what the selected model actually accepts; the prompt is passed through verbatim. Returns paths and metadata rather than bytes — return_image: true opts into the image entering the agent's context. reference_images takes file paths (image-to-image); several may be combined. image_models lists the configured handles and what each one accepts. Hidden unless imageGen is configured. See imagegen.md.
memory (tag)somora_docs_list, somora_docs_readRead somora's own documentation (this directory). Tagged memory for gating (toolset:memory), grouped with the other read-only discovery tools.
memory (tag)resource_list, resource_testDiscover and probe configured remote SSH targets. Tagged memory for gating, like the docs tools.
mcpmcp__<server>__*Tools imported from external MCP servers configured in mcp.servers. Discovered at runtime, schema-sanitized, namespaced per server, and gateable per agent like any built-in tool.
browser (optional)browserA managed Chromium per agent profile, with one window per agent when a profile is shared: open (with per-tab device/locale emulation), snapshot (accessibility tree with element refs), act (click/fill/press/scroll/select), screenshot, tabs, status, request_handoff (the user signs in or decides in the web client, the agent is woken to continue in the same tab), close_tab, stop. Hidden unless browser.enabled; the browser ability is also checked at execution. Handoffs appear in the source chat and taskbar, belong to the requesting agent's window alone, and require an existing, non-archived source session. See browser.md.
buildertodo_write, ask_user, plan_writeThe builder kind's own tools (see builder.md): the session task list shown in the task panel, a question to the person watching (blocks until answered or timed out; only offered while the session is attended), and the plan file (the one write allowed in the plan phase). Offered to builders by their kind allow-list; a chat agent sees them only when its tools.allow names them.
sentinelsentinelTime-based triggers that wake an agent on a schedule — one action-enum tool: create, list (fired one-shots hidden unless include_completed), get, pause, resume, delete, test, history, purge_completed (drop every fired one-shot at once). See sentinel.md.
projects (optional)entity_list, project_list, project_get, project_create, project_update, project_focusPointer-file manifests that bind a session to a real-world thing. Visible only with projects.enabled: true (registered always, hidden by their availability probe otherwise). See projects.md.
skillsskill, skill_listActivate a Markdown skill ("how to do task X with our tools"), or list all skills available to the agent. Skills live at ~/.somora/skills/<slug>/SKILL.md (agentskills.io format). The system prompt carries a name+description registry; skill loads the full body on demand, skill_list re-fetches the registry fresh from disk (useful when an engine froze the system prompt at session start). See skills.md.

Definition shape ​

A tool is a ToolDefinition (see src/tools/types.ts) with:

  • name — globally unique (also the MCP method / OpenAI function name)
  • toolset — grouping tag
  • description — what the LLM sees in the tool list. Long-form preferred: descriptions are policy, not just API docs ("use this INSTEAD of running cat via exec").
  • inputSchema — Zod schema for runtime validation
  • jsonSchema — JSON Schema for MCP / OpenAI tool definitions
  • handler(input, ctx) — receives validated input + a ToolContext ({ agent, getMemoryManager, config })
  • available?(ctx) — optional runtime probe; tools that fail are hidden from the model entirely (no API key → no web_search exposed, etc.)
  • maxResultSizeChars? — cap on the JSON-stringified result; default 100 000 (≈25–35k tokens)

The registry truncates oversized results to a { truncated, preview, hint, … } marker so a runaway tool can't blow context.

Where tools are wired ​

Two ToolRegistry instances exist by necessity:

  • In-process registry in src/server/index.ts — used by the openai-compatible engine's agent-loop and by HTTP debug endpoints (GET /tools).
  • MCP child-process registry in src/mcp/server.ts — spawned per turn by claude-cli and grok-cli (different process, separate memory). codex-cli uses the in-process registry: the app-server asks somora (item/tool/call) and the tool runs in the server.

Both populate from a single registerAllTools(registry) function in src/tools/index.ts. Adding a new tool bundle = ONE new registerMany() line there; both engine surfaces pick it up. The two registries are physically separate (process boundary) but the code that fills them is shared, so the effective tool set is identical across all three engines.

When a call arrives broken ​

A model streams the arguments of a tool call as text, so a call can arrive unusable in two different ways, and they need different answers:

  • Cut off — the text stops mid-value, usually because the answer ran out of output allowance while writing it. Retrying the same request reproduces it exactly, so somora does not retry: it answers the call with "arrived cut off after N characters … call it again with less in one go", and adds the numbers when the round demonstrably spent its whole allowance.
  • Malformed — complete but not valid JSON, the kind of slip a second attempt usually does not repeat. somora silently sends the same request again, up to twice. Only if the model keeps producing invalid JSON does it get told, with the parser's complaint and the reminder that a tool without required parameters takes {}.

The two are told apart structurally: a complete JSON value ends on its closing bracket. Not by the provider's stop reason — routers rewrite it. vLLM replaces length with tool_calls whenever it parsed a tool call, so the stop reason cannot show an output limit at all.

Neither fault ever travels back to the model. Backends parse tool-call arguments while building the next prompt, so returning the text means the whole request is rejected — and since the call sits in the running conversation, every following round of that turn is rejected too. The call keeps its place in the message (dropping it would leave a reply without its call, which backends refuse just as hard) but its arguments are replaced by an empty object. The unparsed text stays in the session record, where it is evidence rather than a payload.

Background reading ​

  • display.md — /show and /verbose toggles for the TUI
  • thinking.md — cross-engine thinking depth control
  • memory.md — per-agent memory inbox + retrieval mechanics
  • wiki.md — shared long-term wiki layer
  • dream-phases.md — REM/Deep/Lucid background consolidation
  • resources.md — SSH targets that file_* / exec / tmux dispatch to
  • skills.md — Markdown skill system + per-agent visibility (deny/allow, Abilities window)
  • files.md — file_* tools + multimodal analyze_file
  • tmux.md — shell-vs-TUI session patterns + capture/send modes