Realtime voice — talking to an agent
A standing, interruptible conversation with one of your agents: it listens while you speak, answers out loud, and you can talk over it. Nothing like the dictation button — see the boundary below.
Not the same as voice.md. That page is press-to-talk: one recording becomes one message you can edit before sending, and the answer can be read back to you. This is a call. The two features share no config, no endpoints and no clients, and both can be on at once.
How it works
Two models, and the split is the whole point:
you ──audio──► somora ──audio──► realtime model (talks)
│
└── "look this up" ──► your agent (knows)
persona, memory, toolsThe realtime model runs the conversation: rhythm, listening, interrupting, speaking. It has no memory, no files and three or four tools: ask the agent, check on the work, fetch an answer that was handed over, and — when allowed — move the call. Everything factual, and every request to act, it hands to the real agent — which answers in its own session with its own model and its own tools. The spoken answer is that answer, shortened for the ear.
That separation is why the voice can be quick without inventing things: the part that talks does not know anything, and the part that knows does not have to be fast at talking.
How it talks
The voice speaks the language from voice.language on the agent, else stt.language, else English — named in full in its instructions, not as a language code, because a two-letter code buried in an English paragraph is a weak signal — the voice would follow the paragraph.
It does not introduce itself. The caller picked this agent in a picker and talks to it daily, so a recital of name and role is a wall in front of the first question. Handed a call, it says one short sentence that it is there and carries on.
Reading what the voice is told
The voice self lives nowhere on disk unless you write a VOICE.md, so the web client shows it: right-click an agent, and the agent window has a Voice prompt tab next to the persona files. It shows the whole instruction the talking model gets, whether it came from VOICE.md or was derived from the persona, plus voice, language and consult policy. The tab appears only for agents that can actually be called.
The same thing over HTTP: GET /voice/instructions?agent=…&session=….
What the agent sees
The question arrives in the bound session as a normal turn with from_system: 'voice' — not as a message from another agent. The agent answers into its own chat and addresses nobody back; the call reads that answer and speaks it. Both clients render the question as its own block, so a reader can tell it apart from something you typed.
The question carries the last few lines of the call with it. Asked "how long does that take", an agent otherwise has no idea what you were talking about, and the session would read like a riddle a week later.
A conversation cannot wait on a build. The question queues on the session like every other turn, and the call waits realtimeVoice.consult.quickAnswerMs (default 8 seconds) in total — for the session lock and for the answer together. An answer inside that window is spoken at once — not read verbatim: it reaches the voice with an English instruction that names the language of the call and the persona's maxSpokenSentences budget (default 4) and asks for the substance, the names, facts and numbers as the agent gave them, without reading lists or paths aloud and without shrinking it to "it is done". Past it, the voice tells you it handed the request over and keeps talking; nothing the agent was doing is cancelled, and the question stays in the queue or keeps running. The answer then reaches you one of two ways: the call reads it out on its own at the next pause in the conversation — once, only in the call it belongs to, and only while that call still talks to the same agent; if the voice happens to be mid-sentence at that moment, the reading is kept and spoken as soon as it falls silent, never dropped. That reading is an instruction too, never a canned sentence: say in the call's language that this is the answer to the earlier question, then give the answer as it is, nothing added — or the voice fetches it when you ask whether it is done (somora_consult_result). Asked how far along it is, the voice reports the running work and each handed-over request with its place in the queue (somora_work_status). A request that was stopped or removed on the way is announced the same way, with the reason. Work the agent started for your question and finished only after its answer was read out — a sub-agent, a question to another agent — reaches you the same way, once, opened as a follow-up to that earlier question and then given as it is; a follow-up for a question this call never asked, or for an agent the call has since left, is not read.
The framing itself travels beside the question, not inside it — in the same field somora uses for the memory block, which every engine puts in front of the user message. The model reads both; the session records only what was asked. Stored inside the text, that boilerplate would be half of every voice turn, and it would be read later by the dream phase and by the recall search, which is exactly where it does damage.
That turn is the whole record. A call leaves the same trace in a session that an agent-to-agent request leaves: the question that reached the agent, and the answer it gave. The talking around it — your sentences as you said them, the spoken rendering of the answer — lives for the length of the call and is not written.
The voice self is told to pass on more than requests: a decision, a date, "that project is history" — a remark that changes what the agent should know goes over as a short note, so it lands in the record and the dream phase can learn it.
That is deliberate. A spoken sentence that never became a question is not a turn: engines that resume their own session drop an unanswered user message, the dream phases would learn every question twice, and a line written while a turn is running can break that turn's pair. One question, one answer, one place.
The voice self
Each agent speaks as itself, in the first person. Its character is derived from its own persona files, so nothing is maintained twice, plus a small voice: block for what only speaking needs.
# ~/.somora/agents/<name>/agent.yaml
voice:
enabled: true # may this agent be called at all
voice: ash # provider voice id (see below)
language: de
style: "trocken, direkt, kein Smalltalk"
consultPolicy: always # auto | substantive | always
maxSpokenSentences: 4The voice self carries a short character sketch, not the whole persona. Who is on the line comes from the agent's own USER.md: its opening — everything before the first ## section, up to 280 characters — is carried along, so put the name and how to address the person there. Anything further about them is a lookup like any other fact. No USER.md, no such line. The voice self is also told the day and the time the call started; for the exact time later in a call it looks it up.
For full control, write ~/.somora/agents/<name>/VOICE.md: its text replaces the derived character. The rules that keep the call honest — ask before answering anything factual, never invent, never refuse work on your own authority, speak in the first person — always stay.
Read what a call would actually send:
curl -s "http://127.0.0.1:18737/voice/instructions?agent=<your-agent>" | jq .Configuration
realtimeVoice:
enabled: true
provider: openai # openai | local (same adapter) — google has no adapter
model: gpt-realtime-2.1-mini
apiKeyFile: ~/.somora/secrets/openai-realtime.key # a FILE, not the key
# url: ws://127.0.0.1:8787/realtime # where to connect; omitted = OpenAI
transport: websocket # default; the only transport somora drives
defaultVoice: alloy
consultPolicy: always
maxCallMinutes: 20 # the meter runs while nobody speaks
allowAgentSwitch: true # move a call to another agent or session mid-conversation
consult: # how long a spoken question waits on the agent
quickAnswerMs: 8000 # an answer inside this is spoken at once; past it the
# request is handed over and read out when it lands
turnDetection: # how easily you can interrupt
threshold: 0.4 # lower = reacts to quieter speech
prefixPaddingMs: 200 # how much run-up counts as speech
silenceDurationMs: 420 # pause that ends YOUR turnThe key lives in a file with 600 permissions, never in config.yaml: a realtime key buys billed minutes, and the config is read by more eyes.
Voices on OpenAI: alloy, ash, ballad, coral, echo, sage, shimmer, verse, marin, cedar. Anything else is refused with that list. They are not tied to a language — each speaks German too — but they sound noticeably different, so give agents different ones.
In the web client
The voice tile appears only when realtime voice is configured and at least one agent may be called. Pick agent and session, press talk, allow the microphone.
- A figure driven by the real audio levels, both sides at once: the agent's from the playback in its colour, yours from the microphone.
- The state line says what is happening: listening, looking it up, talking — and counts the lookups, which is the honest measure of whether the voice self is really delegating.
- The clock runs, amber past three quarters of
maxCallMinutes. - The transcript follows the conversation and shows the sentence being spoken right now, dimmed.
The browser stays dumb: microphone in, speaker out, state on screen. It holds no provider, no key, and never sees a tool call — the server owns all of it.
Moving the call
With allowAgentSwitch: true, say where you want to go. Two directions work, and they are the same operation:
- another agent — "put me through to <agent>"
- another session of the agent you are talking to — "go into your projektA session"
Both conversations keep a line saying where the call went and where it came from.
Name a session and you land in it. Name none and you land in main, always — the session the call is currently in is never carried over to another agent: a same-named session of theirs need not even exist.
Session names are matched by how they sound, not by how they are spelled. A transcript that renders a session called projekt-alpha as Projekt Alpha, projektalpha or just alpha still lands there: case, hyphens, spaces and umlauts are folded away, a fragment is enough, and a near miss still counts. Two sessions that sound equally close are a question, not a guess. This applies to calls only. A session name typed into a slash command or an API call still means exactly what it says.
The name in the window changes when the new agent is actually on the line — not when the handover is decided. The previous agent's last sentence can still be in your speakers at that moment, and a name that changes mid-sentence shows the wrong agent talking.
If the move cannot be made, say because the session does not exist, the call comes back with the agent you had and says so. It does not hang up.
Underneath it is a new connection, because a provider voice cannot be changed once a session has produced audio. That is a deliberate second of transition rather than two agents that sound alike.
Pointing it somewhere else
provider names the protocol, not the vendor. openai and local are the same adapter: one talks to OpenAI, the other to a service of your own that speaks the same session and event language. Set url to that service and nothing else changes — the agents, the tools, the switching and the record all work the way they do here.
realtimeVoice:
enabled: true
provider: local
url: ws://127.0.0.1:8787/realtime
model: my-voice-modelThe model id is appended as ?model=…, the way OpenAI expects it; a service that serves one model can ignore it. Without apiKeyFile the socket carries no Authorization header, so a service on your own machine needs no credential. ws:// is allowed for exactly that case; anything reachable from outside should be wss://. OpenAI's own endpoint still refuses to connect without a key.
What a service has to speak is the contract in src/voice/realtime/types.ts: a session that is configured once, audio in and out as PCM16, transcripts as they arrive, tool calls with results handed back, and the ability to cancel a response mid-sentence. A second adapter for a protocol somora already speaks would drift from this one within a week, which is why there is only one.
What it costs, and what it does not do
A call is billed per minute of connection, including silence — hence maxCallMinutes, which is measured from the start of the call and survives every move to another agent or session. One call runs at a time: a second window is refused, and told who is on the line. The agent's own turns are billed as usual, separately. A ChatGPT or Codex subscription does not cover the realtime API; it needs its own key.
The call lives in the web client only, one call at a time, and always bound to a session. Audio always travels browser → somora → provider over the websocket transport, so tool execution has exactly one path. realtimeVoice.transport defaults to websocket, the only transport somora drives; the key accepts webrtc for forward compatibility but nothing uses it.