Models known to run with somora — per engine, with the settings that work
Configuring a model for somora is not "the model", it is "the model behind this engine": the same GPT-5.6 has a different usable context through Codex than through an API, Qwen wants xhigh where somora says high, Kimi refuses sampling parameters, and a value copied from a model card can quietly mis-size the compaction worker. This page collects what has been verified against real somora installs — the recommended config.yaml block per model, and why each value is what it is. It is maintained by hand; the verified column says when a row was last checked. Corrections and new rows are welcome as pull requests.
Field semantics are explained once, at the end of the page, in What each field does, per engine. Read compaction.md for the contextWindow story in depth.
Aliases below are suggestions — pick your own.
What "verified" means. A dated row went through somora's engine matrix on that day, on a live install: basic tools (time_now, exec, file_read, tmux, memory_search, a deferred tool), tools on an SSH resource (exec, file_write / file_read / file_list, tmux on another host), web_fetch, agent_ask to another agent, spawn_subagent, an image attachment (where the model has vision), /thinking high, an abort mid-turn, and session memory across turns. Last full run: 2026-09-05 — claude-fable-5-1, gpt-5.6-terra, gpt-5.5, deepseek-v4-flash, qwen3.8-flash-next, all green (DeepSeek: image attachment refused by design, no vision).
claude-cli — Claude Code subscription
Runs through the Claude Agent SDK on your claude login. No API key, no sampling knobs (the CLI does not expose them), reasoning via the SDK's effort levels.
providers:
anthropic:
engine: claude-cli
models:
- id: claude-fable-5-1
alias: fable
contextWindow: 1000000
capabilities: [text, image, pdf, reasoning]
- id: claude-opus-5-5
alias: opus
contextWindow: 1000000
capabilities: [text, image, pdf, reasoning]
- id: claude-sonnet-5
alias: sonnet
contextWindow: 1000000
capabilities: [text, image, pdf, reasoning]
- id: claude-haiku-4-5
alias: haiku
contextWindow: 200000
capabilities: [text, image, pdf, reasoning]| model | contextWindow | notes | verified |
|---|---|---|---|
claude-fable-5-1 | 1000000 | Frontier model, adaptive thinking always on. The SDK discloses no thinking text for it — somora shows a placeholder row that the model thought (thinking.md). No separate reasoning-token count (rolled into tokens_out). | 2026-09-03 |
claude-opus-5-5 | 1000000 | Opus 5.5 — Anthropic's recommendation for most workloads; replaces claude-opus-5 (keep your alias, change the id). Thinking text: placeholder, as above. | 2026-09-30 |
claude-sonnet-5 | 1000000 | Same surface as Opus, cheaper on the subscription budget. | 2026-09-03 |
claude-haiku-4-5 | 200000 | Does support extended thinking — keep reasoning in capabilities, it is missing from most example configs. | 2026-09-03 |
Peculiarities of this engine
contextWindowis the model's real window here: Claude Code sessions run against it and compact on their own. The value only feeds the compaction-worker choice and the header percentage.- Tools reach the model through
ToolSearch(Claude Code 2.1.142+): somora's tools are discovered on demand, so the first call in a session is preceded by aToolSearchrow. Normal. thinkinglevels map to the SDK'seffort;offdisables thinking.reasoning.levelsis honoured but rarely needed — the vocabulary is alreadylow | medium | high.- The claude.ai connectors (Gmail, Calendar, Drive) never exist inside a somora session (security.md).
codex-cli — ChatGPT subscription
Runs the bundled Codex (@openai/codex, exact version pinned in somora's package.json) as an app-server per turn, on your codex login. somora hands its tools to Codex as dynamic tools — no MCP child, no deferred-namespace guessing — so every model Codex offers (gpt-5.5, the GPT-5.6 family, GPT-6) reaches the same tool set the other engines see. A global codex on the host is not used; somora codex login signs in with the bundled one, and an existing codex login is picked up automatically (setup.md).
providers:
openai:
engine: codex-cli
models:
- id: gpt-6-astra
alias: astra
contextWindow: 258400 # what codex reports as modelContextWindow (measured 2026-09-11)
capabilities: [text, image, pdf, reasoning]
reasoning:
levels: { "off": low, high: xhigh } # Astra has no `minimal` — map `off` explicitly; `max` deliberately not a default
- id: gpt-5.6-sol
alias: gpt56
contextWindow: 258400 # the codex session window, NOT the 1.05M API window
capabilities: [text, image, pdf, reasoning]
reasoning:
levels: { high: xhigh } # optional: /thinking high → codex xhigh
- id: gpt-5.6-terra
alias: terra
contextWindow: 258400
capabilities: [text, image, pdf, reasoning]
- id: gpt-5.6-luna
alias: luna
contextWindow: 258400
capabilities: [text, image, pdf, reasoning]
- id: gpt-5.5
alias: gpt55
contextWindow: 258400
capabilities: [text, image, pdf, reasoning]| model | contextWindow | notes | verified |
|---|---|---|---|
gpt-6-astra | 258400 | GPT-6 — hardest problems; code-mode-only like the 5.6 family. Every codex model here reports the same modelContextWindow of 258,400 (measured 2026-09-11 against codex 0.153.3; 272000 was a guess and read as an over-full context). Effort vocabulary is `low | medium |
gpt-5.6-sol | 258400 | Flagship — complex coding, research, deepest reasoning. | 2026-09-03 |
gpt-5.6-terra | 258400 | Workhorse; OpenAI positions it as GPT-5.5-class at lower cost. | 2026-09-05 (app-server engine) |
gpt-5.6-luna | 258400 | Fast and cheap — extraction, classification, volume. | 2026-09-03 |
gpt-5.5 | 258400 | Still listed by Codex; the one model here that does not run code-mode-only. Terra is the equivalent at lower cost. | 2026-09-05 (app-server engine) |
gpt-5.4-mini, gpt-5.3-codex | — | Retired for ChatGPT accounts — Codex answers with an error, which surfaces as a failed turn or a crashed compaction worker. Remove them. | 2026-08-31 |
Peculiarities of this engine
- Ask codex for the window instead of copying it from a model card. Codex runs a session against a window it delivers itself and reports on every turn (
modelContextWindow): 258,400 for every model it offers, measured 2026-09-11 on codex 0.153.3, while the API window is 1.05M. somora shows that reported number once the first turn has run, so a wrong configured value only misleads until then — but it also decides whether the model is picked as a compaction summariser, so keep it right. 272000, the figure that circulated in Codex issues and release notes, is 5 % too high and made a long thread read as over-full. - Codex compacts the thread itself; somora's
triggerRatiodoes not apply.contextWindowfeeds worker choice and display only. - Reasoning vocabulary is
minimal | low | medium | high | xhigh | maxfor the GPT-5.6 family. somora'shighis sent ashighunless you map it (levels: { high: xhigh }).maxis documented by OpenAI for the hardest problems with cost and latency to match — map it in a session when you need it, not as a default. - Tools are Codex dynamic tools in the
somoranamespace (somora_directfor image-bearing results,somora_mcp_<server>for external MCP servers). The tools incodexCli.directToolssit in the model's direct list every turn; the rest is deferred and reached via Codex tool search — the developer instructions list their names. - Thinking text: Codex emits reasoning summaries per thinking phase (
summary: autoonturn/start), streamed as the thinking block. - Code Mode. Since Codex 0.153 the GPT-5.6 family and GPT-6 are
code_mode_only: the model writes a JavaScript cell that callstools.somora.<tool>(...)(deferred tools are found throughALL_TOOLS), Codex's built-in host runs the cell, somora runs the tools. That sandbox has norequire,process,fetchor filesystem (verified 2026-09-05). gpt-5.5 calls the same dynamic tools directly.somora codex debug modelsshows thetool_modeper model.
grok-cli — SuperGrok / Premium subscription
Driven over ACP. Community-maintained adapter, text attachments only, no one-shot path (it cannot be a dream or compaction worker).
providers:
xai:
engine: grok-cli
models:
- id: grok-4.5
alias: grok
contextWindow: 500000
capabilities: [text, reasoning]| model | contextWindow | notes | verified |
|---|---|---|---|
grok-4.5 | 500000 | Effort levels pass through --reasoning-effort; reasoning.levels honoured. Thinking-text path unverified (no account on the maintainer's side). | 2026-08-25 |
openai-compatible — self-hosted (vLLM, SGLang, oMLX, Ollama, LM Studio)
The engine somora owns end to end: its own agent loop, sampling parameters, per-model reasoning vocabulary, compaction trigger. The values below are the vendor-recommended defaults for each model family, applied through somora's sampling: block (sampling.md).
providers:
local:
engine: openai-compatible
baseUrl: http://<your-host>:8000/v1
apiKey: "<key or anything for a local server>"
# sendUserTag: true # default — `user: "<agent>/<session>"` on every request so a
# gateway can attribute spend per agent; false to withhold
models:
- id: deepseek-v4-flash # SGLang; behind LiteLLM use the route name (e.g. deepseek-v4-flash-0731)
alias: deep4flash
contextWindow: 700000 # = the server's --context-length, not the model's 1M
capabilities: [text, reasoning]
sampling: { temperature: 1.0, top_p: 0.95 }
reasoning:
levels: { medium: high, high: max }
- id: deepseek-v4.1-flash # SGLang TP=4 (community SM120 runtime), Engram NVMe offload
alias: deep41flash
contextWindow: 700000 # = --context-length
capabilities: [text, image, reasoning]
sampling: { temperature: 1.0, top_p: 0.95 }
maxTokens: 16384 # operational output budget INCLUDING reasoning, not the model's max
reasoning:
levels: { "off": none, low: low, medium: high, high: max }
- id: glm-5.3-flash # vLLM TP=4, native FP8; behind LiteLLM use the route name
alias: glm
contextWindow: 700000 # = --max-model-len (model-native 1M; 262k in the reference recipe)
capabilities: [text, image, reasoning]
sampling: { temperature: 1.0, top_p: 0.95 } # = the model's generation_config.json (vLLM applies it when nothing is sent)
maxTokens: 16384
reasoning:
levels: { "off": low, low: low, medium: high, high: max } # vocabulary is ONLY low/high/max; anything else = max
- id: deepseek-v4-flash-vision-exp # SGLang TP=2, same backbone as deepseek-v4-flash + ViT; experimental
alias: deep4vision
contextWindow: 700000 # = --context-length (KV pool 859k tokens at mem-fraction 0.90)
capabilities: [text, image, reasoning]
sampling: { temperature: 1.0, top_p: 0.95 } # as deepseek-v4-flash
reasoning:
levels: { medium: high, high: max } # as deepseek-v4-flash
- id: qwen3.8-flash-next # vLLM, FP8
alias: qwen38next
contextWindow: 524288 # = --max-model-len (YaRN)
capabilities: [text, image, reasoning]
sampling: { temperature: 0.6, top_p: 0.95, top_k: 20, min_p: 0 }
maxTokens: 16384
reasoning:
levels: { "off": low, high: xhigh }
- id: qwen3.5-397b-a17b-awq # vLLM, AWQ INT4
alias: qwen35big
contextWindow: 262144
capabilities: [text, image, reasoning]
sampling: { temperature: 0.6, top_p: 0.95, top_k: 20, min_p: 0 }
reasoning:
levels: { "off": low, high: xhigh }
- id: qwen3.8-27b-fp8 # vLLM, dense
alias: qwen38small
contextWindow: 262144
capabilities: [text, image, reasoning]
sampling: { temperature: 0.6, top_p: 0.95, top_k: 20, min_p: 0 }
reasoning:
levels: { "off": low, high: xhigh }
- id: gemma-4-31b-it-8bit # oMLX (Apple Silicon)
alias: gemma4big
contextWindow: 131072
capabilities: [text, image]
sampling: { temperature: 1.0, top_p: 0.95, top_k: 64 }
- id: gemma-4-26b-a4b-it-4bit # oMLX
alias: gemma4small
contextWindow: 131072
capabilities: [text, image]
sampling: { temperature: 1.0, top_p: 0.95, top_k: 64 }| model family | server | contextWindow | sampling | reasoning | notes | verified |
|---|---|---|---|---|---|---|
| DeepSeek V4 Flash (284B MoE, 13B active) | SGLang, TP=4 | the server's --context-length (700000 on a two-GPU profile; 1M only with the whole box) | 1.0 / 0.95 (DeepSeek's agentic/coding recommendation) | vocabulary low, high, max; unknown values are ignored, not rejected → map high: max. off: omitting the parameter and sending none both give 0 reasoning tokens (measured 2026-09-05), so no off mapping is needed; low reasons via reasoning_content (84 tokens on a one-line prompt). On longer prompts without the parameter the reasoning can land inline in the text (…</think>); somora splits that into the thinking channel. | no vision; streams reasoning_content, full thinking text in the chat; tool calls parsed server-side | 2026-09-05 |
| DeepSeek V4.1 Flash | SGLang, TP=4 (community SM120 runtime, Engram NVMe offload) | --context-length 700000 | 1.0 / 0.95 (model card; top_k and penalties unset) | vocabulary none, low, high, max → map off: none, medium: high, high: max. Measured in four somora sessions: 0 / 55 / 48 / 92 reasoning tokens — the length is stochastic, it does not rise monotonically with the level | native vision (no analyze_file worker involved); three real tools in one session (time, file read, remote exec) over 2 rounds; maxTokens: 16384 is an operational budget that INCLUDES reasoning | 2026-09-11 (somora: thinking levels, native image, tools, multi-turn; gateway: JSON/tools/streaming; synthetic 697k-token request answered correctly in 107 s) |
| GLM-5.3-Flash (321B MoE, 18B active, KDA + sparse-MLA hybrid) | vLLM, TP=4, native FP8, MTP-4, fp8 KV | --max-model-len 700000 (KV pool 1.16M tokens with --kv-cache-memory 7.5 GiB; needle found at 427k/555k/651k/695k) | 1.0 / 0.95 (generation_config.json; vLLM --generation-config auto applies it server-side when the request omits sampling) | vocabulary low, high, max only; any other value — including none and medium — silently becomes max (measured: none → 2169 reasoning chars, low → 0). Top-level reasoning_effort and chat_template_kwargs.reasoning_effort are equivalent. max makes the model write whole solutions inside reasoning_content and hit output caps (31.9k reasoning / 100 output tokens, finish_reason: length in an OpenCode run) → map high: max only for deliberate use, medium: high, off: low | native vision (1036 image tokens for 1024×768, correct description; declines to hallucinate unreadable text); tool calls via glm47 parser, reasoning via glm45; ~95 tok/s prose, 200–210 tok/s code (MTP), prefill ~45k tok/s; prefix cache reuses >600k tokens (647k cached → 2 s); 4 GPUs = exclusive profile | 2026-09-06 (backend probes: PONG, tool calls, vision, long-context needles; somora engine matrix pending) |
| DeepSeek-V4-Flash-Vision-Exp (284B MoE, 13B active + 0.5B ViT) | SGLang, TP=2 (preview image sgl-project/sglang#37253 + sm_120 patches) | --context-length 700000 (pool 859k at mem-fraction 0.90; 0.85 gives only 203k) | 1.0 / 0.95 | same vocabulary and behaviour as DeepSeek V4 Flash (low/high/max, unknown ignored) | vision verified on synthetic + photo (208–356 image tokens); text/agent quality ≈ Flash-0731 (tool call deepseekv4 parser verified), 81 tok/s single-stream TP=2; on sm_120 image spans use causal windows (local patch) — small-text OCR is shaky, screenshots/charts fine; experimental per DeepSeek | 2026-09-06 (backend probes; somora engine matrix pending) |
| Qwen3.8-Flash-Next FP8 (176B / 6B active) | vLLM, TP=4 | --max-model-len, 524288 with YaRN 2.0 | 0.6 / 0.95 / 20 / min_p 0 (Qwen thinking-mode defaults) | knows no high → 400 unless mapped; none low medium xhigh accepted; unset = model default = thinks (61 reasoning tokens on a one-line prompt, low 55, xhigh 60, none 0, measured 2026-09-05) → map off: none, high: xhigh | vision verified; maxTokens: 16384 because reasoning otherwise eats short answers; parsers qwen3 + qwen3_xml | 2026-09-05 |
| Qwen3.5-397B-A17B AWQ INT4 | vLLM, TP=4 | 262144 (max_model_len) | same as above | as above, but none unverified on this backend — keep off: low until probed | vision verified; hermes tool parser | 2026-09-03 |
| Qwen3.8-27B FP8, dense | vLLM, 1 GPU | 262144 | same as above | as above, none unverified — keep off: low until probed | vision; qwen3_coder tool parser verified | 2026-09-03 |
| Gemma 4 31B / 26B-A4B | oMLX | 131072 | 1.0 / 0.95 / 64 (Gemma team recommendation) | none — no reasoning capability | vision; prefill memory guard answers 400 on long prompts → somora's reactive compaction handles it | 2026-09-03 |
Peculiarities of this engine
contextWindowis the compaction wall here (triggerRatio, default 0.8). Use the server's limit, not the model card's: a 1M value against a 700k backend meant400 ContextWindowExceededErrorbefore somora ever compacted.- A router in front changes the rules. LiteLLM drops
reasoning_effortunless the model'sallowed_openai_paramsincludes it — every level then answers 200 with identical reasoning volume and somora's retry-on-400 never fires. Fix it in the router, then verify with/thinking highvs/thinking lowand the 🧠 count (thinking.md). - A self-hosted build can be a community runtime. The DeepSeek V4.1 row above was measured on a community SM120 image with NVMe offload, not on stock SGLang. The configuration is what was tested there; treat the numbers as that setup's behaviour, not as a property of the model everywhere.
- Reasoning vocabularies differ per family and unknown values are not always ignored. GLM-5.x treats anything but
low/high/maxasmax— thenonethat somora sends Qwen foroffis full reasoning there. Always mapoffexplicitly to the family's lowest level and probe with a one-line prompt (reasoning chars must be 0). - vLLM applies the model's
generation_config.jsonwhen a request sends no sampling params (--generation-config auto, the default) — thesampling:block in somora is then a documented value, not the only place the default lives. SGLang does not do this. - Backends that stream reasoning but report no
reasoning_tokensget an estimated 🧠 count with a tilde. analyze_file(vision worker) uses this engine only — a CLI-engine model cannot be a vision worker. Dream and compaction workers may also run onclaude-cliandcodex-cli; onlygrok-clihas no one-shot path.
openai-compatible — hosted via OpenRouter
Same engine, wire-specific details: reasoning goes as a nested reasoning: { effort } object (reasoning.param: reasoning), PDFs can go native (pdfMode: native), image references travel as data URLs.
providers:
openrouter:
engine: openai-compatible
baseUrl: https://openrouter.ai/api/v1
apiKey: "<your key>"
pdfMode: native
models:
- id: anthropic/claude-haiku-4.5
alias: orhaiku
contextWindow: 200000
capabilities: [text, image, pdf]
- id: minimax/minimax-m3
alias: minimax3
contextWindow: 1048576
capabilities: [text, image, reasoning]
sampling: { temperature: 1.0, top_p: 0.95, top_k: 40 }
reasoning: { param: reasoning }
- id: moonshotai/kimi-k3
alias: kimi3
contextWindow: 1048576
capabilities: [text, image, reasoning]
reasoning: { param: reasoning }
- id: deepseek/deepseek-v4-pro-0813
alias: deep4pro
contextWindow: 1048576
capabilities: [text, reasoning]
sampling: { temperature: 1.0, top_p: 0.95 }
reasoning: { param: reasoning }| model | contextWindow | sampling | notes | verified |
|---|---|---|---|---|
anthropic/claude-haiku-4.5 | 200000 | — | Useful as the last, always-reachable vision worker in a vision.worker chain — the subscription Haiku cannot be one (CLI engines are not vision workers). Costs API money. | 2026-09-03 |
minimax/minimax-m3 | 1048576 | 1.0 / 0.95 / 40 (model card) | image and video input; vocabulary low/medium/high = somora's default, no levels needed | 2026-09-03 |
moonshotai/kimi-k3 | 1048576 | none — on purpose. Moonshot fixes temperature 1.0 / top_p 0.95 server-side and documents "leave the parameters out". | vocabulary low/medium/high | 2026-09-03 |
deepseek/deepseek-v4-pro-0813 | 1048576 | 1.0 / 0.95 | no vision; the hosted big sibling of a local V4 Flash — a good fallback when the local box is off | 2026-09-03 |
Watch the fallback. When a hosted key expires, somora's model fallback chain answers the turn with the next configured model and the chat header shows that model — the engine.fail line in the server log (401 API key expired) is the only loud signal. Check the header, not just the reply.
What each field does, per engine
| field | openai-compatible | claude-cli | codex-cli | grok-cli |
|---|---|---|---|---|
contextWindow | compaction trigger (triggerRatio ×), worker choice, display — use the server's limit | worker choice, display — native window is right | worker choice, display — use the Codex session window (258,400, what codex reports as modelContextWindow), not the API window | worker choice, display |
capabilities | gates attachments (image, pdf) and whether thinking is sent (reasoning) | same | same | same |
reasoning.levels | somora level → wire value (Qwen xhigh, DeepSeek max) | honoured (rarely needed) | honoured — xhigh/max for GPT-5.6 | honoured |
reasoning.param | reasoning_effort (default), nested reasoning for OpenRouter, or chat_template_kwargs | — | — | — |
sampling | sent on every call, dropped once if the backend rejects a key | ignored (not exposed by the CLI) | ignored | ignored |
maxTokens | output cap on every call incl. dream workers | — | — | — |
fallback | availability chain on unreachable / 5xx | same | same | same |
sendUserTag (provider) | user: "<agent>/<session>" on every request, <agent>/rem, <agent>/deep, lucid/<pass>, <agent>/compaction, <agent>/analyze_file for workers — a gateway groups cost per agent and session (LiteLLM stores it in the end_user column of its spend logs, not user); default on, false to withhold | — | — | — |
Minimum versions: Node.js ≥ 22.13 (every somora command refuses an older Node), a current Claude Code. Codex needs no separate install — somora bundles the exact version pinned in its package.json. A CLI engine's tools and lock-down are re-audited after every CLI update (security.md).