Skip to content

Models known to run with somora — per engine, with the settings that work ​

Configuring a model for somora is not "the model", it is "the model behind this engine": the same GPT-5.6 has a different usable context through Codex than through an API, Qwen wants xhigh where somora says high, Kimi refuses sampling parameters, and a value copied from a model card can quietly mis-size the compaction worker. This page collects what has been verified against real somora installs — the recommended config.yaml block per model, and why each value is what it is. It is maintained by hand; the verified column says when a row was last checked. Corrections and new rows are welcome as pull requests.

Field semantics are explained once, at the end of the page, in What each field does, per engine. Read compaction.md for the contextWindow story in depth.

Aliases below are suggestions — pick your own.

What "verified" means. A dated row went through somora's engine matrix on that day, on a live install: basic tools (time_now, exec, file_read, tmux, memory_search, a deferred tool), tools on an SSH resource (exec, file_write / file_read / file_list, tmux on another host), web_fetch, agent_ask to another agent, spawn_subagent, an image attachment (where the model has vision), /thinking high, an abort mid-turn, and session memory across turns. Last full run: 2026-09-05 — claude-fable-5-1, gpt-5.6-terra, gpt-5.5, deepseek-v4-flash, qwen3.8-flash-next, all green (DeepSeek: image attachment refused by design, no vision).


claude-cli — Claude Code subscription ​

Runs through the Claude Agent SDK on your claude login. No API key, no sampling knobs (the CLI does not expose them), reasoning via the SDK's effort levels.

yaml
providers:
  anthropic:
    engine: claude-cli
    models:
      - id: claude-fable-5-1
        alias: fable
        contextWindow: 1000000
        capabilities: [text, image, pdf, reasoning]
      - id: claude-opus-5-5
        alias: opus
        contextWindow: 1000000
        capabilities: [text, image, pdf, reasoning]
      - id: claude-sonnet-5
        alias: sonnet
        contextWindow: 1000000
        capabilities: [text, image, pdf, reasoning]
      - id: claude-haiku-4-5
        alias: haiku
        contextWindow: 200000
        capabilities: [text, image, pdf, reasoning]
modelcontextWindownotesverified
claude-fable-5-11000000Frontier model, adaptive thinking always on. The SDK discloses no thinking text for it — somora shows a placeholder row that the model thought (thinking.md). No separate reasoning-token count (rolled into tokens_out).2026-09-03
claude-opus-5-51000000Opus 5.5 — Anthropic's recommendation for most workloads; replaces claude-opus-5 (keep your alias, change the id). Thinking text: placeholder, as above.2026-09-30
claude-sonnet-51000000Same surface as Opus, cheaper on the subscription budget.2026-09-03
claude-haiku-4-5200000Does support extended thinking — keep reasoning in capabilities, it is missing from most example configs.2026-09-03

Peculiarities of this engine

  • contextWindow is the model's real window here: Claude Code sessions run against it and compact on their own. The value only feeds the compaction-worker choice and the header percentage.
  • Tools reach the model through ToolSearch (Claude Code 2.1.142+): somora's tools are discovered on demand, so the first call in a session is preceded by a ToolSearch row. Normal.
  • thinking levels map to the SDK's effort; off disables thinking. reasoning.levels is honoured but rarely needed — the vocabulary is already low | medium | high.
  • The claude.ai connectors (Gmail, Calendar, Drive) never exist inside a somora session (security.md).

codex-cli — ChatGPT subscription ​

Runs the bundled Codex (@openai/codex, exact version pinned in somora's package.json) as an app-server per turn, on your codex login. somora hands its tools to Codex as dynamic tools — no MCP child, no deferred-namespace guessing — so every model Codex offers (gpt-5.5, the GPT-5.6 family, GPT-6) reaches the same tool set the other engines see. A global codex on the host is not used; somora codex login signs in with the bundled one, and an existing codex login is picked up automatically (setup.md).

yaml
providers:
  openai:
    engine: codex-cli
    models:
      - id: gpt-6-astra
        alias: astra
        contextWindow: 258400          # what codex reports as modelContextWindow (measured 2026-09-11)
        capabilities: [text, image, pdf, reasoning]
        reasoning:
          levels: { "off": low, high: xhigh }   # Astra has no `minimal` — map `off` explicitly; `max` deliberately not a default
      - id: gpt-5.6-sol
        alias: gpt56
        contextWindow: 258400          # the codex session window, NOT the 1.05M API window
        capabilities: [text, image, pdf, reasoning]
        reasoning:
          levels: { high: xhigh }      # optional: /thinking high → codex xhigh
      - id: gpt-5.6-terra
        alias: terra
        contextWindow: 258400
        capabilities: [text, image, pdf, reasoning]
      - id: gpt-5.6-luna
        alias: luna
        contextWindow: 258400
        capabilities: [text, image, pdf, reasoning]
      - id: gpt-5.5
        alias: gpt55
        contextWindow: 258400
        capabilities: [text, image, pdf, reasoning]
modelcontextWindownotesverified
gpt-6-astra258400GPT-6 — hardest problems; code-mode-only like the 5.6 family. Every codex model here reports the same modelContextWindow of 258,400 (measured 2026-09-11 against codex 0.153.3; 272000 was a guess and read as an over-full context). Effort vocabulary is `lowmedium
gpt-5.6-sol258400Flagship — complex coding, research, deepest reasoning.2026-09-03
gpt-5.6-terra258400Workhorse; OpenAI positions it as GPT-5.5-class at lower cost.2026-09-05 (app-server engine)
gpt-5.6-luna258400Fast and cheap — extraction, classification, volume.2026-09-03
gpt-5.5258400Still listed by Codex; the one model here that does not run code-mode-only. Terra is the equivalent at lower cost.2026-09-05 (app-server engine)
gpt-5.4-mini, gpt-5.3-codex—Retired for ChatGPT accounts — Codex answers with an error, which surfaces as a failed turn or a crashed compaction worker. Remove them.2026-08-31

Peculiarities of this engine

  • Ask codex for the window instead of copying it from a model card. Codex runs a session against a window it delivers itself and reports on every turn (modelContextWindow): 258,400 for every model it offers, measured 2026-09-11 on codex 0.153.3, while the API window is 1.05M. somora shows that reported number once the first turn has run, so a wrong configured value only misleads until then — but it also decides whether the model is picked as a compaction summariser, so keep it right. 272000, the figure that circulated in Codex issues and release notes, is 5 % too high and made a long thread read as over-full.
  • Codex compacts the thread itself; somora's triggerRatio does not apply. contextWindow feeds worker choice and display only.
  • Reasoning vocabulary is minimal | low | medium | high | xhigh | max for the GPT-5.6 family. somora's high is sent as high unless you map it (levels: { high: xhigh }). max is documented by OpenAI for the hardest problems with cost and latency to match — map it in a session when you need it, not as a default.
  • Tools are Codex dynamic tools in the somora namespace (somora_direct for image-bearing results, somora_mcp_<server> for external MCP servers). The tools in codexCli.directTools sit in the model's direct list every turn; the rest is deferred and reached via Codex tool search — the developer instructions list their names.
  • Thinking text: Codex emits reasoning summaries per thinking phase (summary: auto on turn/start), streamed as the thinking block.
  • Code Mode. Since Codex 0.153 the GPT-5.6 family and GPT-6 are code_mode_only: the model writes a JavaScript cell that calls tools.somora.<tool>(...) (deferred tools are found through ALL_TOOLS), Codex's built-in host runs the cell, somora runs the tools. That sandbox has no require, process, fetch or filesystem (verified 2026-09-05). gpt-5.5 calls the same dynamic tools directly. somora codex debug models shows the tool_mode per model.

grok-cli — SuperGrok / Premium subscription ​

Driven over ACP. Community-maintained adapter, text attachments only, no one-shot path (it cannot be a dream or compaction worker).

yaml
providers:
  xai:
    engine: grok-cli
    models:
      - id: grok-4.5
        alias: grok
        contextWindow: 500000
        capabilities: [text, reasoning]
modelcontextWindownotesverified
grok-4.5500000Effort levels pass through --reasoning-effort; reasoning.levels honoured. Thinking-text path unverified (no account on the maintainer's side).2026-08-25

openai-compatible — self-hosted (vLLM, SGLang, oMLX, Ollama, LM Studio) ​

The engine somora owns end to end: its own agent loop, sampling parameters, per-model reasoning vocabulary, compaction trigger. The values below are the vendor-recommended defaults for each model family, applied through somora's sampling: block (sampling.md).

yaml
providers:
  local:
    engine: openai-compatible
    baseUrl: http://<your-host>:8000/v1
    apiKey: "<key or anything for a local server>"
    # sendUserTag: true                   # default — `user: "<agent>/<session>"` on every request so a
                                          # gateway can attribute spend per agent; false to withhold
    models:
      - id: deepseek-v4-flash             # SGLang; behind LiteLLM use the route name (e.g. deepseek-v4-flash-0731)
        alias: deep4flash
        contextWindow: 700000             # = the server's --context-length, not the model's 1M
        capabilities: [text, reasoning]
        sampling: { temperature: 1.0, top_p: 0.95 }
        reasoning:
          levels: { medium: high, high: max }
      - id: deepseek-v4.1-flash           # SGLang TP=4 (community SM120 runtime), Engram NVMe offload
        alias: deep41flash
        contextWindow: 700000             # = --context-length
        capabilities: [text, image, reasoning]
        sampling: { temperature: 1.0, top_p: 0.95 }
        maxTokens: 16384                  # operational output budget INCLUDING reasoning, not the model's max
        reasoning:
          levels: { "off": none, low: low, medium: high, high: max }
      - id: glm-5.3-flash                 # vLLM TP=4, native FP8; behind LiteLLM use the route name
        alias: glm
        contextWindow: 700000             # = --max-model-len (model-native 1M; 262k in the reference recipe)
        capabilities: [text, image, reasoning]
        sampling: { temperature: 1.0, top_p: 0.95 }   # = the model's generation_config.json (vLLM applies it when nothing is sent)
        maxTokens: 16384
        reasoning:
          levels: { "off": low, low: low, medium: high, high: max }   # vocabulary is ONLY low/high/max; anything else = max
      - id: deepseek-v4-flash-vision-exp  # SGLang TP=2, same backbone as deepseek-v4-flash + ViT; experimental
        alias: deep4vision
        contextWindow: 700000             # = --context-length (KV pool 859k tokens at mem-fraction 0.90)
        capabilities: [text, image, reasoning]
        sampling: { temperature: 1.0, top_p: 0.95 }   # as deepseek-v4-flash
        reasoning:
          levels: { medium: high, high: max }         # as deepseek-v4-flash
      - id: qwen3.8-flash-next            # vLLM, FP8
        alias: qwen38next
        contextWindow: 524288             # = --max-model-len (YaRN)
        capabilities: [text, image, reasoning]
        sampling: { temperature: 0.6, top_p: 0.95, top_k: 20, min_p: 0 }
        maxTokens: 16384
        reasoning:
          levels: { "off": low, high: xhigh }
      - id: qwen3.5-397b-a17b-awq         # vLLM, AWQ INT4
        alias: qwen35big
        contextWindow: 262144
        capabilities: [text, image, reasoning]
        sampling: { temperature: 0.6, top_p: 0.95, top_k: 20, min_p: 0 }
        reasoning:
          levels: { "off": low, high: xhigh }
      - id: qwen3.8-27b-fp8               # vLLM, dense
        alias: qwen38small
        contextWindow: 262144
        capabilities: [text, image, reasoning]
        sampling: { temperature: 0.6, top_p: 0.95, top_k: 20, min_p: 0 }
        reasoning:
          levels: { "off": low, high: xhigh }
      - id: gemma-4-31b-it-8bit           # oMLX (Apple Silicon)
        alias: gemma4big
        contextWindow: 131072
        capabilities: [text, image]
        sampling: { temperature: 1.0, top_p: 0.95, top_k: 64 }
      - id: gemma-4-26b-a4b-it-4bit       # oMLX
        alias: gemma4small
        contextWindow: 131072
        capabilities: [text, image]
        sampling: { temperature: 1.0, top_p: 0.95, top_k: 64 }
model familyservercontextWindowsamplingreasoningnotesverified
DeepSeek V4 Flash (284B MoE, 13B active)SGLang, TP=4the server's --context-length (700000 on a two-GPU profile; 1M only with the whole box)1.0 / 0.95 (DeepSeek's agentic/coding recommendation)vocabulary low, high, max; unknown values are ignored, not rejected → map high: max. off: omitting the parameter and sending none both give 0 reasoning tokens (measured 2026-09-05), so no off mapping is needed; low reasons via reasoning_content (84 tokens on a one-line prompt). On longer prompts without the parameter the reasoning can land inline in the text (…</think>); somora splits that into the thinking channel.no vision; streams reasoning_content, full thinking text in the chat; tool calls parsed server-side2026-09-05
DeepSeek V4.1 FlashSGLang, TP=4 (community SM120 runtime, Engram NVMe offload)--context-length 7000001.0 / 0.95 (model card; top_k and penalties unset)vocabulary none, low, high, max → map off: none, medium: high, high: max. Measured in four somora sessions: 0 / 55 / 48 / 92 reasoning tokens — the length is stochastic, it does not rise monotonically with the levelnative vision (no analyze_file worker involved); three real tools in one session (time, file read, remote exec) over 2 rounds; maxTokens: 16384 is an operational budget that INCLUDES reasoning2026-09-11 (somora: thinking levels, native image, tools, multi-turn; gateway: JSON/tools/streaming; synthetic 697k-token request answered correctly in 107 s)
GLM-5.3-Flash (321B MoE, 18B active, KDA + sparse-MLA hybrid)vLLM, TP=4, native FP8, MTP-4, fp8 KV--max-model-len 700000 (KV pool 1.16M tokens with --kv-cache-memory 7.5 GiB; needle found at 427k/555k/651k/695k)1.0 / 0.95 (generation_config.json; vLLM --generation-config auto applies it server-side when the request omits sampling)vocabulary low, high, max only; any other value — including none and medium — silently becomes max (measured: none → 2169 reasoning chars, low → 0). Top-level reasoning_effort and chat_template_kwargs.reasoning_effort are equivalent. max makes the model write whole solutions inside reasoning_content and hit output caps (31.9k reasoning / 100 output tokens, finish_reason: length in an OpenCode run) → map high: max only for deliberate use, medium: high, off: lownative vision (1036 image tokens for 1024×768, correct description; declines to hallucinate unreadable text); tool calls via glm47 parser, reasoning via glm45; ~95 tok/s prose, 200–210 tok/s code (MTP), prefill ~45k tok/s; prefix cache reuses >600k tokens (647k cached → 2 s); 4 GPUs = exclusive profile2026-09-06 (backend probes: PONG, tool calls, vision, long-context needles; somora engine matrix pending)
DeepSeek-V4-Flash-Vision-Exp (284B MoE, 13B active + 0.5B ViT)SGLang, TP=2 (preview image sgl-project/sglang#37253 + sm_120 patches)--context-length 700000 (pool 859k at mem-fraction 0.90; 0.85 gives only 203k)1.0 / 0.95same vocabulary and behaviour as DeepSeek V4 Flash (low/high/max, unknown ignored)vision verified on synthetic + photo (208–356 image tokens); text/agent quality ≈ Flash-0731 (tool call deepseekv4 parser verified), 81 tok/s single-stream TP=2; on sm_120 image spans use causal windows (local patch) — small-text OCR is shaky, screenshots/charts fine; experimental per DeepSeek2026-09-06 (backend probes; somora engine matrix pending)
Qwen3.8-Flash-Next FP8 (176B / 6B active)vLLM, TP=4--max-model-len, 524288 with YaRN 2.00.6 / 0.95 / 20 / min_p 0 (Qwen thinking-mode defaults)knows no high → 400 unless mapped; none low medium xhigh accepted; unset = model default = thinks (61 reasoning tokens on a one-line prompt, low 55, xhigh 60, none 0, measured 2026-09-05) → map off: none, high: xhighvision verified; maxTokens: 16384 because reasoning otherwise eats short answers; parsers qwen3 + qwen3_xml2026-09-05
Qwen3.5-397B-A17B AWQ INT4vLLM, TP=4262144 (max_model_len)same as aboveas above, but none unverified on this backend — keep off: low until probedvision verified; hermes tool parser2026-09-03
Qwen3.8-27B FP8, densevLLM, 1 GPU262144same as aboveas above, none unverified — keep off: low until probedvision; qwen3_coder tool parser verified2026-09-03
Gemma 4 31B / 26B-A4BoMLX1310721.0 / 0.95 / 64 (Gemma team recommendation)none — no reasoning capabilityvision; prefill memory guard answers 400 on long prompts → somora's reactive compaction handles it2026-09-03

Peculiarities of this engine

  • contextWindow is the compaction wall here (triggerRatio, default 0.8). Use the server's limit, not the model card's: a 1M value against a 700k backend meant 400 ContextWindowExceededError before somora ever compacted.
  • A router in front changes the rules. LiteLLM drops reasoning_effort unless the model's allowed_openai_params includes it — every level then answers 200 with identical reasoning volume and somora's retry-on-400 never fires. Fix it in the router, then verify with /thinking high vs /thinking low and the 🧠 count (thinking.md).
  • A self-hosted build can be a community runtime. The DeepSeek V4.1 row above was measured on a community SM120 image with NVMe offload, not on stock SGLang. The configuration is what was tested there; treat the numbers as that setup's behaviour, not as a property of the model everywhere.
  • Reasoning vocabularies differ per family and unknown values are not always ignored. GLM-5.x treats anything but low/high/max as max — the none that somora sends Qwen for off is full reasoning there. Always map off explicitly to the family's lowest level and probe with a one-line prompt (reasoning chars must be 0).
  • vLLM applies the model's generation_config.json when a request sends no sampling params (--generation-config auto, the default) — the sampling: block in somora is then a documented value, not the only place the default lives. SGLang does not do this.
  • Backends that stream reasoning but report no reasoning_tokens get an estimated 🧠 count with a tilde.
  • analyze_file (vision worker) uses this engine only — a CLI-engine model cannot be a vision worker. Dream and compaction workers may also run on claude-cli and codex-cli; only grok-cli has no one-shot path.

openai-compatible — hosted via OpenRouter ​

Same engine, wire-specific details: reasoning goes as a nested reasoning: { effort } object (reasoning.param: reasoning), PDFs can go native (pdfMode: native), image references travel as data URLs.

yaml
providers:
  openrouter:
    engine: openai-compatible
    baseUrl: https://openrouter.ai/api/v1
    apiKey: "<your key>"
    pdfMode: native
    models:
      - id: anthropic/claude-haiku-4.5
        alias: orhaiku
        contextWindow: 200000
        capabilities: [text, image, pdf]
      - id: minimax/minimax-m3
        alias: minimax3
        contextWindow: 1048576
        capabilities: [text, image, reasoning]
        sampling: { temperature: 1.0, top_p: 0.95, top_k: 40 }
        reasoning: { param: reasoning }
      - id: moonshotai/kimi-k3
        alias: kimi3
        contextWindow: 1048576
        capabilities: [text, image, reasoning]
        reasoning: { param: reasoning }
      - id: deepseek/deepseek-v4-pro-0813
        alias: deep4pro
        contextWindow: 1048576
        capabilities: [text, reasoning]
        sampling: { temperature: 1.0, top_p: 0.95 }
        reasoning: { param: reasoning }
modelcontextWindowsamplingnotesverified
anthropic/claude-haiku-4.5200000—Useful as the last, always-reachable vision worker in a vision.worker chain — the subscription Haiku cannot be one (CLI engines are not vision workers). Costs API money.2026-09-03
minimax/minimax-m310485761.0 / 0.95 / 40 (model card)image and video input; vocabulary low/medium/high = somora's default, no levels needed2026-09-03
moonshotai/kimi-k31048576none — on purpose. Moonshot fixes temperature 1.0 / top_p 0.95 server-side and documents "leave the parameters out".vocabulary low/medium/high2026-09-03
deepseek/deepseek-v4-pro-081310485761.0 / 0.95no vision; the hosted big sibling of a local V4 Flash — a good fallback when the local box is off2026-09-03

Watch the fallback. When a hosted key expires, somora's model fallback chain answers the turn with the next configured model and the chat header shows that model — the engine.fail line in the server log (401 API key expired) is the only loud signal. Check the header, not just the reply.

What each field does, per engine ​

fieldopenai-compatibleclaude-clicodex-cligrok-cli
contextWindowcompaction trigger (triggerRatio ×), worker choice, display — use the server's limitworker choice, display — native window is rightworker choice, display — use the Codex session window (258,400, what codex reports as modelContextWindow), not the API windowworker choice, display
capabilitiesgates attachments (image, pdf) and whether thinking is sent (reasoning)samesamesame
reasoning.levelssomora level → wire value (Qwen xhigh, DeepSeek max)honoured (rarely needed)honoured — xhigh/max for GPT-5.6honoured
reasoning.paramreasoning_effort (default), nested reasoning for OpenRouter, or chat_template_kwargs———
samplingsent on every call, dropped once if the backend rejects a keyignored (not exposed by the CLI)ignoredignored
maxTokensoutput cap on every call incl. dream workers———
fallbackavailability chain on unreachable / 5xxsamesamesame
sendUserTag (provider)user: "<agent>/<session>" on every request, <agent>/rem, <agent>/deep, lucid/<pass>, <agent>/compaction, <agent>/analyze_file for workers — a gateway groups cost per agent and session (LiteLLM stores it in the end_user column of its spend logs, not user); default on, false to withhold———

Minimum versions: Node.js ≥ 22.13 (every somora command refuses an older Node), a current Claude Code. Codex needs no separate install — somora bundles the exact version pinned in its package.json. A CLI engine's tools and lock-down are re-audited after every CLI update (security.md).