Models

curated 2026-07-19 · live benchmark feeds later

What each model in the catalog is genuinely good at, in plain language. Cost per typical message assumes ~15.5k input + 150 output tokens, uncached, computed from the live rates in openclaw.json (the same source the spend log uses). No benchmark scores here — just honest guidance.

deepseek-v4-flash

Eco gear

$0.0022 per typical message

$0.14 in / $0.28 out per 1M tokens · cached read $0.028

Good at

The bargain bin that isn't junk: greetings, quick one-fact lookups, short rewrites and translations, high-volume chatter where 'competent and instant' beats 'brilliant'.

Overkill for

Nothing — it's the floor. The real risk is the opposite: it underpowers tool calls, code, and anything subtle. That's why Cruise never sends tool-y messages here.

Speed feel

Snappy — answers feel instant.

Ask it to…

settle quick facts all afternoon — 'what's the capital of X?' — for a fifth of a cent each.

gemini-3-5-flash

Comfort gear

$0.0050 per typical message

$0.3 in / $2.5 out per 1M tokens · cached read $0.03

Good at

The everyday workhorse. Reliable tool calling (Cruise's tool floor), summaries, drafting, solid reasoning at about a tenth of Sport-gear prices.

Overkill for

'hey' and one-word lookups — DeepSeek handles those for half the price. Also not the pick for the subtlest judgment calls; that's what Sport gears are for.

Speed feel

Fast — keeps up with conversation easily.

Ask it to…

read a file, search the web, summarize a thread — routine tool-y tasks it just handles.

gpt-5-6-luna

$0.016 per typical message

$1 in / $6 out per 1M tokens · cached read $0.1

Good at

The budget GPT: general chat, drafting, light reasoning with GPT polish at a fraction of the flagship's price.

Overkill for

Single-fact lookups — DeepSeek and Flash do those for less, just as well.

Speed feel

Fast.

Ask it to…

draft the email or brainstorm twenty names — GPT flavor on a budget.

gemini-3-1-pro

$0.033 per typical message

$2 in / $12 out per 1M tokens · cached read $0.2

Good at

Long-context and document work — Google's pro tier shines when you feed it a lot: big PDFs, whole repos, multimodal inputs.

Overkill for

Casual chat and short tasks, where Flash is nearly as good for a sixth of the price.

Speed feel

Moderate.

Ask it to…

chew through this long document and give me the structured brief.

gpt-5-6-terra

$0.041 per typical message

$2.5 in / $15 out per 1M tokens · cached read $0.25

Good at

The balanced GPT: most reasoning tasks, coding help, analysis — a sweet-spot price/quality point for everyday serious work.

Overkill for

Small talk and one-liners; the lower gears earn their keep there.

Speed feel

Moderate — you can watch it think, but barely.

Ask it to…

review this function, explain the bug, then rewrite it cleaner.

kimi-k3

Sport — dev primarySport — dev-desktop primarySport — dev-visual primary

$0.049 per typical message

$3 in / $15 out per 1M tokens · cached read $0.3

Good at

Strong reasoning and long context at a near-flagship price. The dev agent's daily driver — comfortable with code, debugging, and multi-step analysis.

Overkill for

Chatter and one-liners; Cruise sends those down-tier automatically.

Speed feel

Moderate.

Ask it to…

debug this hairy issue and plan the fix step by step.

claude-opus-4-8

Sport — default primarySport — main primary

$0.081 per typical message

$5 in / $25 out per 1M tokens · cached read $0.5

Good at

Top-tier reasoning, writing, and judgment with a distinctive voice. The main agent's primary — what 'Sport' means day to day. Nuanced long-form, careful decisions, code that reads like a human wrote it.

Overkill for

Greetings, lookups, anything routine — that's exactly what Cruise routes away from it.

Speed feel

Deliberate — takes its time and spends it well.

Ask it to…

write the message that has to land exactly right.

gpt-5-6-sol

$0.082 per typical message

$5 in / $30 out per 1M tokens · cached read $0.5

Good at

The GPT flagship: hardest reasoning, high-stakes drafting, complex code — when the answer has to be right and the problem has teeth.

Overkill for

Anything short or routine. A 'thanks' costs flagship money here.

Speed feel

Deliberate — flagship pace.

Ask it to…

reason through the gnarly architecture decision and stress-test my plan.

claude-fable-5

Max gear

$0.163 per typical message

$10 in / $50 out per 1M tokens · cached read $1

Good at

The top shelf — the hardest, longest, most nuanced work: big refactors, high-stakes documents, problems where the second-best answer isn't good enough.

Overkill for

Almost everything daily — it costs double Opus. Manual-only gear for a reason.

Speed feel

The slowest — big-model pace, worth it when it matters.

Ask it to…

take the whole problem and do it right, cost be damned.

Chevy — Models