1.3 Inference engines and models

Book 2 · The Delegation ContractChapter 1 · section 3 of 14

An inference engine interprets context and produces language, structured data, classifications, plans, code, or other predictions. The form it takes in 2026 is a large language model — hosted or local, multimodal or text-only — plus the supporting cast that shapes what it produces: embedding models that turn text into vectors, rerankers that order retrieved documents, classifiers that route work between systems. An inference engine has no organizational authority on its own. Its output becomes part of delegated intelligence only when another system gives it a role, an output contract, and a way to evaluate the result. In the component language of Chapter 3, the engine supplies the nondeterministic interpretation; everything that bounds it — tools, memory, directives, permission envelope — is supplied by the layers surveyed in the rest of this section.

1.3.1 The proprietary frontier

The proprietary frontier is a longer list than it was a year ago, and the conversation around it is louder than the conversation around any other layer in this chapter. These are the models that get the keynote slots, the magazine covers, and the breathless weekly launch coverage. They are also the models that providers frame as the most capable, the most dangerous, and the most worth paying for. It is what gets talked about, so the survey starts here.

The list of vendors with a credible proprietary frontier line at this writing runs from the American and Chinese platforms most readers have heard of, through a set of mostly Chinese labs whose products compete in the same league but rarely surface in English-language press, to one or two commercial holdouts that are positioning around specialization rather than raw capability:

  • OpenAI. GPT-6 Astra, released September 3, 2026, is the new flagship — a one-million-token-class model OpenAI priced at $10 input and $50 output per million tokens. The GPT-6 family filled out fast: GPT-6 Sol ($2 input, $10 output per million tokens) and GPT-6 Luna ($0.10 and $0.50) arrived September 22 at half their GPT-5.6 predecessors’ prices, leaving the GPT-5.6 family (Sol $4/$20 on a promotional rate, Terra $2/$12, Luna $0.20/$1.20) as the predecessor generation, with the “ultra” reasoning effort that runs four sub-agents in parallel.10
  • Anthropic. Claude Fable 5.1, generally available September 1, remains the premium rung, with list rates still $10 and $50 per million tokens and cheaper cache reads that Anthropic says cut typical bills about 25 percent and highly agentic ones as much as 45 percent. The Claude 5.5 family now fills most of the public ladder: Opus 5.5 (September 22) matches Fable 5.1 on most work at $4 and $20 per million tokens — about 40 percent less than Opus 5 to run on typical workloads, because it both costs less per token and uses fewer of them — and Sonnet 5.5 (September 28) holds Sonnet 5’s $2 and $10 while running more than 30 percent faster and costing up to 30 percent less per task. Haiku 5.5 is announced for the coming weeks; until it ships, Haiku 4.5 remains the floor. Mythos 5.1 is the same base model with different safeguards, reserved for trusted cybersecurity and life-sciences access.
  • Google. The Gemini 3 family still spans Pro through Flash. Gemini 3.8 Flash, released September 2, is the current workhorse Flash — $0.75 input and $3.75 output per million tokens on an introductory rate through the end of 2026 — with a restricted 3.8 Flash Cyber variant sold only through Google’s Fairwind program for trusted defenders. Gemini 3.1 Pro remains the enterprise pricing anchor in most cost comparisons.
  • xAI. Grok 4.7, released September 21 at Grok 4.6’s prices, tops out the line — a half-million-token context window and positioning aimed squarely at long-running agents.
  • Microsoft. The MAI family launched at Build 2026 in June, with MAI-Thinking-1 as the reasoning flagship. Microsoft reports the reasoning model matching Claude Opus 4.6 on SWE-Bench Pro at 53 percent — a vendor-reported comparison, but a respectable debut and a sign that the company is no longer content to resell someone else’s frontier.11
  • Amazon. The original Nova family shipped four text tiers (Micro, Lite, Pro, Premier) plus Canvas and Reel for image and video, all running natively inside AWS. Nova Forge lets enterprises start training from early checkpoints and blend their own data with Nova-curated training data on managed SageMaker infrastructure.
  • Baidu. ERNIE 5.0, released at Baidu World in November 2025 and updated through 2026, is a 2.4-trillion-parameter unified multimodal model. It ranks eighth globally on the LMArena Text Arena as of January 2026 — the highest-placed Chinese model on that leaderboard.
  • ByteDance. Seed 2.0 (released February 14, 2026) is the foundation behind the Doubao app, which serves more than two hundred million users. The flagship Pro variant is positioned as a direct competitor to GPT-5.2, Claude Opus 4.5, and Gemini 3 Pro, with reported pricing roughly a tenth of the Western frontier.
  • Alibaba. Qwen 3.8 Max is the flagship in the Qwen family at 2.4 trillion parameters (weights released August 2026), but the proprietary Qwen Plus / Max tier sold through Alibaba Cloud is the production entry point for most enterprise buyers.

What sets these models apart from the rest of the market is not one capability. It is the combination of three things: training compute measured in tens of millions of GPU-hours, a long context window with reliable tool use at the top end, and the engineering investment to keep the harnesses, SDKs, and hosted services coherent. Frontier work also costs a lot of money, and the frontier is no longer a corner of the technology industry — it is becoming a major component of the human economy. The five largest cloud and AI infrastructure providers have collectively committed between roughly $630 billion and $700 billion in capital spending for 2026 alone, most of it for data centers, chips, and networking — a sixty percent jump over the record set the year before.12 The money does not stop at the end of the fiscal year: those same companies have signed about $1.09 trillion in future lease payments for data centers that have not yet been built, commitments measured in trillions rather than billions, and the leases run for up to nineteen years.

A new headline-tier model arrives every week from somewhere on this list. Every release is covered as the new state of the art; every prior release is framed as having been surpassed. The conversation in industry coverage tracks the leaderboard more closely than it tracks what the systems actually do, and a working assumption — that the “best” model this month is meaningfully different from the “best” model last month — is mostly a marketing artifact. In practice the gap between adjacent flagship releases is small, the gaps between vendors at any given moment are smaller still, and the differences an orchestrator actually cares about — price per token, context window size, tool-use behavior, latency, deployment surface — are easier to read from a price sheet than from a benchmark chart.

1.3.2 The open-weight frontier

Open-weight models are now competing directly with the proprietary frontier models, and that competition is putting significant pressure on the frontier providers — to keep innovating, and to reduce pricing. The story is not only that the models have caught up. It is that the arrival of always-on open agents — most visibly OpenClaw, a local-first open-source runtime with hundreds of thousands of GitHub stars and a heartbeat scheduler that wakes the agent on a configurable interval — created a sustained demand for models that were cheap enough to call every few minutes without a human noticing the bill.13 A consumer chatbot that costs a dollar per long conversation is one thing; an always-on agent that exchanges dozens of small messages an hour, every hour, is another. OpenClaw users documented the shift publicly in 2026 — default routing to DeepSeek V3 or Kimi K2.5 for routine heartbeats, escalation to a Western flagship only for complex reasoning — and the resulting demand pull helped push Chinese and open-weight vendors to publish pricing at one-tenth to one-thirtieth of the Western flagships. The pressure on the proprietary frontier is not just competitive; it is structural. Once the always-on agent category existed, frontier-class models sold at $30 per million output tokens stopped being the default, and the frontier vendors responded with smaller, cheaper tiers — and, by late September 2026, with flagship-adjacent price cuts: GPT-6 Sol and Luna halved their predecessors’ rates the same week Anthropic launched Opus 5.5 at $20 per million output tokens.

The open-weight frontier spans several deployment classes, each with a different buyer in mind:

  • Full scale. DeepSeek’s V4-Pro, generally available August 13, 2026 under an MIT license at roughly 1.6 trillion parameters, with output pricing around $0.87 per million tokens; DeepSeek also published MIT weights for V4-Flash-Vision-Exp at the end of August, moving that multimodal checkpoint off API-only access, and retired the V4 Flash tier on September 10 in favor of V4.1 Flash; Moonshot AI’s Kimi K3 at 2.8 trillion parameters — the largest open-weight model released to date — with output pricing at $15 per million tokens against $3 per million input; Alibaba’s 2.4-trillion-parameter Qwen 3.8 Max (weights released August 12, 2026); Z.ai’s MIT-licensed GLM line including GLM-5.3, a 750-billion-parameter model that Z.ai reported surpassing Kimi K3 on many agentic benchmarks at release, with evaluations run inside Claude Code.
  • Mid scale. NVIDIA’s Nemotron 3 Ultra at 550 billion parameters with training data and recipes released under the NVIDIA Open Model License; Mistral’s Medium 3.5; Baidu’s open-weight ERNIE distillations.
  • Small and on-device. Google’s Gemma 4 (Apache 2.0); Microsoft’s Phi-4 family (MIT); Mistral’s Small 4 (Apache 2.0); OpenAI’s gpt-oss at 120 billion and 20 billion parameters (Apache 2.0 — still OpenAI’s most recent open-weight release); Meta’s Muse Glimmer 30 billion parameters (Apache 2.0); Nous Research’s Hermes 4.3 at 36 billion parameters, trained on the Psyche decentralized network and MIT-licensed.14

None of these are toys or research curiosities: the open-weight models now sit within a few months of the closed frontier on most agentic benchmarks.15 The differences the orchestrator cares about are still about deployment rather than raw capability: where the weights live (Hugging Face, ModelScope, the vendor’s own CDN), what license governs commercial use, what hardware the model actually fits on, and whether the vendor publishes the recipes needed to fine-tune or distill on top of it.

The shift shows up in what teams run by default. Many that paid frontier prices through 2025 have moved their default model to an open-weight one — GLM-5.3, Qwen, DeepSeek — because proprietary frontier pricing stayed high and the open models became too good to ignore. The frontier providers have answered the way providers under price pressure always do: cutting list prices, in some cases by half, and running promotions on top. GPT-6 Luna is the clearest response — after OpenAI’s July 2026 cut took GPT-5.6 Luna down 80 percent, the September 22 release halved it again to $0.10 per million input tokens and $0.50 per million output,16 which puts a frontier model in the same price range as some of the open-weight alternatives and makes it competitive on functionality and price at the same time. It is an interesting moment to be choosing a default model: the open weights have never been more capable, and the proprietary tiers have never been cheaper.

1.3.3 The economics of inference

Not every component in this survey gets its own economics discussion — inference does, because the models an orchestrated system runs are one of the biggest costs of the whole arrangement, and there will rarely be just one: most real systems call several models, each priced differently. Choosing them well is one of the most consequential decisions an orchestrator makes, and it moves cost and quality at the same time.

The economics reach deeper than model choice, and they change system design. The frontier list rates, per million tokens: GPT-6 Astra and Claude Fable 5.1 both at $10 input and $50 output; the tier just below — GPT-6 Sol, Opus 5.5, Sonnet 5.5 — running $2 to $4 input and $10 to $20 output; Gemini 3.1 Pro at $2 and $12; Gemini 3.8 Flash at an introductory $0.75 and $3.75. Claude Mythos 5.1 sits outside the public price sheet — it is sold only through trusted-access programs.17 Frontier open-weight output runs from under a dollar to about $15, with DeepSeek’s V4.1 Flash at $0.60 per million output tokens off-peak and GLM-5.3 at $4.40. That is close to an order of magnitude at the median and close to two at the cheap end, and the gap is the reason the always-on agent category exists at all. Say a component has an hourly heartbeat, and each heartbeat is an inference call that re-reads a large context, on the order of 300,000 input tokens with a short output. On DeepSeek V4.1 Flash off-peak that call costs about five cents; on Fable 5.1 at list rates, about three dollars. Run it 24 times a day and you are looking at about a dollar a day versus about seventy-five dollars a day, for the same job.18 In the context of a large business, seventy-five dollars a day might not seem like much — but if you have a fleet of a thousand agents running simultaneously, it is the difference between about a thousand dollars a day and about seventy-five thousand. These are the sorts of arithmetic an orchestrator does when designing a system.

Price is also not one dial. The bill depends on the reasoning effort (the same model at low effort answers cheaply; at high or ultra effort it burns thinking tokens that can dominate the cost of the call), on the processing path (batch and flex tiers run at roughly half the realtime price), on caching (cached reads are discounted as much as 90 percent), and on context length (long-context runs re-read the same material and pay for it every time). Two systems on the same model and the same workload can land an order of magnitude apart on these dials alone. Model choice sets the price range; effort, batching, and caching decide where in that range the bill actually lands.

Pricing changes move adoption, and the clearest window on that behavior is OpenRouter, the routing service that puts hundreds of models behind one API and publishes what its users actually run. Its token rankings are topped not by the most capable models but by the cheap ones, and the rankings page itself carries the caveat that this measures adoption, not quality. Users move fast. When DeepSeek released the V4 family, the company’s share of OpenRouter token flow went from under ten percent at the start of the year to roughly twenty percent by June, making it the top model on the service.19 Providers respond, and pricing has become promotional, with discounts clustering around releases and competitive moments, when a provider wants trial volume. The practical rule: budget on list price, treat a promotion as a free trial, and expect the discount to vanish when the next release needs the same attention.

Take a modest production workload — an internal documentation agent answering employee questions, moving thirty million input tokens and three million output tokens a month. Run everything on DeepSeek V4.1 Flash off-peak ($0.15 per million input, $0.60 per million output) and the month costs about six dollars. Run everything on Claude Fable 5.1 at its $10 per million input and $50 per million output list rates, and the same month costs four hundred and fifty dollars before the cheaper cache reads.20 That is a seventy-five-fold difference on identical work.

The capability gap is real, and stating it fairly is part of the job. Fable 5 was reported at 80.3 percent on SWE-Bench Pro against 58.6 for GPT-5.5; the line is built for the long-horizon, multi-step work where cheap models fall apart. DeepSeek V4-Flash is a fast tier tuned for routine traffic — routing, classification, summarization, heartbeat checks — and it will lose on hard reasoning. So the seventy-five-fold premium buys genuine capability. The orchestrator’s question is what fraction of the workload needs it. If five percent of the documentation questions genuinely require frontier reasoning, route those to Fable 5.1 and the rest to V4.1 Flash: the month costs about twenty-seven dollars instead of four hundred and fifty — frontier capability where it matters, a 94 percent reduction everywhere else.

At that price gap, work that would be prohibitively expensive on a frontier model becomes realistic on a small or open-weight one. Fan-out that runs ten attempts on a cheap model and keeps the best one stops being expensive. An evaluator model that critiques every output of a stronger model becomes a standard pattern instead of a luxury. A routing layer that sends easy work to a small local model and hard work to the frontier becomes the default architecture, not an optimization. Cache-hit pricing on the open-weight side drops input costs by another order of magnitude, which is why long-context agent systems — which re-read the same documents many times — increasingly run on cheap open-weight models with caches warm.

None of this makes a model swap free. A system wired to one provider’s tool-calling quirks and prompt formats does not move by changing a model string. Sometimes a swap works with no visible degradation, but more often than not, when you switch between two high-priced frontier models, you learn that the memory, the skills, and the context need to be tuned in subtle and nuanced ways to adapt to the new model. The harnesses and orchestration systems surveyed in this chapter make the swap feasible, and reusable evaluation sets tell you whether the new model passes your actual work before you commit.

Every model has a personality. Two inference engines might sit next to each other on the same benchmark leaderboard, but use them at their limits and you find each one is better at some things than others, in ways the leaderboard does not show. An orchestrator running several models at once ends up in an odd role: you turn into a kind of psychiatrist for a new model. One of the first things you have to do is figure out its hangups — what it is good at, what it refuses, what it overcomplicates, where it quietly drifts. This is a significant amount of effort, and it is part of the real cost of a model swap.

And cost awareness comes with the role. For most agent systems, model spend is the largest line item on the bill by a long shot — ahead of hosting, storage, and tooling — and the orchestrator is the person positioned to see it whole: the application team sees per-call costs, finance sees an invoice, and the person who owns the routing, the effort dials, and the evaluation results sees all three at once.

1.3.4 Workload types

Not every part of an orchestrated system needs the strongest model available, and knowing which parts need what is one of the ways the bill stays sane. The work that inference engines do inside one of these systems falls into a handful of recognizably different kinds:

  • Routine calls. Classification, extraction, summarization, formatting, the hourly heartbeat that checks whether anything changed. The output is short, the stakes are low, and a small or open-weight model does the job. This is also most of the call volume, which is why routing it cheaply matters more than any single clever optimization.
  • Tool use and function calling. The model has to emit a correctly formatted call to the right tool with the right arguments, every time, and not invent calls that do not exist. This is a reliability test more than an intelligence test, and some cheap models are excellent at it while some expensive ones are erratic.
  • Coding and repository work. Understanding a codebase, planning a multi-file change, writing the diff, running the tests. This is the workload SWE-bench family benchmarks measure, and it is where the capability gap between frontier and cheap models is widest.
  • Reasoning-heavy analysis. Long chains of inference over evidence — technical due diligence, incident analysis, cross-document synthesis. The general-reasoning benchmarks measure this, and it is the workload where overdelegation without a strong model produces confident nonsense.
  • High-frequency agent loops. Dozens or hundreds of small calls per hour, every hour: the always-on agents. The per-call capability bar is modest; the cost bar is everything.
  • Specialist domains. Medicine, law, finance, science. Domain benchmarks exist, and the choice here is often dictated less by benchmark scores than by jurisdiction, licensing, and regulatory rules about which models are allowed to touch which data.

The job is matching the engine to the workload, and the two levers are cost and quality: a cheap model on a hard workload fails the workload, and an expensive model on a trivial workload fails the budget.

1.3.5 Benchmarks

Benchmarks matter to orchestration for one reason: they are the only shared vocabulary for talking about capability across models and vendors. When a vendor says a model is “smarter,” the orchestrator’s next question is on which benchmark, against whom, at what price — and the answer to that question is the difference between routing decisions that hold up and routing decisions that collapse the first time the work gets hard.

The benchmarks an orchestrator wants to know are the ones that match the workloads above:

  • SWE-bench (Verified, and the newer Pro). Real GitHub issues in real repositories, multi-file diffs, tests must pass. The standard measure of coding ability.
  • GPQA Diamond. Graduate-level, Google-proof multiple-choice questions in physics, chemistry, and biology — the standard measure of hard general reasoning.
  • BFCL (Berkeley Function-Calling Leaderboard). Graded function calls: right function, right arguments, right types, no invented calls. The measure of tool-use reliability, and one where some cheap models beat expensive ones.
  • tau-bench. A model plays a support agent with a tool panel and a demanding user, scored on whether the task actually completes. The measure of multi-turn tool use under friction.
  • Terminal-Bench. Real tasks in a real terminal: install, configure, fix, build. The measure of agentic computer use, and the hardest one for everyone.
  • HLE (Humanity’s Last Exam). Expert-written questions at the frontier of every field; the go-to example of a benchmark so hard that every model scores embarrassingly low. Useful as a ceiling check, not a selector.
Figure 15. Where the models sit, August 2026. SWE-bench Pro scores for a selection of models, from the published leaderboards.

None of these leaderboards is static — model releases and score updates land weekly, and a figure like this one is stale within months of publication. The durability is not in the numbers. It is in the pattern they keep showing: the benchmarks that match your workloads are the ones worth tracking, and the models that lead them are worth watching, even when you are routing most of your volume to something cheaper.

1.3.6 Reading a model as a component

The properties an orchestrator reads off a model are the same ones an architect reads off any other component.

Property What to ask Why it matters for orchestration
Context window How many tokens fit in one call? Determines whether memory management is a convenience or a survival skill
Tool-use behavior Does it call tools reliably, and can it be constrained to a schema? The difference between a component and a liability in a multi-step run
Deployment target Hosted API, private cloud, on-device? Sets the data-sensitivity ceiling and the exit options
Cost per million tokens Input and output, with and without caching? At a 10x spread, architecture becomes an economic decision
License Open weights, API-only, restricted program? Determines inspectability, self-hosting, and how long the choice stays yours
Multimodality Does it take images, audio, PDFs natively? Decides whether preprocessing is a separate system or a feature
Sovereignty Which jurisdiction does it run in, and are you allowed to use it where you operate? Organizations and governments increasingly have rules that dictate which models may touch which data, and the list of compliant options can be short

None of the rows is about which model is smartest. All of them are about what the engine can be connected to — which is the whole job.

There are likely another 20 or 30 factors that could go into the choice — some are legal restrictions, some are emerging regulatory concerns, some are vendor-specific gotchas that only surface in production. The space is incredibly active, the factors multiply every quarter, and the practical move is to be clear about your own constraints (jurisdiction, budget, data sensitivity, exit options) and to re-check the market when those constraints change.

1.3.7 Where models show up for an orchestrator

An orchestrator encounters models through harnesses rather than raw APIs. Every runtime surveyed later in this chapter — Cursor, Claude Code, OpenClaw, Hermes Agent, Goose, Pi — puts a model picker in front of the user, and most now accept several providers plus local models. The consequence is that model choice is a standing configuration decision inside the orchestrator’s system, not an annual procurement decision. Chapter 7 returns to this at the level of architecture — deciding which step needs which caliber of judgment is the same decision as deciding which step needs a model at all.


  1. Vendor lineups as of September 29, 2026: OpenAI, “GPT-6 Astra,” https://openai.com/index/gpt-6-astra/, “Introducing GPT-6 Sol and Luna,” September 22, https://openai.com/index/introducing-gpt-6-sol-and-luna, and model documentation, https://developers.openai.com/api/docs/models; Anthropic, “Claude Fable 5.1 and Claude Mythos 5.1,” https://www.anthropic.com/claude-fable-and-mythos-5-1, “Introducing Claude Opus 5.5,” September 22, https://www.anthropic.com/claude-opus-5-5, and “Introducing Claude Sonnet 5.5,” September 28, https://www.anthropic.com/claude-sonnet-5-5 (Haiku 5.5 announced, not yet released); Google, “Gemini 3.8 Flash and 3.8 Flash Cyber,” https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/, and API changelog, https://ai.google.dev/gemini-api/docs/changelog; xAI, “Introducing Grok 4.7,” September 21, https://x.ai/news/grok-4-7. Capabilities and positioning are vendor claims.↩︎

  2. Proprietary field as of August 29, 2026: Microsoft’s MAI family and vendor-reported benchmark, https://blogs.microsoft.com/blog/2026/06/02/microsoft-build-2026-be-yourself-at-work/; Amazon Nova and Nova 2, https://aws.amazon.com/nova and https://docs.aws.amazon.com/nova/latest/nova2-userguide/what-is-nova-2.html; Baidu ERNIE, https://ernie.baidu.com/blog; ByteDance Seed 2.0 / Doubao, https://www.digitalapplied.com/blog/bytedance-seed-2-doubao-ai-model-benchmarks-guide; and Alibaba’s proprietary Qwen tiers. Product counts, rankings, and positioning are vendor-reported snapshots.↩︎

  3. Reuters, “AI data-centre race builds $1 trillion lease burden for Big Tech,” August 4, 2026, https://www.reuters.com/business/retail-consumer/ai-data-centre-race-builds-1-trillion-lease-burden-big-tech-2026-08-04 (five companies — Microsoft, Meta, Oracle, Amazon, Alphabet — with about $1.09 trillion in future payments under leases that have not yet begun, mostly for data centers, with runs of 15 to 19 years; Microsoft $329.1 billion, Meta $278.99 billion plus $68 billion of further leases signed in July, Oracle $260 billion, Amazon $137.21 billion, Alphabet $85.2 billion). 2026 capital-spending commitments: CNBC, “Tech AI spending approaches $700 billion in 2026,” February 6, 2026, https://www.cnbc.com/2026/02/06/google-microsoft-meta-amazon-ai-cash.html; Data Center Frontier, “Hyperscalers Plan $630 Billion in 2026 CapEx,” https://datacenterrichness.substack.com/p/hyperscalers-plan-630-billion-in (Big Four 2026 capex of up to $630 billion, up about 62 percent from $388 billion in 2025).↩︎

  4. OpenClaw, https://docs.openclaw.ai/ and https://openclaw.ai (verified August 29, 2026). Open-source MIT-licensed local-first AI agent runtime created by Peter Steinberger; reached 346,000+ GitHub stars within five months of release; the OpenClaw Foundation became a non-profit in June 2026. The always-on agent cost discussion is documented in the OpenClaw community’s own guides and the cost surveys cited by the project, including the OpenClaw pricing guide at https://openclawai.io/blog (DeepSeek V3 at ~$0.27 per million input tokens used as the default routing target; Claude Haiku 4.5 cited as a budget daily-driver at ~$6 per million combined tokens) and https://www.glbgpt.com/hub/openclaw-cost-pricing-guide (24/7 OpenClaw deployments $15–$300+ per month depending on model choice). The pressure-on-pricing claim is this book’s synthesis, not a single vendor’s statement.↩︎

  5. Open-weight landscape as of August 28, 2026: Z.ai GLM, https://z.ai/blog/glm-5.2; DeepSeek, https://api-docs.deepseek.com/updates; Qwen, https://qwen.ai/blog?id=qwen3.8; Google Gemma, https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4; Microsoft Phi, https://huggingface.co/microsoft/phi-4; Mistral, https://docs.mistral.ai; Thinking Machines Inkling, https://thinkingmachines.ai/news/introducing-inkling; and OpenAI gpt-oss, https://openai.com/index/gpt-oss-model-card. Licenses and rankings vary by model and change quickly.↩︎

  6. List rates per million input/output tokens as of September 29, 2026: GPT-6 Astra $10/$50, https://openai.com/index/gpt-6-astra/; GPT-6 Sol $2/$10 and GPT-6 Luna $0.10/$0.50 (September 22), https://openai.com/index/introducing-gpt-6-sol-and-luna; Claude Fable 5.1 $10/$50, https://www.anthropic.com/claude-fable-and-mythos-5-1; Claude Opus 5.5 $4/$20 (September 22), https://www.anthropic.com/claude-opus-5-5; Claude Sonnet 5.5 $2/$10 (September 28), https://www.anthropic.com/claude-sonnet-5-5; Gemini 3.1 Pro $2/$12 below 200,000 tokens and $4/$18 above, https://ai.google.dev/gemini-api/docs/pricing; Gemini 3.8 Flash $0.75/$3.75 introductory through December 31, 2026, https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/. DeepSeek V4.1 Flash $0.15/$0.60 off-peak ($0.30/$1.20 peak) after the September 10 retirement of V4 Flash. Open-weight rates ranged from under $1 to roughly $15 output (Kimi K3). Prices are moving snapshots; the order-of-magnitude spread is the durable claim.↩︎

  7. Reasoning-effort pricing: OpenAI, “Advancing the price-performance frontier with GPT-5.6,” July 30, 2026, https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6 — Luna down 80 percent and Terra down 20 percent from the July 9 GA prices (Sol unchanged at $5/$30, Terra to $2/$12, Luna to $0.20/$1.20 per million tokens; all three superseded on September 22 by the GPT-6 Sol and Luna releases at $2/$10 and $0.10/$0.50, https://openai.com/index/introducing-gpt-6-sol-and-luna); batch/flex tiers at half the realtime rate; cache reads discounted 90 percent (Luna cached reads at $0.02 per million). Tier pricing and the four-parallel-sub-agent “ultra” reasoning effort per OpenAI’s GPT-5.6 announcement and developer documentation, https://openai.com/index/gpt-5-6. Effort and processing-path economics also per vendor pricing pages and OpenRouter model pages (https://openrouter.ai/openai/gpt-5.6-terra), verified August 29, 2026.↩︎

  8. Anthropic gives Fable 5.1 a public $10/$50 per-million-token list price but sells Mythos 5.1 only through trusted-access programs, https://www.anthropic.com/claude-fable-and-mythos-5-1. Finout reports higher effective Mythos rates, https://www.finout.io/blog/claude-fable-5-mythos-5-pricing-benchmarks. Because Mythos has no public price sheet, those figures cannot be independently checked.↩︎

  9. Heartbeat arithmetic at September 2026 list rates. A call of 300,000 input tokens and 2,000 output tokens costs about $0.05 on DeepSeek V4.1 Flash off-peak ($0.15/$0.60 per million) and $3.10 on Claude Fable 5.1 ($10/$50). Twenty-four runs a day is about $1.10 versus $74; a thousand-agent fleet is about $1,100 versus $74,000 a day. Peak-hours billing doubles the DeepSeek side; caching and batch discounts reduce both totals. The point is the spread, not the exact bill.↩︎

  10. OpenRouter rankings, https://openrouter.ai/rankings, and “DeepSeek V4 Is Earning Agentic Token Share,” https://openrouter.ai/blog/insights/deepseek-v4-adoption. Rankings measure adoption within OpenRouter traffic, not quality or the whole market. The broader usage study is “State of AI: An Empirical 100 Trillion Token Study with OpenRouter,” arXiv:2601.10088, https://arxiv.org/abs/2601.10088. Discount listings, https://openrouter.ai/collections/discounted-models, are temporary promotions rather than durable prices.↩︎

  11. September 2026 list rates: Claude Fable 5.1 at $10/$50 per million input/output tokens, https://www.anthropic.com/claude-fable-and-mythos-5-1; DeepSeek V4.1 Flash off-peak at $0.15/$0.60 (the V4 Flash tier was retired September 10, 2026, and its successor rates are used here; V4 Flash’s pre-August 2026 rate of $0.14/$0.28 was the basis of earlier drafts). At those rates, 30 million input plus 3 million output tokens cost about $450 versus $6; routing 95 percent to DeepSeek and 5 percent to Fable costs about $27. Arithmetic by this book, before caching or batch discounts.↩︎