llm-providers-wiki
Synthesis — LLM Providers
The evolving thesis. Spun out of the hub _inbox llm-providers cluster on 2026-06-01, seeded
by deepseek-api-docs and grown with three landscape sources the router went and found
(open-weight roundup, API-pricing comparison, the llm-stats leaderboard).
Current thesis
The 2026 model market is defined by a collapsing price floor meeting a still-premium frontier, with open weights as the force pushing the two apart.
- A ~600× cost spread (llm-api-pricing): budget APIs at ~$0.10/1M input vs frontier reasoning at $30/$180. deepseek disrupted the bottom — capable open weights plus an OpenAI-compatible ultra-cheap API — dragging the floor down and forcing proprietary labs to justify their premium.
- Open-weight has caught up enough for production (open-weight-models): Llama 4, Qwen3, gemma-4, DeepSeek, Kimi, GLM. Two structural enablers — sparse MoE (frontier capability at practical inference cost) and Apache-2.0/MIT licensing clarity — plus a context explosion (256K → 1M → 10M tokens), and now multimodality at the local tier: gemma-4 12B runs encoder-free vision + native audio on 16GB (gemma-4-12b-announcement), making memory footprint a competitive axis alongside capability/cost/context. quantization is the second footprint lever (after sparse MoE): Gemma 4’s QAT checkpoints push the floor to a sub-1GB E2B and a 2-bit mobile scheme, claimed above PTQ quality gemma-4-qat — footprint is now contested by training technique, not just architecture. ternary-weights takes that one step further (2026-08-02): neutrino-1 claims an 8B-class model at 2.56 GB download / 3.88 GB on disk with weights trained in a ternary format rather than quantized into one, so there is no float model anywhere in its lineage. If it holds, footprint stops being a treatment applied to a model and becomes a property of it — but every number supporting it is first-party and the method is unpublished.
- “Best model” is multi-axis (llm-benchmarks): composite scores over capability × speed × price × context. Proprietary labs (anthropic, OpenAI) still lead reasoning; open/Chinese labs (Qwen, DeepSeek) lead cost-per-quality; xAI leads context — and with Grok 4.5 now also contests cost-per-quality from inside the closed frontier (grok-4-5-price-vs-benchmarks). A sharper cost axis emerges from that source: cost-per-task = per-token price × tokens-per-task, so a token-efficient model can be cheaper in practice even at a higher sticker rate (and vice-versa — Fable 5 is both pricey per token, $10/$50, and no cheaper per task).
The unifying tension: capability still concentrates at a few proprietary frontier labs, while cost and access are being democratized from below by open weights. Whether the frontier premium holds depends on whether open-weight reasoning closes the gap — the live question of the wiki.
Open weights have a third face (2026-07-26): the behaviour is editable too. The wiki has read open weights as licensing (who may run it) and footprint (who can afford to). abliteration adds a lever the license grants without naming: gemma-4-31b-it-scotoma is Gemma 4 31B-it with its refusal behaviour surgically weakened in layers 7–41 and republished under Google’s own Apache-2.0. Alignment shipped as weights is alignment shipped as something a third party can rewrite in an afternoon and redistribute on the same terms. The contrast with the closed tier is exact: a Claude refusal arrives as an HTTP 200 you can route around but never delete (claude-refusals-and-fallback). And the instrument doing the cutting is Anthropic’s own Jacobian lens — an interpretability tool built for auditing models, pointed at modifying one.
The provider map now sorts into five positions plus a reseller layer. The market is no longer a simple open-vs-closed binary; the players cluster by strategy:
- Consumer frontier (closed): openai (GPT/o-series; the API shape everyone else clones) and anthropic hold the reasoning lead and the $30/$180 premium. xai-grok sits here too — proprietary, reasoning-focused — distinctive for a 2M-token context (the “xAI leads context” claim) and an open→closed arc (Grok-1 was Apache-2.0; everything since closed), the inverse of the open-first labs. But Grok 4.5 (mid-2026) breaks the premium framing (grok-4-5-price-vs-benchmarks): at $2/$6 it prices below the whole field and runs the China playbook — good-enough benchmarks, win on cost-per-task (xAI claims 4.2× fewer tokens than Opus 4.8). A closed US frontier lab adopting the price-floor strategy is new, and it blurs the “premium frontier vs cheap challengers” split the thesis rests on.
- Open-weight challengers: qwen (Alibaba) is the strongest evidence open weights are closing from below — the deepest open model ladder (0.6–32B dense + MoE; reasoning/VL/audio/Coder/omni), mostly Apache-2.0, 200k+ Hugging Face derivatives, and open-first (unlike dual-track google) — alongside deepseek, Meta’s llama (the family that catalyzed the wave), and Z.ai, whose GLM-5.2 (MIT, 753B-A40B MoE, text-only) became the #1 open-weight model on artificial-analysis in June 2026 — the lead in the open field is no longer DeepSeek/Qwen alone. InternScience‘s Agents-A1 (late June 2026, Apache-2.0, 35B MoE on Qwen3 lineage) joins as an agentic challenger claiming trillion-parameter-class agentic performance by training technique, not scale (self-reported).
- US open-weight, customization-as-product (new, 2026-07): Thinking Machines Lab shipped Inkling (Apache-2.0, 975B-A41B multimodal MoE, 45T tokens, 1M ctx) while saying outright it is “not the strongest overall model available today, open or closed” inkling-announcement. That is a position no one else on this map holds. It doesn’t compete on the capability × price × context board at all: the weights are free, the pitch is a base you fine-tune, and the money is one layer up on Tinker, its fine-tuning platform. Where google gives away gemma-4 to seed mindshare for a paid gemini frontier, Thinking Machines has no frontier tier to seed — adaptation is the product. It also fills a geographic hole: per Willison inkling-willison-review, the US open-weight bench is now Inkling + NVIDIA Nemotron + Gemma 4, and Inkling is the only one from a company whose open weights are the strategy rather than a side-bet. Note the seam with mistral-ai‘s Forge below — two labs converging on training-as-a-service over their own open weights from opposite directions.
- European open-weight → sovereign platform: mistral-ai (Apache-2.0 Mixtral MoE) — the open axis is no longer China-only. Its mid-2026 site (mistral-platform) shows the position evolving up-stack: from a model lab to a sovereign, EU-hosted full-stack platform (agents Vibe / Vibe for Code, Studio, custom-training Forge, dedicated Compute; self-hosted / EU-cloud / partner deployment). That collapses the gap to the enterprise/sovereign row below — Mistral reaches the same data-residency play as cohere but from the open-weight side, so “European open-weight” and “sovereign enterprise” are becoming one position, not two.
- Enterprise / sovereign: cohere competes not on consumer reach but on data residency and compliance (private cloud, on-premise, regional EU/APAC/UK) for regulated industries — the founders’ Transformer-paper roots (Aidan Gomez) behind an early “same architecture, different deployment model” bet. Crucially its North Mini Code (9 Jun 2026, the first North-family model) is Apache-2.0 open-weight (30B MoE / 3B active, runs on a single H100): Cohere pursues sovereignty through open weights and local deployment, so it sits on the open-weight axis too, not opposite it — the sovereign angle is go-to-market, not licensing. It scores 33.4 on the independent artificial-analysis Coding Index, opening a developer-coding play on top of the enterprise base.
- Dual-track: google ships both a closed frontier (gemini) and permissive open weights (gemma-4) — see Recurring reads.
Llama also sharpens the licensing thread: open-weight but not OSI-open (custom community license, OSI-disputed), against the Apache-2.0 cleanliness of gemma-4/mistral-ai/qwen — “open” is a spectrum, not a binary. Across the top sits a cloud-reseller layer: amazon-bedrock is not a model maker but a one-API aggregator of Claude/Llama/Mistral/Cohere/Nova — the demand-side mirror of the map, where model choice becomes a config parameter and differentiation moves up to routing/caching/data. Through a reseller, open and closed models are the same API call; the buyer meets the whole price/openness spectrum through one pane.
That layer has a second, self-hosted shape: the ai-gateway you run yourself. omniroute (MIT, ~30K stars) fronts a claimed 290+ providers from a process on your laptop, cascading from subscription seats down through cheap providers to free ones as each quota dies, and it pushes the demand-side logic past routing into two levers the reseller never offered — client-side prompt compression (12 stacked engines, 15–95% claimed) and free-tier aggregation (~1.53B tokens/month pooled across 43 provider free tiers). Both are unilateral: neither needs the provider’s agreement, and the second arguably runs against its interest, which is why the project itself flags 15 of those pools as ToS-questionable. The reseller commoditized model choice; the local router commoditizes the provider’s generosity.
Recurring reads
- Compatibility as strategy — challengers (deepseek) conform to OpenAI-compatible (and now Anthropic-compatible) API shapes to erase switching costs; the incumbent API is becoming a de-facto standard.
- Cost is dominated by output tokens (2–6× input) and slashed by caching/batch/routing — so
engineering, not just model choice, sets real cost. The caching lever cuts twice over:
Anthropic’s cache-aware ITPM (claude-api-rate-limits) means cached input tokens don’t
count toward the rate limit, so prompt caching raises effective throughput (a 2M ITPM limit at
an 80% hit rate processes ~10M tokens/min) on top of the ~90%-off read price — the engineering
lever now buys headroom, not just discount. Model routing is now a first-class
API primitive, not just app glue (claude-refusals-and-fallback): Anthropic ships
server-side
fallbacksthat retry a safety-declined request down a model chain (fable-5 → opus-4.8) inside one call, with per-attempt billing (usage.iterations[]), afallback-creditbeta to avoid double-paying prompt cache, and sticky routing that pins a conversation to whoever accepted. The “engineering sets real cost” thread hardens into the provider’s own surface — and note refusals are an HTTP 200, invisible to error-rate monitoring. - Dual-track labs — google ships both a closed frontier (gemini) and a permissive open-weight family (gemma-4), unlike the open-only Chinese labs or closed-only frontier labs. Open weights double as developer-mindshare seeding, not just a product. The May 2026 roundup google-ai-updates-may-2026 adds a wrinkle: both tracks are now pitched at the same use-case — agents and coding (Gemini 3.5 “frontier intelligence for agents and coding”; Gemma 4 12B “agentic workflows”). The open/closed split is about licensing and footprint, not target workload — a lab can run one agent-and-coding pitch across the whole price/openness spectrum. Q2 2026 earnings (alphabet-q2-2026-earnings) add the usage-scale dimension both prior roundups lacked. The story stops being only which model and becomes how much volume: ~22B tokens/minute across Google’s APIs (up ~38% Q/Q from ~16B), a 950M-MAU Gemini app, 9M+ developers, and ~900M Gemma downloads — the dual track measured on both rails at once (closed tokens + open downloads). The model side widened too: a full cost-tier ladder (Gemini 3.6 Flash / 3.5 Flash-Lite) plus the first domain-specialized Gemini (3.5 Flash Cyber, “comparable to larger models”) and Gemini 4 in pre-training — so the provider is now differentiating by tier and vertical, not just a single frontier number. All company-reported (T1 provenance, unaudited here); the durable read is that the market’s growth axis is now consumption volume, the demand-side complement to the output-token cost thread. Cross-spoke: Cloud +82% / $514B backlog / TPU 8t-8i / Axion / the million-accelerator Virgo Network (cloud-wiki), AI Mode past 1B MAU merged with AI Overviews (search-marketing-wiki), and ADK ~70M downloads / Gemini Enterprise in ~90% of the Fortune 100 (agentic-tooling-wiki) are logged there.
Generative-media sub-theme (watch for spin-out)
The gemini line now has sourced generative-media tiers with their own pricing units, not just $/token: Nano Banana 2 Lite (text-to-image, $0.034/1k images, ~4 s) and Gemini Omni Flash (video, $0.10/sec, Veo-3.1-Fast parity, 10 s max). This is a distinct modality from the text/multimodal LLMs at the spoke’s core — and a parallel to how audio-gen was carved out to speech-audio-wiki. For now it stays here as part of the Gemini family; if non-Google image/video-gen sources accumulate (≥3), a dedicated generative-media spoke is the spin-out candidate. Logged so the cluster is visible rather than buried.
Source 2 — and the first non-Google one (2026-07-22). mage-flow-turbo is Microsoft’s MIT-licensed 4B text-to-image model (rectified flow matching, 4-step distilled, native 512–2048px), claiming 0.59 s per 1024² image on an A100 and parity-or-better against Qwen-Image 20B / Z-Image 6B / FLUX.2 32B. It moves the sub-theme in two ways. It splits the modality along the same open-vs-closed axis the text market runs on — Google’s image tier is priced per image, this one is weights you download, so the generative-media cluster now has both a hosted-API instance and an open-weight one, and only the first has a price. And its argument is the spoke’s own recurring one, one modality over: 4B beating 20–32B is agents-a1‘s “not by scale” claim in image form, bought with tokenizer–backbone co-design and step distillation instead of agent horizon. Same T3 caveat as A1: first-party card, first-party comparison set, no independent re-run — and the corpus still has no artificial-analysis equivalent for image models, which is the missing referee for this whole sub-theme. Cluster now at 2 of 3.
Open questions
- What is a capability claim worth when the artifact is checkable but the model isn’t? (opened 2026-08-03)
Every model page in this wiki rests on numbers the vendor produced and nobody outside can re-run — the
standing complaint behind the benchmark tensions below. astra breaks the pattern in an unexpected
direction: OpenAI introduced its unreleased next model in a mathematics paper, and shipped
ten machine-checkable Lean proofs with it (
../research-wiki: ten-proofs, lean-certificate). A stranger can verify the mathematics on a laptop with no access to the model. What that stranger cannot verify is anything about Astra — not how many attempts the proofs took, not how much human steering went in, not that a shipped version behaves comparably. So the field now has a claim shape where the output is falsifiable and the capability still isn’t, which is better than a vendor eval table and is not the same as a benchmark. Watch whether other labs copy the form (verifiable artifacts alongside announcements) or whether it stays confined to domains with formal checkers — mathematics and program verification, and very little else this wiki tracks. The related tell: Astra’s one published quantity, ~$2,000 of tokens, is priced at Sol rates, so even the cost figure is denominated in a different model. - How long do the free tiers survive being pooled? (opened 2026-07-26) A free tier is a marketing cost the provider prices against expected casual use. omniroute and its class turn 43 of them into a single ~1.53B-token/month budget with automatic failover between them, which is a different product than the one each provider thought it was giving away. Either the pools shrink and get fingerprinted (per-key heuristics, gateway detection), or free access stops working as an acquisition channel at all. The project’s own “15 providers ToS-flagged so you decide” is the tell that this is already contested. No data yet on whether any named pool has tightened in response.
- Does compression become a provider feature or stay client-side? Prompt compression is currently something a gateway does to a provider — omniroute claims 15–95% fewer tokens billed. Providers already sell the cooperative version (prompt caching), which pays them; the adversarial version doesn’t. Watch whether the labs answer with cheaper caching or with terms against it.
- What is a safety claim worth on a downloadable model? (opened 2026-07-26) google releases Gemma 4 with a refusal posture; ReadyArt republishes it with that posture edited and the Apache-2.0 license intact. The lab’s alignment work survives exactly as long as nobody bothers to remove it, and the removal is a weight merge, not a retrain. Two things to watch: whether open-weight licenses start carrying behavioural terms (they can’t be enforced on weights already downloaded), and whether labs shift safety into things an editor can’t reach — pre-training data, or serving-side classifiers that only exist on the API. Note the one honest complication in the founding source: scotoma‘s own card says it “refuses basically as much as its base model,” so the edit may be less powerful than the framing. No evals either way — that’s the gap.
- Does the frontier premium survive? If open-weight reasoning (DeepSeek/Qwen) closes on
Anthropic/OpenAI, the $30/$180 tier loses its moat. Watch llm-benchmarks over time. Latest data
point (2026-06-18): GLM-5.2 tops the open-weight artificial-analysis index and sits
#2 to Claude Fable 5 on Code Arena WebDev at ~1/5 the price — the closest an open model has come to the
closed coding lead. The standing caveat: it’s token-hungry (~43K output tokens/task), so the cheap
per-token rate overstates the real-cost gap — the “output tokens dominate cost” thread cutting the other
way. Same caveat hits Sonnet 5 (2026-06-30): its prompting guide
notes a new tokenizer emitting ~30% more tokens for the same text, so the headline per-token price
understates real cost relative to 4.6 — token-accounting, not just rate, sets the bill. The effort
parameter cuts the other way: medium on Sonnet 5 ≈ high on Sonnet 4.6, so you can buy the old
intelligence tier at a lower effort/token setting. Net cost depends on which lever dominates per
workload. New angle (2026-06-30): the squeeze now also comes from inside the frontier labs.
Claude Sonnet 5 is pitched as performing close to Opus 4.8
at the Sonnet price ($3/$15 standard, $2/$10 introductory), with safety hardening and an explicitly
agentic design. A lab compressing its own mid-tier toward the flagship erodes the $30/$180 premium
without any help from open weights — and feeds the agentic-tooling thesis that “Sonnet + harness beats
raw Opus.” Vendor-reported (HLE 43.2% / 57.4% w-tools, OSWorld-Verified 81.2%, SWE-bench Verified
85.2% per the system card); awaits an independent
artificial-analysis re-run. Caveat from the system card: Anthropic itself says Sonnet 5
“trails our Opus and Mythos-class models” on almost all benchmarks, so “close to Opus” is a
price-adjusted claim, not parity — the premium is squeezed, not erased. The card also discloses a
model class above Opus, “Mythos” (Claude Mythos 5), treated as Anthropic’s current capability
frontier — the proprietary ladder has a rung the public market hasn’t priced yet. New mechanism
(2026-06-30): Agents-A1 adds a third way the premium could erode — neither cheaper
tokens nor a flagship squeezing its own mid-tier, but training technique: a 35B open-weight MoE
claiming frontier-class agentic scores (GAIA 96.0, IFEval 94.8) by multi-teacher domain distillation
- long agentic horizons rather than scale. Self-reported (T3), so it awaits an artificial-analysis re-run; if it holds, “frontier agentic behavior at mid-size open weights” is the strongest version of the closing-from-below thread yet. Strongest data point yet, and it’s the lab’s own (2026-07-24): Opus 5 ships at $5/$25 — the exact price of Opus 4.8 — while Anthropic claims it more than doubles 4.8 on Frontier-Bench and lands within 0.5% of Fable 5 on CursorBench 3.2 (claude-opus-5-announcement). A generation of claimed capability at a flat rate is the intra-lab squeeze running a second time, one rung higher: after compressing the mid-tier toward the flagship with Sonnet 5, Anthropic now compresses the flagship toward its own premium tier and markets the compression as the product (“close to Fable 5 at half the price”). The premium isn’t being competed away from below so much as discounted from inside. Two corrections to earlier readings: the proprietary ladder’s top rung is not what the system card implied — Mythos 5 now reads as a domain specialist (ahead only on cybersecurity and biology research, “substantially behind” Opus 5 elsewhere), so the unpriced rung above Opus is narrower than assumed; and the pitch is stated almost entirely as cost-per-task ratios with no absolute scores, which is the cost-per-task frame becoming the vendor’s default unit of self-description. The strongest closing-from-below data point yet, with a catch in the licence (2026-07-28, kimi-k3). Moonshot’s K3 ships 2.8T/104B-active weights publicly and posts a split decision against the current flagships rather than last generation’s: ahead on GPQA Diamond (93.5 vs Opus 4.8’s 91.0), SWE-Marathon (42.0 vs Fable 5’s 35.0), BrowseComp and Video-MME; behind on DeepSWE (67.5 vs 70.0). Trading wins with the closed frontier — not approaching it — is a harder data point than glm-52 or agents-a1, and the benchmark table is unusually candid about its confounds (different harnesses per model, a modified SWE-Marathon branch, and a note that Fable 5 hit fallbacks on 35% of tasks in their run, which would have hurt the rival’s score). Still first-party; awaits artificial-analysis. The catch is that the terms tightened as the capability closed. K2.6 was Modified MIT; K3 ships a bespoke licence that is MIT-shaped except for a Model-as-a-Service revenue threshold (US$20M / 12 months) requiring a separate commercial agreement — structurally Llama’s move, gating the clouds and API resellers rather than users (open-weight-models). So the answer to this question is becoming two-sided: capability is being competed away from below faster than the openness is. The open-weight wedge stays sharp on price and access, while the leading open labs quietly reserve the serving business for themselves.
- Is caching efficiency an intra-lab cost axis? A T4 practitioner anecdote
(Njenga, 2026-07) times a single
"Hello"in Claude Code and finds Fable 5 costing $0.4745 vs ~3.2× less on Sonnet 5 — because Sonnet read more from cache (24.8k vs 15.6k tokens) and cache reads bill below fresh input. If it holds, two models from the same lab diverge on real cost through how well each caches a session, not just their sticker rates — the llm-api-pricing caching lever operating within a provider. But it’s one unrepeated run behind a paywall, and it cuts against Fable being the fast, coding-strong pick (Code Arena WebDev leader), so treat it as a flag, not a finding. Corroborated on the sticker side (2026-07-08): The Decoder’s Grok-4.5 comparison (grok-4-5-price-vs-benchmarks) lists Fable 5 at $10/$50 — ~3.3× Sonnet 5’s $3/$15, matching Njenga’s ~3.2×. So the gap is at least partly a list-price difference, not only caching — Fable 5 is genuinely the dearer model. That reframes the “fast/cheap” expectation: Fable 5 may be a premium creative/coding tier, not a budget one. Third read (2026-07-09): Hassid‘s buyer’s guide (dont-use-fable-5-hassid) again lists $10/$50, so three independent third-party sources now agree on the sticker — well-corroborated, though still wanting a T1 Anthropic sheet. He adds a usage-economics layer the rate alone misses: because the model re-reads the whole thread each turn, real cost scales with conversation length (his tests: ≈$0.15 short, ≈$6 at 19 turns, ≈$14 at 40) and Fable reportedly shifts to pay-per-use credits after 2026-07-12. His buyer’s ladder — Fable 5 only for hard goals at high/max effort (1–2 turns), Opus 4.8 as the cheaper-and-smarter everyday workhorse, Sonnet 5 not justified over Opus — is a practitioner inversion of Anthropic’s own mid-tier push (see the Sonnet-5 “close to Opus” thread above): a non-Anthropic voice arguing the flagship, not the mid-tier, is the value pick. Opinion (T3), but it sharpens the open “which lever dominates net cost” question toward turns/context length, not just per-token rate. Sticker now implied by a T1 source (2026-07-24): Anthropic prices Opus 5 at $5/$25 and calls it “half the price” of Fable 5 (claude-opus-5-announcement), which arithmetically puts Fable 5 at $10/$50 — the figure three third-party reads gave. Inference from a first-party claim rather than a published sheet, but the three T3/T4 reads are no longer alone. The caching half of this question is untouched by it and still rests on Njenga’s single run. Closed — the rate was published all along (2026-08-03). Anthropic’s 9 June launch announcement states “$10 / $50 per Mtok” directly; the hub holds it as research-wiki/claude-fable-5-mythos-5-announcement, created a month before this spoke’s Fable 5 page. So this thread’s four-source triangulation was corroborating a published figure, not substituting for one — and it landed on the number exactly, which is the stronger result. The correction is narrow: only the claim that no first-party rate existed was wrong. What this says about the corpus is the larger point — the answer sat in a sibling spoke for two months and this spoke never saw it. Entity dedup hasentity-indexto prevent exactly that; claims have no equivalent, so a cross-spoke check belongs in the ingest gate for any question a page marks as open. Context window, max output and intended tier remain genuinely undisclosed by Anthropic, and the caching half still rests on Njenga’s single run. - Is “open weight” still one category? (new, 2026-07-16) The thesis treats open weights as a single force pushing the price floor down, on the assumption that downloadable ⇒ self-hostable ⇒ cheap. Inkling strains that. It is Apache-2.0 and encoder-free-multimodal like gemma-4 — and it needs 2TB+ of aggregated VRAM in BF16, ~600GB at NVFP4 inkling-announcement, against Gemma 4 12B’s 16GB. Same licence, ~125× the hardware. If the open field is really two fields — local-tier weights (a footprint play, where quantization is the lever) and cluster-tier weights (a customization play, where you rent GPUs or the maker’s platform anyway) — then only the first actually attacks the price floor, and the second competes with the closed frontier on control and fine-tunability instead. Watch whether later releases sort into those two poles or fill the middle.
- Does giving the weights away and selling the tuning work? thinking-machines-lab is the clean test: no frontier tier, no inference margin to protect, revenue at Tinker. If it holds, the open-weight wedge stops being only a price attack on incumbents and becomes a business model of its own. One release deep, so this is a question, not a finding — there’s no pricing sheet, no independent benchmark re-run, and no track record.
- How stale is this? Pricing and rankings churn weekly; every page here is a dated snapshot.
Vendor/SEO bias — pricing-comparison and leaderboard sources have incentives; numbers are indicative. A neutral, reproducible benchmark would be the highest-value next source.Addressed (2026-06-12): artificial-analysis is the methodology-disclosed, continuously-re-run independent platform (Intelligence Index v4.0 = composite of GPQA Diamond / HLE / τ²-Bench / Terminal-Bench / SciCode; blended price at a 7:2:1 cache:input:output ratio; live TTFT) — the reproducible yardstick to track the “does the frontier premium survive?” question against. It also makes the host/reseller layer measurable (same weights across providers). Residual caveat: composite weighting + the 7:2:1 blend are disclosed editorial choices, and it’s still a churning snapshot. Watch item (2026-07-23, weak): Kimi K3 is claimed at 2.8T open-weight with public weights on 2026-07-27, and #1 editorial writing + edging Fable 5 on frontend coding — but the only source is a T4 creator-opinion essay (fable5-gpt56-kimi-k3-creators) with no scores. If it lands as described it’s another closing-from-below data point beside glm-52/agents-a1; until the 27th it’s an unverified anchor, not evidence.- Geopolitics/licensing — Chinese open-weight leaders (DeepSeek, Qwen, Kimi) dominate the open field; enterprise/regulatory acceptance is unmodeled here.
Growth edges
Ranked; each names the kind of source that would close it (see ../QUALITY.md → Growth edges).
- An independent re-run of any vendor benchmark claim. Every model page here rests on numbers the vendor produced and nobody outside can reproduce; it is the standing complaint behind the spoke’s benchmark tensions. — needs: a T1/T2 third-party evaluation of a model this wiki already pages.
- Whether the verifiable-artifact form spreads. astra shipped machine-checkable Lean proofs with a capability claim — falsifiable output, unfalsifiable capability. — needs: a second lab doing the same, or an analysis of the form.
- Free-tier economics. How long a pooled free tier survives is asked and unsourced. — needs: a first-party pricing/policy change with stated reasoning (T3 acceptable — it is a first-party fact).
Coverage edges (added 2026-08-08, at the curator’s request for a wider backlog). These widen what the spoke covers instead of answering an open question above; one ordinary solid source closes any.
- The terms nobody reads. Data retention, whether prompts train the model, and zero-retention tiers decide which provider an enterprise may legally use, and no page scores that dimension — llm-api-pricing prices tokens only. — needs: the providers’ own terms and enterprise addenda (T3, first-party facts), read side by side.
- Batch and off-peak pricing. Most providers sell asynchronous batch work at roughly half price, which changes the cost comparisons this spoke keeps making. — needs: first-party batch pricing and its stated turnaround guarantees.
- Embeddings and rerankers. A whole product line with its own pricing and leaderboards, absent from a spoke that otherwise tracks every model release. — needs: provider docs plus MTEB.
- Availability, not just capability. No page holds an outage record, an SLA, or a status-page history, so a model can be scored here without anyone asking whether it stays up. — needs: published status-page history or a postmortem.
The tier ladder doesn’t order everything
The spoke’s working model of Anthropic’s line is a price/capability ladder: claude-sonnet-5 mid, claude-opus-5 flagship-by-price, claude-fable-5 premium, Mythos a domain specialist. Every comparison held here — the benchmark ratios, the cost-per-task framing, the $10/$50-vs-$5/$25 split — sorts on that one axis.
Anthropic’s own documentation shows one dimension running the other way. Per anthropic-model-docs, Opus 5’s reliable knowledge cutoff is May 2026; Fable 5’s and Sonnet 5’s are January 2026. The tier sold as below the frontier carries four months more world knowledge than the tier above it.
That is not a contradiction — nothing claims Fable 5 is newer — but it is a hole in how this spoke has been reasoning. Recency is a real axis of model quality for anything touching current events, current prices, or recent releases, and no source in the corpus prices it. A buyer choosing Fable 5 for “the best model” gets a better reasoner with an older world. Watch whether the next Anthropic release keeps the flagship-lags-by-recency shape or closes it.
It also sharpens the standing caution about where these facts come from. The knowledge cutoff appears in the docs and in no launch post — same place the Opus 5 effort ladder turned out to be mis-stated (see anthropic-model-docs). Launch posts are the right source for positioning and the wrong one for anything a buyer or developer would act on; this spoke had been building almost entirely from them.
Contradictions / tensions
None internal yet. Cross-source tension to watch: leaderboards crown proprietary reasoning while the open-weight roundup argues practical fit beats rank — a framing disagreement, not a fact conflict. Fable-5-is-expensive, now on two reads (2026-07-08): Njenga’s T4 test put Fable 5 at ~3.2× Sonnet 5 on a trivial prompt; The Decoder’s rate card (grok-4-5-price-vs-benchmarks, T3) independently lists Fable 5 at $10/$50 ≈ 3.3× Sonnet’s $3/$15. Two non-Anthropic sources agree it’s the dear model — corroboration, not proof. Odd for a model pitched as the fast pick; likely a premium tier. Still pending a T1 Anthropic price sheet before recording as fact.
Efficiency is now the shared pitch — and tiering goes intra-family
openai-gpt56-grok45-clash shows OpenAI and xAI timing rival launches hours apart, both leading with token efficiency. Two structural reads land on the thesis. (1) Tiering moved inside the family: OpenAI’s GPT-5.6 ships as Sol / Terra / Luna (flagship / everyday / speed-&-affordability) — the capability×price split the market already had across labs is now explicit within one lineup, with Luna as the cheap-fast tier. (2) Cost-per-task is the industry frame, not one vendor’s spin: Musk pitches Grok 4.5 as “Opus-class but faster, more token-efficient and lower cost” — the same cost-per-task argument, in the maker’s own words, as buyers watch token spend. Gemini 3.5 Pro is expected the same month. So the price-floor pressure the thesis pins on open weights now has a second driver: efficiency competition among the closed frontier labs themselves.
A new availability axis — export controls. The same source reports the Trump administration export-banned Anthropic’s Fable and Mythos models (cyberattack-misuse risk); Anthropic suspended them and redeployed Fable 5 last week. For the first time in this spoke a model’s availability turned on government policy, not price or licensing — a governance variable sitting above the market map (cross-spoke: ai-governance-wiki).
A second pricing shape — the speed provider. The market’s price story so far is a per-token rate card (the ~600× spread, output 2–6× input, cost-per-task as the hidden multiplier). Cerebras breaks the shape: a wafer-scale inference provider (cerebras-inference) reselling open models on speed, it publishes no per-token price and sells daily-token subscription buckets instead (Cerebras Code Pro $50/mo→24M tok/day, Max $200/mo→120M tok/day; $5 free credit; Developer from $10). Both coding tiers were sold out (mid-2026) — capacity is rationed, so what’s scarce and priced here is throughput, not model quality. This adds an axis the rate-card view misses: a provider can compete on how fast and how much per day rather than how much per token. It also sharpens the boundary with llm-inference-wiki — Cerebras the mechanism/company lives there (cerebras-inference, cerebras-systems); what lives here is its place in the provider market and its unusual price structure.
A third shape — speed as a paid mode on one model (2026-07-24). Opus 5 sells a fast mode at 2× the base rate for ~2.5× the default speed (claude-opus-5-announcement). Cerebras prices speed by building different hardware; OpenAI’s Luna tier prices it by shipping a different model. Anthropic prices it as a runtime toggle on identical weights, which puts latency alongside effort as a per-request dial the buyer pays for directly. Three ways to sell speed now coexist — separate silicon, separate model, same model different meter — and only the third leaves the capability question untouched. Whether the 2× premium prices real serving cost or just willingness to pay is not answerable from the announcement.
The first independent capability read, and why it doesn’t settle much
Almost every capability claim in this spoke is the lab’s own. opus-5-arc-agi-3 is the exception: ARC Prize scores Opus 5 at 30.2% on ARC-AGI-3 against a 7.8% prior record (GPT-5.6 Sol Max), with “Fable-class” at ~20%, plus 90.4% / 97.5% on ARC-AGI-2 and -1. Three things it changes and one it doesn’t.
A benchmark nobody has saturated ranks differently from the composites. The indices in llm-benchmarks score what the vendor ships and crowd into a narrow band at the top; ARC-AGI-3 scores the model without a harness on tasks it has never seen, and the field is spread from 7.8% to 30.2%. It also inverts the family ordering Anthropic sells: Fable 5 is the premium tier at 2× the price, and here the cheaper flagship beats it.
A vendor ratio that doesn’t reconcile. The launch post claims “3× the next-best model on ARC-AGI 3” (claude-opus-5-announcement). Against the 7.8% record that’s ~3.9×; against ARC Prize’s ~20% Fable-class figure it’s ~1.5×. Both claims can’t describe the same comparison. Recorded as a tension, not resolved — this is the same shape as the Fable-5 pricing case, where third-party numbers arrived before the vendor’s own.
Novelty benchmarks decay once labs aim at them. Opus 5 postdates ARC-AGI-3’s public format, and the jump does not reproduce on Guanghan Ning’s private Witness benchmark, where Opus 5 (43.4) ties Fable 5 and kimi-k3 and falls below Opus 4.8 on the one unfamiliar rule combination. ARC Prize’s Greg Kamradt argues that isn’t disqualifying. Either way the useful read is structural: an evaluation whose value is novelty has a shelf life measured from the day its format goes public, so private and rotating benchmarks buy resistance to targeting at the cost of reproducibility — the axis the artificial-analysis disclosure story sits at the other end of.
What it doesn’t change: the open question below about the frontier premium. A 30.2% that may be partly benchmark-targeted, on a test the vendor’s own harness would score higher on, is not evidence about capability-per-dollar in production.
”Model” is becoming a product boundary, not a technical one
Sakana Fugu is a multi-agent system sold as a model: OpenAI-compatible endpoint, a model string, a reasoning-effort dial, and — in the docs’ own first sentence — an agent system underneath that routes across all supported providers by default, with the pool narrowed per API key rather than per request fugu-get-started. Three of this spoke’s working assumptions break at once.
You can’t say which model answered, so a capability benchmark measures a routing policy on the day it ran. You can’t state a per-token price for a named model, because the composition is the product. And the reseller/lab distinction collapses: Sakana presents as a lab with its own model family while being, underneath, a composer of other labs’ models — llm-provider‘s “frontier lab” and “cloud reseller” categories arriving at one endpoint.
Read it beside claude-opus-5‘s fast mode (a runtime toggle on identical weights, 2× price for ~2.5× speed) and the trend is the same from both directions: the SKU has stopped corresponding to a set of weights. One vendor sells two products off one model; another sells one product over many models. In both cases what the buyer purchases is a configuration, and the rate-card view this spoke was built on describes less of the market each month. The ai-gateway page carries the mechanism; the honest consequence for the map is that capability and price claims now need to name what they measured, not just which model string was passed.
Two smaller markers from the same source. Cyber is now its own SKU across three labs —
Anthropic’s Mythos (claude-opus-5), Google’s Gemini 3.5 Flash Cyber
(alphabet-q2-2026-earnings) and now fugu-cyber — with none of them publishing what makes a
model a cyber model, against a backdrop where Anthropic’s cyber-capable line drew an export ban
(claude-fable-5). Domain specialization has quietly become a product axis beside size, speed
and price, and the security one is the axis regulators already watch. And the effort dial starts at
high: no low, no medium, with a Codex integration that raises the stream idle timeout to two
hours. A model whose cheapest setting is “high” and whose turns run for hours is not competing in
the same market as a chat completion.
The architecture has stopped corresponding to a lab
neutrino-1 states an architecture matching qwen‘s published Qwen3-8B/config.json on every
field — 36 layers, 4,096 hidden, 32 query over 8 KV heads, 12,288 feed-forward, 151,936 vocabulary,
40,960 positions, untied embeddings, and a matching parameter count — and its release post never
mentions Qwen. The 0.6B matches Qwen3-0.6B the same way.
Nothing improper follows. Configurations are not proprietary, Qwen3 is Apache-2.0, and holding the geometry fixed while changing the weight format is the right experiment when the format is what you are testing. The point for this map is different: Qwen’s contribution to the open ecosystem is now a default shape as much as a set of checkpoints, and a shape leaves no attribution trail. The 200,000+ Hugging Face derivatives are countable; adopted geometries are not.
This is the erosion logged in “Model is becoming a product boundary”, one layer down. There, a model string stopped identifying a set of weights. Here, a set of weights stops identifying a design lineage. The spoke holds four positions on that gradient now: agents-a1 discloses its Qwen3 lineage, readyart publishes openly-labelled derivatives of someone else’s checkpoints, sakana-fugu presents other labs’ models as a family of its own, and neutrino-1 sits at a fourth point — plausibly original weights on a borrowed skeleton, with the borrowing unstated.
The practical consequence is a caveat on footprint claims, which are always comparative. “An 8B at 2.56 GB” means something different if the 8B in question is a known-good geometry someone else validated. This wiki records the match, not a verdict, and needs a second Fermion release or an independent evaluation to say more.
Whose price is on the sticker?
An open-weight model doesn’t have a price — it has a vendor rate and a set of independent hosts serving the same weights, and the corpus now holds two sources that disagree about which is which. glm-52 (T2, simon-willison) puts $1.40 / $4.40 per 1M at OpenRouter hosts; deepseek-glm-qwen-price-gap (T4) calls those figures Z.ai’s list price and reports resale separately at ~$0.55 / $1.85. Neither is first-party, and the T4 article’s headline claim — a 10x gap to DeepSeek V4 Flash at $0.14 input — only holds on its own reading.
Left open, because the underlying point is more interesting than the arithmetic: open weights turn a model’s price into a market rather than a number. Anyone can serve them, so “what does GLM-5.2 cost” has no single answer, and comparisons against closed models (one vendor, one rate) are not like-for-like. That belongs alongside the tokens-per-task multiplier in llm-api-pricing — two different ways a headline per-token rate fails to predict a bill.
Method note worth keeping: this conflict surfaced because a T4 aggregator was ingested into a spoke already holding T1 and T2 on the same subjects. Its spec sheet checked out against both, its GLM release date was three days off, and its pricing framing broke exactly where the headline lived. Weak sources are cheap to ingest and worth grading against strong ones already in the corpus.
Cross-spoke adjacency
- research-wiki — owns anthropic & claude-opus-4-8 as the model substrate under the tools-for-thought ecosystem (capability+cost lens). Here they’re entries in the broader market.
- llm-inference-wiki — how these models run (MoE efficiency, kv-cache, serving). This spoke is who makes them and what they cost.
- agentic-tooling-wiki — consumes providers (model selection, “Sonnet+harness > raw Opus”); the cost/capability frontier here sets the harness economics there.
- ai-governance-wiki — AI export controls / regulation (the Fable/Mythos export ban openai-gpt56-grok45-clash) govern which models are available at all; this spoke owns the market those rules constrain.
Index — LLM Providers Wiki
Catalog of every page, grouped by schema.org
@type. Spine: synthesis (thesis),log.md(history), this file (catalog). Some wiki-links resolve cross-wiki (research-wiki, llm-inference-wiki) — intentional bridge links. Pricing/ranking facts are dated snapshots.
DefinedTerm (concepts)
- llm-provider — umbrella: the kinds of providers (frontier labs, open-weight labs, cloud resellers) and their axes · domain
- open-weight-models — the open-weight wave: MoE, Apache-2.0/MIT licensing, local deploy · concept
- llm-api-pricing — per-token pricing tiers, the ~600× spread, caching/batch/routing levers · concept
- llm-benchmarks — composite multi-axis leaderboards (capability × speed × price × context) · concept
- quantization — int4/2-bit precision as a footprint lever; QAT vs PTQ, plus native-format training as a third case; sub-1GB open models · mechanism
- ternary-weights — minus/zero/plus at ~1/8 the bytes of fp16; trained in, not quantized into; the format as market lever · mechanism
- abliteration — weight-editing away a model’s refusal direction; only possible on open weights; the coherence-cost trade · mechanism
- arc-agi — the novelty benchmark: unseen tasks, ARC-AGI-3’s interactive-game format, the model-alone (no-harness) rule, and the contamination problem a novelty test can’t escape · benchmark
- ai-gateway — the one-endpoint layer over many providers: managed resellers vs self-hosted routers; fallback, cost and token-volume policy · concept
Organization (providers)
- deepseek — Chinese lab; low-cost OpenAI-compatible API + strong open weights; the seed subject
- google — dual-track: proprietary Gemini frontier + open-weight Gemma family
- openai — frontier incumbent; proprietary GPT/o-series; its API is the de-facto compatibility standard ·
source - mistral-ai — Europe’s leading lab; Apache-2.0 open weights (Mixtral MoE) + proprietary API; 2026-07 pivot to a sovereign full-stack platform ·
source - qwen — Alibaba; the leading Chinese open-first family (Apache-2.0; deepest ladder; 200k+ HF derivatives) ·
source - xai-grok — Elon Musk’s xAI; proprietary frontier; 2M-token context; open→closed arc (Grok-1 was Apache-2.0) ·
source - amazon-bedrock — AWS reseller/aggregator: many providers’ models behind one API (the cloud-reseller axis) ·
source - sakana-ai — Tokyo lab behind sakana-fugu; this spoke’s first Japanese entry, and a composer of other providers’ models presenting as a lab with its own family; sells into Codex/Claude Code rather than a chat app
- cohere — Canadian enterprise/sovereign-AI lab; Command + North model families; no consumer product
- z-ai — Chinese open-weight lab (formerly Zhipu AI) behind the GLM family; text-only flagship + separate vision models
- internscience — open-weight, agentic, science-leaning lab behind agents-a1 (built on Qwen3 lineage)
- thinking-machines-lab — US open-weight startup behind inkling + tinker; concedes the frontier, sells the fine-tuning
- arc-prize — runs the arc-agi benchmark + leaderboard; scores the model alone, no harness; the spoke’s nearest thing to an independent referee
- fermion-research — ships models, not an API: three Apache-2.0 ternary models + engine, no waitlist; publishes measurement depth while keeping format and training method closed. No founders, funding or independent eval known
- readyart — HF org publishing weight-edited derivatives (gemma-4-31b-it-scotoma); trains nothing, ships modified checkpoints — the spoke’s first derivative publisher
SoftwareApplication (models)
- gemma-4 — Google’s open-weight family (E4B / 12B / 26B-A4B); Apache-2.0; encoder-free multimodal 12B
- gemma-4-31b-it-scotoma — ReadyArt’s abliteration of Gemma 4 31B-it; Apache-2.0 33B BF16, Jacobian-lens projection keeping ~22% of the edit re-applied at 1.5× over layers 7–41; no evals, and its own card says it “refuses basically as much as its base model” ·
source· T3 · huggingface.co - gemini — Google’s closed-weight frontier family (Gemini 3.5 agents/coding, Omni video, for Science)
- llama — Meta’s open-weight family (Llama 1→4); catalyzed the open wave; custom (non-OSI) license ·
source - cohere-north-mini-code — Cohere’s first agentic coding model; Apache-2.0 open-weight 30B MoE (3B active), single-H100, 256K ctx; SWE-Bench 83.2% ·
source· T3 · cohere.com - kimi-k3 — Moonshot’s open-weight multimodal agentic flagship, released 2026-07-28: 2.8T total / 104B active (16 of 896 experts), 1M ctx, native MXFP4 QAT, MoonViT-V2 vision; splits decisions with the closed frontier (GPQA 93.5, SWE-Marathon 42.0 ahead; DeepSWE 67.5 behind Fable 5’s 70.0) on a table that documents its own confounds; bespoke licence with a $20M MaaS revenue threshold · T1 · github.com/MoonshotAI
- deepseek-glm-qwen-price-gap — Tech Insider: DeepSeek V4 / GLM-5.2 / Qwen3.6 open-weight price comparison; T4 SEO+affiliate template (“X vs Y vs Z: $N Price Gap [2026]”), specs corroborate against glm-52/deepseek-api-docs but the headline 10x rests on calling $1.40/$4.40 GLM’s vendor list where this wiki has it as reseller — conflict flagged ·
source· T4 · tech-insider.org - glm-52 — Z.ai’s text-only open-weight flagship; MIT, 753B-A40B MoE, 1M ctx; #1 open-weight on Artificial Analysis (v4.1), #2 Code Arena WebDev ·
source· T2 · simonwillison.net - claude-sonnet-5 — Anthropic’s agentic mid-tier (
claude-sonnet-5, 2026-06-30); pitched near Opus 4.8 at Sonnet price ($3/$15 std, $2/$10 intro); SWE-bench Verified 85.2%, HLE 57.4% w-tools, OSWorld 81.2%, 1M ctx ·source· T1 · anthropic.com - mage-flow-turbo — Microsoft’s open-weight text-to-image model; MIT, 4B NR-MMDiT, rectified flow matching + 4-step distillation, native 512–2048px; claims 0.59 s/1024² on one A100 and parity-or-better vs Qwen-Image 20B / FLUX.2 32B (GenEval 0.88, self-reported); 2nd generative-media source, 1st non-Google ·
source· T3 · huggingface.co - agents-a1 — InternScience’s open-weight agentic model; Apache-2.0, 35B MoE (Qwen3 lineage), 262K ctx; claims trillion-param-class agentic perf by distillation not scale (GAIA 96.0, IFEval 94.8, self-reported) ·
source· T3 · github.com - claude-fable-5 — Anthropic’s
claude-fable-5, Claude 5-family fast/coding-strong model; $10/$50 published by Anthropic at launch (found 2026-08-03 in research-wiki’s announcement page; three third-party reads had independently corroborated it) + pay-per-use after 2026-07-12; Anthropic’s own Opus 5 post places it above Opus 5 on intelligence at 2× the price - claude-opus-5 — Anthropic’s
claude-opus-5(2026-07-24); near-frontier flagship at $5/$25, same as Opus 4.8, pitched as “close to Fable 5 at half the price”; fast mode 2×price/~2.5×speed; ratio-only vendor benchmarks; context window unstated ·source· T1 · anthropic.com - sakana-fugu — a multi-agent system sold as a model: OpenAI-compatible endpoint,
fugu/fugu-ultra-v1.1/fugu-cyber, effort high/xhigh/max (no low or medium), routes across all providers by default with the pool set per API key; the ai-gateway that calls itself a model · T1 · fugu.sakana.ai - neutrino-1 — Fermion Research’s Apache-2.0 ternary family (8B / 0.6B / 0.6B-Chat, 2026-07-27): an 8B-class model in a 2.56 GB download, weights trained in ternary rather than quantized; MMLU 72.1 beats fp16 8Bs at 1/6 the bytes, but loses IFEval and BFCL, ships a 40,960 context (smallest in its own table), and every number is first-party. Architecture matches Qwen3-8B field for field, unmentioned ·
source· T1 · fermionresearch.com - astra — OpenAI‘s unreleased “next major model”, named 2026-08-01 in a mathematics paper rather than a model launch: an internal version produced ten results in maths/TCS and wrote their Lean proofs; search cost ~$2,000 of tokens priced at Sol rates (a volume proxy — Astra has no price card). No params, context, benchmarks, or date. Its public evidence is machine-checkable proofs, not a vendor eval table — see ten-proofs · T1 · openai.com
- inkling — Thinking Machines Lab’s open-weights base (2026-07-15); Apache-2.0, 975B-A41B multimodal MoE, 45T tokens, 1M ctx; self-declared not the strongest — a base to fine-tune; needs 2TB+ VRAM (600GB NVFP4)
- tinker — Thinking Machines’ fine-tuning platform + OpenAI-compatible API for inkling; the business model behind the free weights
TechArticle / BlogPosting / Article / Dataset (sources)
- anthropic-model-docs — Anthropic’s platform docs (models overview + effort reference): the operational T1 against the launch posts’ promotional T1. Supplies what the Opus 5 announcement omitted — 1M context, 128k/300k output, May 2026 knowledge cutoff (four months newer than Fable 5’s), Bedrock/AWS/Google/Foundry availability — and corrects the effort ladder from four levels to five: the launch post drops
medium, the level the docs tell developers to use “liberally” as the primary cost control ·source· T1 · platform.claude.com - deepseek-api-docs — DeepSeek’s API reference (V4 Flash/Pro, thinking mode, context caching) ·
source· api-docs.deepseek.com - fable5-gpt56-kimi-k3-creators — Medium (Christie C.): creator-opinion Fable 5 vs GPT-5.6 vs Kimi K3; benchmark-vs-real-utility thesis; sole source for the kimi-k3 claim ·
source· T4 · medium.com - alphabet-q2-2026-earnings — Google (Pichai Q2 2026 CEO message): Gemini lineup spread — 3.6 Flash / 3.5 Flash-Lite / 3.5 Flash Cyber (domain-specialized) / 3.5 Pro testing / Gemini 4 pre-training; scale = ~22B tokens/min, 950M Gemini-app MAU, ~900M Gemma downloads; cloud/search/ADK facets cross-linked ·
source· T1 · blog.google - open-source-llms-2026 — Hugging Face: the 2026 open-weight roundup (Llama 4, Qwen3, Gemma 4, Kimi, …) ·
source· huggingface.co - gemma-4-12b-announcement — Google: Gemma 4 12B, encoder-free multimodal (vision+audio), 16GB local ·
source· blog.google - google-ai-updates-may-2026 — Google: May 2026 AI roundup (Gemini 3.5, Omni, for Science; + out-of-scope products) ·
source· blog.google - gemma-4-qat — Google: Gemma 4 quantization-aware-training checkpoints (Q4_0 + 2-bit mobile; sub-1GB E2B) ·
source· blog.google - llm-api-pricing-comparison — CloudZero: every major model ranked by cost; the ~600× spread ·
source· cloudzero.com - llm-leaderboard-stats — llm-stats.com: 300+ models by composite intelligence/speed/price ·
source· llm-stats.com - artificial-analysis — independent, methodology-disclosed benchmark (Intelligence Index v4.1: 9 evals, four weighted categories Agents 34% / Coding 24% / SciReasoning 24% / General 18%; blended price 7:2:1, TTFT) ·
source· T2 · artificialanalysis.ai - moe-architecture — Wikipedia: Mixture-of-Experts definition (gating/router, top-k experts, total-vs-active params, shared experts) ·
source· T2 · en.wikipedia.org - hf-quantization-concepts — Hugging Face Transformers: quantization concepts (int8 ~4× / int4 packing & bandwidth, FP8 E4M3/E5M2, affine, PTQ vs QAT) ·
source· T1 · huggingface.co - claude-refusals-and-fallback — Anthropic API: the
refusalstop_reason + model-fallback (server-side/SDK/manual), fallback-credit billing, sticky routing ·source· platform.claude.com - claude-api-rate-limits — Anthropic API: usage tiers, spend caps, RPM/ITPM/OTPM, token-bucket pacing, and cache-aware ITPM (cached input doesn’t count) ·
source· T1 · platform.claude.com - gemini-omni-flash-nano-banana-2-lite — Google: two Gemini generative-media models — Nano Banana 2 Lite (text-to-image, $0.034/1k, 4s) + Omni Flash (video, $0.10/sec, 10s) ·
source· T3 · blog.google - prompting-claude-sonnet-5 — Anthropic docs: prompting/migration guide for Sonnet 5 — effort param (medium S5 ≈ high S4.6), adaptive thinking default, no sampling params (400), new tokenizer (~30% more tokens), literal instruction-following ·
source· T1 · platform.claude.com - claude-sonnet-5-system-card — Anthropic: 145-page pre-deployment report for Sonnet 5 — full benchmark table (vs GPT-5.5 / Gemini 3.5 Flash), 1M/10M context, RSP determination, alignment & welfare; reveals a Mythos class above Opus ·
source· T1 · anthropic.com - fable-5-vs-sonnet-5-token-cost — Njenga (Medium, member-only): Fable 5 vs Sonnet 5 cost in Claude Code — a
"Hello"cost $0.4745 on Fable vs ~3.2× less on Sonnet; Sonnet caches more (24.8k vs 15.6k) so bills less ·source· T4 · medium.com - grok-4-5-price-vs-benchmarks — The Decoder: Grok 4.5 at $2/$6 undercuts Opus 4.8/GPT-5.5/Fable 5 ($10/$50); trails on coding benchmarks but xAI argues cost-per-task (4.2× fewer tokens) wins — the China price-play from a US lab ·
source· T3 · the-decoder.com - openai-gpt56-grok45-clash — Business Insider: OpenAI GPT-5.6 (Sol/Terra/Luna) rolls out Thu (staggered per Trump admin) vs Musk’s “Opus-class” Grok 4.5 days later; both pitch token efficiency; Fable/Mythos export-ban backdrop ·
source· T3 · businessinsider.com - dont-use-fable-5-hassid — Ruben Hassid (“How to AI”): buyer’s guide to Fable 5 — 3rd read of $10/$50 + pay-per-use after 2026-07-12, per-turn cost blowup (≈$14/40 turns), and a Fable-for-hard-goals / Opus-4.8-as-workhorse ladder ·
source· T3 · ruben.substack.com - inkling-willison-review — Simon Willison on Inkling: the release facts, thin training-data docs, the US open-weights bench (Inkling + Nemotron + Gemma 4), and a pelican the model called a stork ·
source· T2 · simonwillison.net - inkling-announcement — Thinking Machines Lab: the Inkling announcement + model card — 66-layer MoE (256+2 experts, 6 active), 1M ctx, benchmark table, VRAM needs ·
source· T1 · thinkingmachines.ai - mistral-platform — Mistral’s own site (2026-07): the model lab now a sovereign EU-hosted full-stack platform — agents (Vibe / Vibe for Code), Studio, Forge, Compute; Medium 3.5 / Small 4 / OCR 4 / Voxtral TTS named ·
source· T2 · mistral.ai - claude-opus-5-announcement — Anthropic’s Opus 5 launch post: $5/$25 (flat vs 4.8), paid fast mode (2× price / ~2.5× speed), effort low→max, six cost-per-task benchmark ratios with no absolute scores, and Mythos 5 named as a cyber/biology specialist rather than the ceiling ·
source· T1 · anthropic.com - omniroute — MIT self-hosted ai-gateway (v3.8.49, ~30K stars): one local endpoint over a claimed 290+ providers / 500+ models, 4-tier cascade fallback, 19 routing strategies, 12-engine prompt compression (15–95%), ~1.53B free tokens/month pooled from 43 tiers — all self-reported ·
source· T1 · github.com - opus-5-arc-agi-3 — The Decoder: ARC Prize scores Opus 5 30.2% on ARC-AGI-3 vs a 7.8% record (Fable-class ~20%), 90.4%/97.5% on -2/-1; carries its own counter-evidence — the gain shrinks to a 3-way tie on the private Witness benchmark, and Opus 5 postdates the benchmark’s public format ·
source· T3 · the-decoder.com - fugu-get-started — Sakana’s Fugu onboarding docs: endpoint/keys/effort levels, per-key provider pools, Codex + Claude Code install; ships agent-conduct guards in
base_instructionsand a 2-hour stream idle timeout (vs Codex’s ~5-min default); no pricing, no benchmark, no context window ·source· T1 · fugu.sakana.ai - cerebras-pricing — Cerebras’ pricing page: a speed-priced inference provider with no per-token rate card — daily-token subscription buckets (Code Pro $50/24M-day, Max $200/120M-day, both sold out), $5 free credit, Developer from $10 ·
source· T2 · cerebras.ai
Person
- simon-willison — independent dev/writer; the methodology-transparent practitioner voice behind the GLM-5.2 read (and recurring cross-wiki)
- joe-njenga — Medium AI-tooling writer; practitioner “I tested X” anecdotes (source of the Fable 5 vs Sonnet 5 cost test) · T4
- ruben-hassid — “How to AI” Substack; non-technical AI consultant translating docs into buyer’s-guide advice (Fable 5 usage/pricing) · T3
- diego-souza — GitHub
diegosouzapw; owner/maintainer of omniroute, the spoke’s first self-hosted gateway
Synthesis
- synthesis — the thesis: collapsing price floor vs. premium frontier, open weights as the wedge
Bridge nodes (live in sibling wikis, linked cross-wiki)
anthropic · claude-opus-4-8 (research-wiki) · llm-inference · kv-cache · cerebras-inference · cerebras-systems (llm-inference-wiki)