Log — LLM Providers Wiki
Append-only history. Each entry starts with ## [YYYY-MM-DD] <op> | <title> where
<op> is ingest, query, lint, or split, so grep "^## \[" log.md | tail -5 works.
[2026-07-19] ingest | Cerebras pricing (hub-routed, URL-only, T2)
Second Cerebras source from Telegram, hub-routed here (runner-up last time; the pricing/market angle is this spoke’s turf — “API access & pricing” — while the mechanism/company sit in llm-inference-wiki). Cerebras is not a model maker but a speed-differentiated inference provider reselling open models on wafer-scale HW; its pricing page publishes no per-token rate card and instead sells daily-token subscription buckets (Cerebras Code Pro $50/mo→24M tok/day, Max $200/mo→120M tok/day, both sold out; $5 free credit; Developer from $10; Enterprise = contact sales). Created source summary cerebras-pricing (WebPage, url-only). Folded a new “speed-priced, no rate card” axis into llm-api-pricing and a “second pricing shape — the speed provider” paragraph into synthesis. No duplicate org node — the provider/company cerebras-systems and the service cerebras-inference already live cross-wiki in llm-inference-wiki; linked as bridge nodes (added to the index footer), not duplicated. Cross-spoke note: the Cerebras Code subscription is a coding-agent product (agentic-tooling-wiki) — recorded as context, not routed away. Volatile T2 snapshot (tiers/prices/“sold out” will drift). Runner-up: llm-inference-wiki.
[2026-06-05] ingest | Gemma 4 with Quantization-Aware Training (hub-routed)
Hub-routed Telegram drop. Clear match (existing gemma-4); runner-up llm-inference-wiki (quantization-as-mechanism) declined — that spoke owns sampling/KV-cache/batching, not quantization, and the source’s angle is deployable footprint in the open-weight market. Ingested:
- New source gemma-4-qat (
source, url-only, BlogPosting): QAT checkpoints for E2B/E4B/26B-MoE; Q4_0 4-bit + a targeted 2-bit mobile scheme (mixed precision, channel-wise, static activation scaling); E2B text < 1 GB; QAT claimed above PTQ quality; GGUF/compressed-tensor on HF; broad local runtime list (llama.cpp/Ollama/LM Studio/MLX/vLLM/SGLang/Transformers.js/LiteRT-LM/Unsloth). - New concept quantization (DefinedTerm) — int4/2-bit footprint lever, QAT vs PTQ; scoped to the market/deployability angle, with the deeper inference-mechanics depth flagged as llm-inference-wiki’s (llm-inference) territory (bridge, no dup).
- Updated gemma-4 (QAT bullet + expanded deployment list; also tidied a stray
</content>tag). - Updated synthesis point 2: quantization is the second footprint lever after sparse MoE — footprint now contested by training technique, not just architecture.
- Updated index (added quantization + the source). Volatile caveat: the “<1GB” / “higher-than-PTQ quality” figures are vendor claims, undated vs neutral benchmarks — snapshots. Site rebuilt + verified.
[2026-06-01] split | llm-providers-wiki created from _inbox cluster llm-providers (4 sources)
The cluster was seeded by the parked deepseek-api-docs stub; on user instruction (“seek for content about llm-providers” → option (a)), the router web-searched the model/provider landscape and curated 3 strong sources to ingest, taking the cluster to 4 and triggering this spin-out.
Scaffolded from CLAUDE.template.md; domain = the LLM model & provider landscape. Ingested 4
sources (all URL-only): deepseek-api-docs (TechArticle), open-source-llms-2026 (HF,
BlogPosting), llm-api-pricing-comparison (CloudZero, Article), llm-leaderboard-stats
(llm-stats.com, Dataset). Created 4 concept pages (llm-provider, open-weight-models,
llm-api-pricing, llm-benchmarks) + 1 provider page (deepseek). Synthesis thesis:
collapsing price floor (open weights + DeepSeek) vs. still-premium proprietary frontier; “best
model” is multi-axis. Cross-wiki bridges: anthropic/claude-opus-4-8 (research-wiki),
llm-inference/kv-cache (llm-inference-wiki). Deleted the parked _inbox deepseek record.
Caveat baked into the schema: pricing/rankings are volatile — pages are dated snapshots.
Spoke count 8 → 9.
Note on sourcing: 3 of the 4 founding sources were router-found (web search), not human-dropped — a departure from the usual human-curated flow, done at explicit user request.
[2026-06-04] ingest | Introducing Gemma 4 12B (Google blog)
Hub-routed (Telegram). New source summary gemma-4-12b-announcement (BlogPosting, url-only). Gemma 4 12B: a dense, mid-size, encoder-free multimodal open-weight model — vision via a single matrix-mul embedding, native audio projected straight into the token space, no separate encoders. Benchmarks “nearing” the 26B-A4B at <½ the memory; runs on 16GB VRAM/unified memory; Apache-2.0; HF/Kaggle; llama.cpp/MLX/vLLM/Transformers. New Thing pages: gemma-4 (SoftwareApplication — family: E4B / 12B / 26B-A4B) and google (Organization — first provider page beyond deepseek; dual-track Gemini + Gemma). Updated open-weight-models (added a multimodality at the local tier trend + footprint-as-battleground) and synthesis (modality/memory as a new axis; “dual-track labs” recurring read). Index updated with a new SoftwareApplication group. Gemma 4 was previously only name-dropped (the 26B-A4B) in open-source-llms-2026; this gives the family a real page + the new 12B variant. Snapshot caveat: benchmark/footprint claims are the vendor’s own.
[2026-06-04] lint | full pass (13 pages)
Triggered after the Gemma 4 ingest. Structure green: no dangling wikilinks (all 12 local slugs resolve; the 4 non-local targets — anthropic, claude-opus-4-8, llm-inference, kv-cache — are the declared cross-wiki bridges); no orphans (every page linked ≥3×); index catalogs all pages + synthesis; @types correct (providers=Organization, models=SoftwareApplication, concepts=DefinedTerm, sources=CreativeWork subtypes) — none need narrowing.
Fixed — missing cross-links (created by the new google/gemma-4 pages): linked previously
plain-text mentions → google in llm-provider, llm-api-pricing, llm-api-pricing-comparison;
gemma-4+google in open-source-llms-2026. Bumped those 4 pages’ updated: to 2026-06-04.
Flagged, not auto-fixed (need a source / judgment):
- Opus version drift: pricing pages list Claude Opus 4.7 ($5/$25) while benchmarks/leaderboard/ provider pages reference claude-opus-4-8 as current #2. Both are dated mid-2026 snapshots from different sources (not a contradiction), but Opus 4.8 pricing is uncaptured — refresh when a pricing source for 4.8 arrives. Don’t invent a number.
- Gemini snapshot lag: llm-api-pricing-comparison quotes CloudZero’s Gemini 2.5 Flash $0.15/$0.60 while llm-api-pricing cites Gemini 3.1 Pro — faithful to each source, but the 2.5 figure is stale; flag on next pricing refresh.
- Thin anchor: google (1.1KB) is the lightest page — a deliberate brief org anchor; optional enrichment (name the Gemini model line + its llm-benchmarks placement). Not an error.
[2026-06-05] ingest | Google AI updates — May 2026 (hub-routed)
Hub-routed Telegram drop (broad Google “what we shipped in May” roundup). Routed here over runners-up search-marketing-wiki (Universal Cart / agentic-commerce, AI-search) and agentic-tooling-wiki (Android Halo / agent surfaces) because the source’s dominant in-scope substance is model releases. Ingested the model/provider signal; explicitly scoped out the consumer-product/hardware/search items (noted as cross-spoke context in the source page, not parked separately — facets of one roundup).
- New source summary google-ai-updates-may-2026 (
source, url-only, BlogPosting). - New Thing page gemini (SoftwareApplication) — the closed-weight frontier family that the wiki kept referencing but had no page for (filled the gap the 2026-06-05-earlier lint flagged on google). Anchors Gemini 3.5 (“frontier intelligence for agents and coding”), Gemini Omni (multimodal video gen), Gemini for Science.
- Updated google (Gemini section + dual-track-converges-on-agents/coding point; linked gemini).
- Updated synthesis “Dual-track labs” recurring read: open/closed split is licensing+footprint, not target workload — both Gemini and Gemma now pitched at agents-and-coding.
- Updated index (added gemini model + the source). Volatile caveat: Gemini 3.5/Omni facts are marketing-roundup positioning, undated benchmarks — treated as snapshots. No pricing given. Site rebuilt + verified.
[2026-06-09] ingest | +3 major providers (OpenAI, Mistral AI, Llama) — all-spokes cron test
Filled the most glaring landscape gaps — the frontier incumbent and the non-Chinese open-weight pole: openai (Organization, src — proprietary GPT/o-series; the de-facto OpenAI-compatible API standard; MS partnership), mistral-ai (Organization, src — Europe’s leading lab; Apache-2.0 Mixtral MoE + proprietary API), llama (SoftwareApplication, src — Meta’s open-weight family that catalyzed the wave; open-weight but non-OSI custom license). Sharpens the licensing-spectrum thread (Llama vs Apache-2.0 gemma-4/mistral-ai). Wikipedia-sourced background; pricing/rankings remain dated snapshots. 16 → 19 pages.
[2026-06-10] ingest | Qwen + xAI/Grok + Amazon Bedrock — all-spokes pass
Three new pages filling named-but-unpaged gaps. qwen (Org/SoftwareApplication, url, Wikipedia) — Alibaba’s leading open-first family the thesis kept citing: 0.6–32B dense + MoE (30B-A3B), reasoning/VL/audio/Coder/omni, mostly Apache-2.0; 200k+ HF derivatives, 234M app users; the strongest open-weight challenger to deepseek. xai-grok (Org/SoftwareApplication, url, Wikipedia) — xAI’s proprietary frontier model; 2M-token context (the “xAI leads context” claim), and an open→closed arc (Grok-1 Apache-2.0 → all later closed), the inverse of llama/qwen → sharpens the “open is a spectrum/strategy-over-time” thread. amazon-bedrock (Org/SoftwareApplication, url, AWS) — the wiki’s first cloud reseller: one-API aggregator of Claude/Llama/Mistral/Cohere/Nova (not a maker); the demand-side mirror where model choice is a config param and differentiation moves to routing/caching/data. Folded into synthesis (new 2026-06-10 section) + index (Organization rows). All vendor/encyclopedic snapshots — standing volatility caveat. No contradictions. 19 → 22 pages.
[2026-06-11] ingest | Claude API — Refusals and Fallback (hub-routed)
Hub-routed Telegram drop (platform.claude.com). Classified against wikis.md: dominant substance
is the Anthropic provider API surface + model routing + billing, so routed here over the
runner-up agentic-tooling-wiki (the SDK-middleware / sub-agent-fallback angle — noted as
cross-spoke adjacency, not split). Not research-wiki (that’s the model-substrate capability/cost
bridge, not API mechanics) and not ai-governance-wiki (this is API-level classifier behavior, not
policy/GRC). Ingested:
- New source claude-refusals-and-fallback (
source, url-only, TechArticle):stop_reason: "refusal"as an HTTP 200;stop_details.category(cyber/bio/frontier_llm/reasoning_extraction); three fallback paths (server-sidefallbacksbeta, SDKBetaRefusalFallbackMiddleware, manual);fable-5 → opus-4.8chain; per-attempt billing viausage.iterations[];fallback-creditbeta; sticky routing; batch behavior; “refusals are invisible to error-rate monitoring” pitfall. - Updated synthesis recurring-read “engineering sets real cost”: model routing is now a first-class API primitive (server-side fallback + fallback-credit + sticky routing), not just app glue — hardening the cost-is-engineering thread into the provider’s own surface.
- Updated index (added the source under TechArticle/sources).
- Linked bridge nodes anthropic & claude-opus-4-8 cross-wiki (no dup).
[2026-06-12] ingest | Artificial Analysis — independent LLM benchmark platform
All-spokes daily expansion. Added artificial-analysis (@type WebPage) — the neutral, reproducible, methodology-disclosed benchmark the open question explicitly wanted (a rigor step up from the flagged-as- biased llm-leaderboard-stats/llm-api-pricing-comparison). Captured the four-axis frame (quality/ speed/price/context), the Intelligence Index v4.0 composite (GPQA Diamond, HLE, τ²-Bench, Terminal-Bench, SciCode, IFBench, …), the blended price at 7:2:1 cache:input:output, and live TTFT — incl. that it benchmarks API providers (the host/reseller layer, amazon-bedrock) not just labs. Resolves the “vendor/SEO bias — want a neutral benchmark” open question (residual caveat: disclosed editorial weighting + volatile snapshot). Wired to llm-benchmarks / llm-api-pricing; synthesis + index updated. 1 new page.
[2026-06-13] quality | maintenance cycle (tier + freshness backfill)
Quality Cycle (not expansion — spoke grown 06-12, never pad). Backfilled tier: + freshness: volatile
on all 15 source pages: T1×2 (official Anthropic/DeepSeek docs), T2×8 (Artificial Analysis + Wikipedia
entities + HF/llm-stats), T3×5 (Google/AWS/CloudZero vendor pages). Operationalizes the volatile-snapshot
staleness tracking the spoke most needs (pricing/leaderboards churn weekly). Scorecard → hub quality-log.md.
0 new pages. No synthesis/index change (frontmatter-only).
[2026-06-15] ingest | Cohere / North Mini Code — thenewstack.io (honest stub)
Article body JS-gated; fetch failed. From headline + URL: Cohere is pivoting from enterprise sovereign-AI to developer-targeted coding models with “North Mini Code.” T4 (The New Stack trade press). Created cohere (Organization, first Cohere page) + cohere-north-mini-code (TechArticle stub). Entity discovery: Cohere is a new org node (no match in entity-index), but entity pages for a new org from a T4 trade-press stub without confirmed founding facts felt over-reaching — noted as a gap; a T1 primary source would warrant a full entity page. Synthesis updated with sovereign-AI niche. Runner-up: none.
[2026-06-14] dedup | Google — made canonical here (merged the duplicate agentic-tooling node)
The entity-index fuzzy audit flagged Google paged twice (agentic-tooling + here). Per ENTITIES.md’s
canonical-node rule, made this the one Google node (provider landscape is Google’s core identity;
richer + url-sourced), folding in the agent-platform-vendor facet (ADK, agentskills adoption) with
cross-links to adk/agentskills-spec, + aka: aliases. The agentic-tooling copy was deleted; its
google links now resolve cross-wiki here.
[2026-06-15] lint | North Mini Code — T4 stub upgraded to T3 primary source
Quality cycle: retried the JS-gated The New Stack stub, then sourced the primary instead — cohere‘s own blog (cohere.com/blog/north-mini-code) + the Cohere Labs Hugging Face model card (found via WebSearch). Rewrote cohere-north-mini-code from a headline-only T4 stub into a full T3 model page (retyped TechArticle → SoftwareApplication): 30B MoE / 3B active, 128 experts (8/tok), 256K ctx / 64K gen, Apache-2.0, single-H100 FP8; SWE-Bench Verified 83.2% pass@1, Terminal-Bench v2 62.8%, Artificial Analysis Coding Index 33.4. Correction: the earlier stub framing (and the synthesis provider-map) called Cohere a non-open-weight sovereign player — North Mini Code is in fact Apache-2.0 open-weight, so cohere, the synthesis “enterprise/sovereign” bullet, and the index were corrected to put Cohere on the open-weight axis (sovereignty via open weights, not opposite them). url swapped thenewstack.io → cohere.com.
[2026-06-18] ingest | GLM-5.2 (Z.ai) — Simon Willison review (hub-routed)
Hub-routed Telegram drop (simon-willison‘s post; T2, independent practitioner). Clear match — the open-weight text-LLM market; no runner-up (text-only model, so not speech-audio; not inference mechanics). GLM/Z.ai were plain-text mentions before this; now paged. New:
- glm-52 (
source, url-only, SoftwareApplication): MIT, 753B-A40B MoE, 1.51TB, 1M ctx, text-only. #1 open-weight on artificial-analysis Intelligence Index v4.1 (51, ahead of MiniMax-M3 44 / DeepSeek V4 Pro 44); #2 Code Arena WebDev behind Claude Fable 5. ~$1.40/$4.40 per 1M via OpenRouter. Caveat recorded: token-hungry (~43K output tok/task), so the cheap rate overstates the real-cost gap. - New entities: z-ai (Organization, fmr Zhipu AI — the GLM maker, open-weight challenger slot) and simon-willison (Person, recurring cross-wiki practitioner voice; the SVG-pelican tester).
- Folded into synthesis (open-weight-challenger bullet + the “does the frontier premium survive?” open question — strongest open-vs-closed data point yet, an open model #2 on a coding board at ~1/5 the price). Updated open-weight-models (GLM-5.1→5.2) and artificial-analysis (noted v4.1). No contradiction with existing pages. Volatile snapshot dated 2026-06-18.
[2026-06-18] ingest | DefinedTerm enrichment pass (subagent)
Deepened the spoke’s thinnest DefinedTerm pages with sourced mechanism detail (all five were ~32–40 lines). Two new source pages + one refreshed in place:
- hf-quantization-concepts (TechArticle, T1, huggingface.co) — int8 ~4× smaller than FP32 / int4 halves again + weight-packing & memory-bandwidth payoff; FP8 E4M3/E5M2; affine scale+zero-point; per-tensor vs per-channel; PTQ vs QAT (QAT better at low bit-widths). Drove quantization: added “size math”, “why int4 is the workhorse + FP8 alternative”, and grounded the QAT-vs-PTQ claim.
- moe-architecture (Article, T2, en.wikipedia.org) — gating/router picks top-k experts (k=1/2), conditional computation, total-vs-active split (less compute despite 30× more params), shared experts (DeepSeek). Drove open-weight-models “MoE dominance” bullet (now explains the mechanism behind large-total/small-active serving).
- artificial-analysis refreshed in place to v4.1 composition: 9 evals, four weighted categories (Agents 34% / Coding 24% / SciReasoning 24% / General 18%; agents+coding = 58%), ±1% 95% CI. Drove llm-benchmarks: new “Anatomy of a composite” section.
- llm-provider and llm-api-pricing already developed with Related present — left untouched (no padding). Index updated with the two new source rows. One blocked fetch: en.wikipedia.org/wiki/Quantization_(machine_learning) 404’d; substituted the HF concepts doc. All volatile figures dated 2026-06-18; cross-wiki bridges (llm-inference, kv-cache, claude-opus-4-8) left as-is.
[2026-06-28] ingest | Claude API — Rate limits
New source page claude-api-rate-limits (TechArticle, T1, platform.claude.com), routed from the
hub. Anthropic’s rate-limit reference: two limit types (monthly spend caps — Start $500 / Build
$1,000 / Scale $200,000 / Custom none — and per-model RPM/ITPM/OTPM), token-bucket pacing
(continuous replenish, bursts trip 429s), auto tier promotion, separate pools for Batches / Managed
Agents / Fast mode, and the full anthropic-ratelimit-* header set + per-workspace limits. Dominant
hook for this spoke = cache-aware ITPM: cached input tokens (cache_read_input_tokens) don’t count
toward the rate limit (except Haiku 3.5 †), so prompt caching is a throughput lever, not only a price
cut (2M ITPM @ 80% hit ≈ 10M tok/min). Folded into llm-api-pricing (caching bullet) and the
synthesis “engineering sets real cost” recurring read; sits beside claude-refusals-and-fallback as
the provider-API-access pair. No new entities (publisher Anthropic already a cross-wiki bridge node).
All tier numbers dated 2026-06-28 (volatile snapshot).
[2026-06-30] ingest | Claude Sonnet 5 (Anthropic model release)
Routed by the hub router from Telegram. New model claude-sonnet-5 (claude-sonnet-5), released
2026-06-30 — Anthropic’s “most agentic Sonnet yet,” pitched as performing close to
Opus 4.8 at the Sonnet price. Created a single SoftwareApplication + source:true page (the model-
announcement precedent set by cohere-north-mini-code / glm-52). T1 (primary vendor announcement),
freshness volatile. Facts: pricing $2/$10 intro through 2026-08-31, $3/$15 standard; HLE 34.6% (no
tools) / 46.8% (w-tools); OSWorld-Verified 78.5%; default for Free/Pro, on Claude Code/Platform/API;
safety reported lower undesirable-behavior rate than Sonnet 4.6, weak on cyber exploit dev vs Opus;
context window not stated. Folded into synthesis open question “does the frontier premium survive?”
as a new angle — the squeeze coming from inside a frontier lab, not only open weights. Index updated.
No new entities (Anthropic + Opus 4.8 already cross-wiki bridge nodes in research-wiki). Vendor-reported
numbers await an independent artificial-analysis re-run.
[2026-06-30] ingest | Claude Sonnet 5 System Card (PDF, 145pp)
Follow-up to the same-day claude-sonnet-5 announcement — the primary pre-deployment report (routed from Telegram). Created source page claude-sonnet-5-system-card (@type Report, T1) with the full capability table (Sonnet 5 vs Sonnet 4.6 / GPT-5.5 / Gemini 3.5 Flash), RSP determination, and alignment/welfare assessment. Correction: the announcement-page figures (HLE 34.6/46.8, OSWorld 78.5) were Sonnet 4.6’s column in the card — a lossy-fetch misattribution; corrected the claude-sonnet-5 model page to the card’s Sonnet 5 numbers (SWE-bench Verified 85.2, HLE 43.2/57.4, OSWorld 81.2, FrontierCode v1 38.8) and flagged the correction in place per record-don’t-overwrite. Filled the open context-window field: 1M standard, 10M via context compaction (trigger 200k). Two notable new facts folded into synthesis: (1) the card admits Sonnet 5 “trails our Opus and Mythos-class models” on almost all benchmarks — so “close to Opus” is price-adjusted, not parity; (2) it discloses a Mythos class above Opus (Claude Mythos 5) as Anthropic’s current frontier. Cross-spoke: the RSP/threshold/safeguard machinery is ai-governance-wiki substance (noted in the source page; runner-up that spoke). No new entities (Anthropic + Opus 4.8 already cross-wiki bridge nodes; Mythos 5 not yet paged — left as a named future node). All numbers dated 2026-06-30 (volatile).
[2026-06-30] ingest | Agents-A1 (InternScience) — open-weight 35B agentic MoE
T3 first-party (GitHub repo + model card; self-reported benchmarks). Routed here by the hub router:
it’s a model (weights + training method), not a harness, so it’s the model-market’s subject — the
agentic capability angle is cross-spoke context with agentic-tooling-wiki, noted, not split. Apache-2.0,
35B MoE (active params undisclosed), 262K context, built on Qwen3
lineage (card ships a qwen3 reasoning parser). Pitch: “trillion-parameter performance with a 35B agent”
— close the gap by scaling the agent horizon, not parameters, via a 3-stage supervised method
(full-domain SFT → per-domain teachers → multi-teacher domain-routed on-policy distillation), not RL.
Six unified domains (long-horizon search / engineering / scientific research / instruction-following /
tool-calling + general). Self-reported: GAIA 96.04, IFEval 94.82, Seal-0 56.36, SciCode 44.33,
FrontierScience-Olympiad 79.0. New pages agents-a1 (model+source), internscience (Organization).
Touched synthesis — added a third mechanism to “does the frontier premium survive?” (training
technique, after cheaper tokens and flagship-self-squeeze) and listed A1 among the open-weight challengers.
arXiv tech report 2606.30616 (primary, but benchmark claims still self-run); awaits an independent
artificial-analysis re-run. Volatile snapshot.
[2026-06-30] ingest | Prompting Claude Sonnet 5 (Anthropic docs) — refresh of an existing model
T1 primary (platform.claude.com). Dedup: subject claude-sonnet-5 already paged → enriched in place
(new “API behavior & migration” section) + new source summary prompting-claude-sonnet-5. Key facts:
sampling params (temperature/top_p/top_k) now rejected with 400; manual extended thinking
(budget_tokens) removed; adaptive thinking on by default; effort param (low→max, default
high, xhigh for hardest) with the cross-gen mapping medium S5 ≈ high S4.6 / high S5 ≈ max S4.6; new
tokenizer ~30% more tokens (truncates 4.6-tuned max_tokens; raises effective cost); more agentic + more
literal instruction-following; code-review recall drop = harness effect not regression; computer use
computer_20251124. Touched synthesis — the tokenizer adds a Sonnet-5 instance of the “token
accounting, not rate, sets the bill” thread (alongside GLM-5.2 token-hunger), with the effort-tier uplift
as the offsetting lever. Cross-spoke: harness-facing advice → agentic-tooling-wiki; adaptive-thinking
mechanics → llm-inference-wiki (noted, not split). Volatile snapshot.
[2026-06-30] ingest | Nano Banana 2 Lite & Gemini Omni Flash (Google) — Gemini generative-media tiers
T3 vendor blog (blog.google). Two gemini-family generative-media models, now with concrete IDs +
pricing (the May roundup only had positioning). Nano Banana 2 Lite (gemini-3.1-flash-lite-image):
text-to-image, ~4s, $0.034/1k images, replaces gen-1 Nano Banana; claims prompt adherence, character
consistency, legible in-image text (Elo-vs-latency-vs-cost chart). Gemini Omni Flash
(gemini-omni-flash-preview): video gen + conversational editing, $0.10/sec = Veo 3.1 Fast, 10s max,
preview (no audio refs, no API scene-extension, character-consistency wobble). Pitched as an image→video
pipeline (Nano Banana still → Omni Flash animation; Interactions API ≤3 sequential edits). On AI Studio /
Gemini API / Enterprise Agent Platform + consumer surfaces. New source gemini-omni-flash-nano-banana-2-lite;
enriched gemini (two release entries). Touched synthesis — added a generative-media sub-theme
note: new pricing units ($/image, $/sec) and a spin-out watch (a dedicated image/video-gen spoke if
≥3 non-Google generative-media sources land; parallels the audio-gen→speech-audio carve-out). Volatile.
[2026-07-05] ingest | Fable 5 vs Sonnet 5 — token cost (Njenga, Medium) → llm-providers-wiki
Routed from hub. Member-only Medium blog (T4, paywalled past the baseline) comparing Fable 5
and Sonnet 5 cost inside Claude Code. Only the baseline test is readable: a single "Hello"
cost $0.4745 on Fable 5 vs ~3.2× less on Sonnet 5, with Sonnet reading more from cache (24.8k vs 15.6k) and
so billing below fresh input. New source page fable-5-vs-sonnet-5-token-cost; new Thing page
claude-fable-5 (first page for it — folds in the existing Code Arena WebDev coding-lead datapoint, so “fast,
coding-strong”, which makes the cost anomaly notable); new joe-njenga Person node (entity discovery).
Synthesis: added a caching-efficiency-as-intra-lab-cost-axis open question + logged the unverified anomaly under
Contradictions (held pending a T1 Anthropic price sheet). Cross-spoke context noted on the source page (Claude
Code → agentic-tooling; prompt caching → llm-inference kv-cache) — not re-paged. Numbers are one unrepeated,
harness-inflated run; recorded as claim, not fact.
[2026-07-08] ingest | Grok 4.5 cheap vs Fable 5 / GPT 5.5 (The Decoder) → llm-providers-wiki
Routed from hub (Telegram). New source grok-4-5-price-vs-benchmarks (Article, T3, the-decoder.com): Grok 4.5 at $2/$6 per 1M undercuts Opus 4.8 ($5/$25), GPT-5.5 ($5/$30), and Fable 5 ($10/$50); it trails on coding benchmarks (DeepSWE 53% vs Fable 5 70%, Terminal Bench 83.3%, SWE Bench Pro 64.7%) but xAI argues cost-per-task wins (claims 4.2× fewer tokens than Opus 4.8, ~80 tok/s) — the “good-enough + cheap” China/deepseek play from a closed US lab. Touched 4 pages: xai-grok (repositioned from premium to price-leader; added Grok 4.5 pricing/benchmarks/strategy), claude-fable-5 ($10/$50 rate corroborates Njenga’s ~3.2× anomaly — Fable is genuinely dear, likely a premium tier not a budget one), llm-api-pricing (frontier snapshot + the cost-per-task = price × tokens-per-task lever), and synthesis (xAI blurs the premium-vs-cheap split; cost-per-task axis; Fable-expensive now on two reads → moved from unverified anomaly to two-source corroboration in Contradictions). Still no T1 Anthropic price sheet for Fable 5 — reads remain third-party. Entity (The Decoder outlet) deferred. T3 (AI-news, specific numbers but not primary rate cards; token-efficiency figures are xAI’s). Ran avoid-ai-writing. +1 page.
[2026-07-08] ingest | OpenAI GPT-5.6 (Sol/Terra/Luna) vs Grok 4.5 clash (Business Insider) → llm-providers-wiki
Routed from hub (Telegram). Fetched via firecrawl (WebFetch hard-blocked by BI; paywalled subscriber page). New source openai-gpt56-grok45-clash (NewsArticle, T3, Hugh Langley, 2026-07-08). Substance: OpenAI GPT-5.6 rolls out wider Thursday (staggered at the Trump administration’s request) in three flavors — Sol (flagship; agentic coding/biology/cybersecurity), Terra (everyday), Luna (speed+affordability); Musk’s Grok 4.5 launches “in coming days,” billed “Opus-class but faster, more token-efficient and lower cost” (his own words = 2nd source on grok-4-5-price-vs-benchmarks‘s cost-per-task framing); Gemini 3.5 Pro later this month. Governance backdrop: Trump admin export-banned Anthropic’s Fable & Mythos (cyberattack-misuse risk) → Anthropic suspended → Fable 5 redeployed last week. Musk–Altman rivalry (Musk lost his 2024 suit). Touched 4 pages: openai (added GPT-5.6 Sol/Terra/Luna
- Luna token-economy), xai-grok (Musk’s framing + launch timing), claude-fable-5 (export-ban/suspend/ redeploy availability shock + Mythos), synthesis (efficiency now the shared closed-lab pitch; tiering goes intra-family; new export-control availability axis + ai-governance cross-spoke adjacency). T3 (reputable journalism; launch specs are pre-release vendor claims). Ran avoid-ai-writing. +1 page.
[2026-07-09] ingest | “Don’t use Claude Fable 5” (Ruben Hassid, Substack) → llm-providers-wiki
Hub-routed (Telegram). New source dont-use-fable-5-hassid (BlogPosting, T3, opinion; non-technical AI consultant translating docs). Value: a third independent read of $10/$50 (with grok-4-5-price-vs-benchmarks
- fable-5-vs-sonnet-5-token-cost), so the sticker is now well-corroborated (still no T1 Anthropic sheet). Adds two new facts: pay-per-use credits after 2026-07-12, and usage economics — cost scales with conversation length (whole thread re-read each turn: ≈$0.15 short / ≈$6 at 19 turns / ≈$14 at 40), reframing the open “which lever dominates net cost” toward turns/context, not just per-token rate. Buyer’s ladder: Fable 5 for hard goals at high/max effort (1–2 turns), Opus 4.8 as cheaper+smarter workhorse, Sonnet 5 not justified over Opus — a practitioner inversion of Anthropic’s mid-tier push. Touched claude-fable-5 (3rd read + usage economics + positioning; refreshed index one-liner), synthesis (Fable pricing bullet). Paged author ruben-hassid (Person). Cross-spoke: the tasks→goals + effort framing echoes agentic-tooling-wiki’s effort-level (linked). Ran avoid-ai-writing. +2 pages.
[2026-07-10] ingest | Mistral homepage — model lab → sovereign full-stack platform (hub-routed, Telegram)
Ingested mistral.ai (first-party site). New source summary mistral-platform (WebSite, source, T2 — first-party marketing, volatile) and refreshed the mistral-ai org page in place (its first update since the 2026-06-09 Wikipedia ingest). The story: Mistral now presents as a sovereign, EU-hosted full-stack platform, not just a model lab — product suite Vibe (work agent; the renamed Le Chat) / Vibe for Code / Studio / Forge / Compute over the model line; deployment self-hosted / EU-cloud / partner (AWS/GCP/Azure/IBM/SAP/Snowflake/NVIDIA/ Outscale); customers HSBC, ASML, Stellantis, EPO. New models named (no specs): Medium 3.5, Small 4, OCR 4, Voxtral TTS — named inline, not stubbed (proportionate to a spec-less marketing page). Synthesis: expanded the “European open-weight” position to “European open-weight → sovereign platform,” collapsing the gap to the cohere enterprise/sovereign row (Mistral reaches data-residency from the open-weight side). Cross-wiki adjacencies flagged: agentic-tooling (Vibe for Code / Studio agents), speech-audio (Voxtral TTS). No structural move → no verify. +1 page, 1 refreshed.
[2026-07-16] ingest | Inkling — Thinking Machines Lab’s open-weights model (hub-routed, Telegram)
Routed source: inkling-willison-review (simonwillison.net, 2026-07-16, T2). Pulled the first-party announcement + model card it links as a second source, inkling-announcement (T1), because the practitioner post carries the framing but not the architecture, benchmark table, or hardware numbers.
New pages (5): inkling (SoftwareApplication — Apache-2.0, 975B-A41B MoE, 45T tokens of text/image/audio/
video, 1M ctx, 66 layers, 256 routed + 2 shared experts, relative pos-embeddings not RoPE; Inkling-Small
276B-A12B in testing), thinking-machines-lab (Organization), tinker (SoftwareApplication — the
fine-tuning platform), plus the two source summaries. Refreshed open-weight-models and
simon-willison in place. No mira-murati page: leadership didn’t survive the evidence-only check
against the fetched text — left for a later source.
Synthesis: new position on the provider map — US open-weight, customization-as-product. Thinking Machines ships Apache-2.0 weights while stating outright the model is “not the strongest overall model available today, open or closed,” and monetizes fine-tuning on Tinker. No frontier tier to seed, unlike google‘s Gemma/Gemini dual track; nearest neighbour is Mistral’s Forge, reached from the opposite direction. Two new open questions: (1) is “open weight” still one category — Inkling wants 2TB+ VRAM (600GB at NVFP4) against Gemma 4 12B’s 16GB, same licence, ~125× the hardware, so only the local tier really attacks the price floor; (2) does free-weights-paid-tuning work as a business. Also fed the compatibility-as-strategy thread: a brand-new lab still ships an OpenAI-compatible API.
Caveats recorded: benchmarks self-reported at effort=0.99, no artificial-analysis re-run; training data documented only as “publicly accessible data repositories”; model card admits role-play/indirect-framing safety leakage. No structural move → no verify. +5 pages, 2 refreshed.
[2026-07-22] ingest | Mage-Flow-Turbo (Microsoft, Hugging Face)
Routed from the hub (route in ../log.md). URL-only, T3 — first-party model card, self-reported
benchmarks against a self-chosen (all-open) comparison set, no independent re-run.
New page: mage-flow-turbo (SoftwareApplication) — MIT-licensed 4B NR-MMDiT text-to-image, rectified
flow matching in a custom Mage-VAE latent space, 4-step distilled Turbo variant, native-resolution
packing 512–2048px at any aspect ratio, 0.59 s per 1024² on a single A100, GenEval 0.88, claimed
parity-or-better vs Qwen-Image 20B / Z-Image 6B / FLUX.2 32B. arXiv 2607.19064; diffusers + mage-flow
CLI + Gradio app.
Dedup: no mage/microsoft/flux page in wiki/. No microsoft Organization page created — a
canonical microsoft node already exists in agentic-tooling-wiki, so it is linked cross-wiki per the
bridge-node convention in CLAUDE.md rather than duplicated. No other entities paged.
Gap-relevance: feeds two live threads. (1) Generative-media sub-theme — this is source 2 of 3 and
the first non-Google one, the exact trigger condition gemini-omni-flash-nano-banana-2-lite wrote
down for spinning out an image/video-gen spoke; synthesis section updated, cluster tally recorded, not yet
spun out. (2) “Parameter count is the wrong axis” — 4B vs 20–32B rivals is agents-a1‘s claim in a
different modality, reached by co-design + step distillation instead of agent horizon.
Also noted: it adds no pricing unit (open weights, self-hosted), which sharpens the open-vs-hosted
split inside the generative-media cluster — Google’s tier has $/image, this has none. And the corpus has
no artificial-analysis-style independent referee for image models, logged as the sub-theme’s gap.
Verify deferred per hub policy (content-only). avoid-ai-writing run.
52 → 53 pages.
[2026-07-23] ingest | Fable 5 vs GPT-5.6 vs Kimi K3 for creators (Medium)
Routed from the hub (route in ../log.md). URL-only, T4 — Medium creator-opinion essay, benchmark
positions asserted with no scores/leaderboard/methodology.
New pages: fable5-gpt56-kimi-k3-creators (OpinionNewsArticle source) and kimi-k3 (thin anchor).
Kept the K3 claim recorded but did not mint a real model page off one weak source: kimi-k3 is an
explicit thin, unverified anchor (2.8T open-weight, weights due 2026-07-27, claimed #1 editorial writing
- edging Fable 5 on frontend coding) flagged for upgrade on the 27th. No Moonshot-AI org page created (delinked — no distinct evidence; Moonshot already appears inline across open-weight-models/ open-source-llms-2026/llm-provider). Dedup: Kimi was a recurring inline mention (K2.6, ~1.1T) with no page; linked the K2.6 mention in open-source-llms-2026 forward to kimi-k3 as the lineage successor. Gap-relevance: feeds the “does the frontier premium survive?” open question (an open model claimed at/above a closed frontier model on writing+coding) and restates the benchmark-vs-real-utility gap from a creator’s seat (cousin of dont-use-fable-5-hassid/fable-5-vs-sonnet-5-token-cost). Synthesis: added a flagged weak watch item under the frontier-premium question, not a finding. Entities: none paged (author a Medium byline; Moonshot AI evidence-thin here). Verify deferred per hub policy (content-only). avoid-ai-writing run.
[2026-07-23] ingest | Alphabet Q2 2026 earnings — Pichai’s remarks (Google)
Routed from the hub (route in ../log.md). URL-only, T1 for provenance (first-party CEO message);
all figures company-reported and unaudited here — dated snapshot.
Broad multi-spoke source handled per HUB “three+ spokes” rule: routed whole to llm-providers-wiki as the
dominant in-scope substance (the Gemini model lineup + Google’s model-scale metrics), with other facets
recorded as cross-spoke context inside the source page, not fragmented into other spokes.
New page: alphabet-q2-2026-earnings (BlogPosting source). In-scope core: Gemini 3.6 Flash, 3.5
Flash-Lite, 3.5 Flash Cyber (first domain-specialized Gemini logged here), 3.5 Pro testing,
Gemini 4 pre-training; consumer Omni video-in-app (+40% DAU). Scale: ~22B tokens/min (↑~38% Q/Q),
950M Gemini-app MAU, 9M+ developers, ~900M Gemma downloads (Gemma 4 ~300M), enterprise
token-concentration (500 customers >1T tokens/yr).
Updated gemini (Q2 lineup bullet incl. the Cyber domain-variant + Gemini 4 pretraining) and google
(new “Scale (Q2 2026)” section — the usage-volume view of the dual track). Synthesis: extended the
dual-track “Recurring reads” bullet with the usage-scale dimension — the market’s growth axis is now
consumption volume (the demand-side complement to the output-token cost thread), plus the tier/vertical
differentiation (cost ladder + Cyber).
Dedup: no Gemini-3.6/3.5-tier page existed; folded into the existing gemini anchor rather than a page per
model (dated-snapshot convention). Cross-spoke facets logged (see below), not paged.
Entities: none new — google canonical node updated; Pichai not paged (spoke convention: page the org/model,
not the exec byline).
Verify deferred per hub policy (content-only). avoid-ai-writing run.
[2026-07-24] ingest | Introducing Claude Opus 5 (Anthropic)
Routed from the hub (route in ../log.md). URL-only, T1 for provenance (first-party launch post);
the benchmark claims inside are vendor-self-reported and stated as ratios with no absolute scores, so
they’re treated as T3-grade evidence on a T1 document.
New pages: claude-opus-5-announcement (BlogPosting source) + claude-opus-5 (SoftwareApplication).
Core facts: claude-opus-5, $5/$25 per 1M — identical to Opus 4.8; fast mode at 2× price for ~2.5×
speed; effort levels written out (low/high/xhigh/max); pitched as “close to the frontier intelligence of
Fable 5 at half the price”; Frontier-Bench v0.1 >2× Opus 4.8, CursorBench 3.2 within 0.5% of Fable 5,
ARC-AGI 3 3× next-best, Zapier AutomationBench ~1.5×, OSWorld 2.0 past Fable 5 at ~1/3 the cost, organic
chemistry +10.2 pts. Available on API / Claude.ai / Claude Code / Claude Cowork; AWS, Vertex AI, Microsoft
Foundry named as partners. Opus 4.8 not deprecated — it’s the fallback for cyber-flagged Cowork requests.
Context window is never stated — logged as an open gap on both pages.
Updated claude-fable-5: the post is the first T1 source placing Fable 5 above Opus 5 on intelligence,
and “half of $5/$25” implies the $10/$50 rate the three third-party reads gave; tier field split
(T1 positioning / T3 rate).
Synthesis: three folds — (1) the frontier-premium question gains its strongest data point, a flat
generation-over-generation price and the lab discounting its own flagship from inside; Mythos 5
downgraded from “capability frontier” (the Sonnet-5 system-card reading) to a cyber/biology specialist;
(2) the Fable-5-pricing question gains a T1-implied sticker; (3) a third shape of selling speed —
a paid runtime mode on identical weights, next to Cerebras’ silicon and OpenAI’s Luna tier.
Dedup: no Opus-5 page existed; claude-opus-4-8 stays a research-wiki bridge node, not duplicated here.
Entities: none new — anthropic is the canonical cross-wiki node; no byline on the post.
Verify deferred per hub policy (content-only). avoid-ai-writing run.
[2026-07-26] ingest | gemma-4-31B-it-scotoma (ReadyArt) — abliteration of Gemma 4 31B-it
Routed from the hub (Telegram). Hugging Face model card; T3 — it is the official artifact page,
so the facts about what it is are primary, but every behavioural claim is the author’s unmeasured
self-report on their own model. No benchmarks, no evals, no dataset.
New pages: gemma-4-31b-it-scotoma (SoftwareApplication source), abliteration (DefinedTerm,
mechanism), readyart (Organization).
Core facts: Apache-2.0, 33B params BF16, base gemma-4-31B-it; method = locate refusal directions
with heretic abliteration → Jacobian-lens projection keeping ~22% of abliteration magnitude
→ merge at 1.5× scaling into BF16 across layers 7–41; 367 downloads/month; “loosened, not
lobotomized”; self-described research artifact.
Recorded contradiction: the card claims it “loosens gemma-4-31B-it’s cautious reflex” and that
“scotoma refuses basically as much as its base model.” Both kept, flagged on the page — a
refusal-removal technique disclaiming that refusal rates moved.
Updated gemma-4: a 31B-it size appears that no Google source in the corpus mentions — added to
the family table marked reported-not-confirmed (the 31B name vs 33B param count is the card’s own,
too). Also a new bullet: derivatives edit the safety layer, not just precision.
Updated open-weight-models: fourth structural consequence — a permissive license on downloadable
weights is permission to rewrite and redistribute, safety posture included; Apache-2.0 doesn’t
distinguish running from editing.
Synthesis: new thesis paragraph (“open weights have a third face: the behaviour is editable too”),
pairing this against claude-refusals-and-fallback — a closed refusal can be routed around, an open
one deleted. New open question: what is a safety claim worth on a downloadable model?
Dedup: no abliteration/refusal-editing page existed; gemma-4 refreshed in place, not duplicated.
Entities: readyart (publisher — in-domain in its own right as a model publisher; the llm-provider
typology has no slot for a train-nothing derivative shop, which is itself the finding). No author node —
the card carries no byline.
Cross-spoke: the Jacobian lens is the parked hub _inbox ai-interpretability stray (tally 1,
2026-07-10). This is sighting 2 and the first applied use — noted on the page, not parked; cluster
still short of the ≥3 spin-out trigger.
Verify deferred per hub policy (content-only). avoid-ai-writing run.
[2026-07-26] ingest | OmniRoute — the free MIT AI gateway (github.com/diegosouzapw/OmniRoute)
Routed here by the hub (runner-up: agentic-tooling-wiki, which owns the 33+ coding agents and the 104 MCP tools it exposes — but the substance is provider aggregation, quota and cost, i.e. this spoke’s API-access-and-pricing clause). T1 by origin (project’s own repo), with the standing caveat written into the page: every figure is self-reported — 290+ providers, ~1.53B free tokens/month, 89% average compression, and a self-scored comparison table against OpenRouter/LiteLLM. New: omniroute (source), ai-gateway (concept), diego-souza (entity). Updated: llm-api-pricing — the “model routing” cost lever now names the gateway layer, plus two new levers the corpus hadn’t recorded: client-side prompt compression and free-tier aggregation. amazon-bedrock reframed as the managed shape of the same layer. Synthesis: the reseller paragraph gains its self-hosted twin — the reseller commoditized model choice, the local router commoditizes the provider’s generosity. Two new open questions: how long free tiers survive being pooled, and whether compression becomes a provider feature or stays adversarial. Dedup: no gateway/router/OpenRouter page existed; the concept was only implicit in amazon-bedrock. The nearest prior art was the one-line “model routing” lever in llm-api-pricing, now expanded. Gap-relevance: gives the demand-side thread its second data point and its first non-cloud one. Verify deferred per hub policy (content-only). avoid-ai-writing run.
[2026-07-27] ingest | Opus 5 blows past Fable 5 and GPT-5.6 Sol on ARC-AGI-3 (The Decoder, 2026-07-26)
Routed here by the hub (runner-up: none — a benchmark result on the model market is this spoke’s “capability/cost benchmarks” axis; research-wiki owns Anthropic-as-substrate, not leaderboards). T3 tech-press reporting on ARC Prize’s analysis; the underlying scores would be T2 read direct. New: opus-5-arc-agi-3 (source), arc-agi (benchmark concept), arc-prize (entity). Updated: claude-opus-5 (first third-party numbers + the unreconciled 3× claim), claude-fable-5 (first non-Anthropic capability read on it — every earlier one was price), llm-benchmarks (novelty benchmarks as a second kind), synthesis (new tension section). What it adds:
- An independent referee. Nearly every capability figure in this spoke is the lab’s own. ARC Prize scores the model with no harness, publishes results/replays/code, and hedges its own finding (Kamradt: Opus 4.8 won some environments outright).
- A vendor ratio that doesn’t reconcile. Anthropic’s launch post says “3× the next-best model on ARC-AGI 3.” Against the 7.8% record that’s ~3.9×; against ARC Prize’s ~20% Fable-class figure it’s ~1.5×. Recorded as a tension, both numbers kept.
- The decay of a novelty benchmark. Opus 5 postdates ARC-AGI-3’s public format; the ~4× jump doesn’t reproduce on Ning’s private Witness (43.4, tied with Fable 5 and Kimi K3, and below Opus 4.8 on the one unfamiliar rule combination). The generalizable read is that a novelty test’s shelf life starts the day its format goes public. Entities deferred: Matthias Bastian (author), Greg Kamradt, Guanghan Ning, The Decoder (publisher — 2nd source from it after grok-4-5-price-vs-benchmarks; page it on a 3rd). Witness benchmark noted on arc-agi rather than paged; page it if it recurs. Verify deferred per hub policy (content-only). avoid-ai-writing run.
[2026-07-28] ingest | Sakana Fugu — Get Started (fugu.sakana.ai)
Routed here by the hub (runner-up: agentic-tooling-wiki — Fugu is a multi-agent system and installs
into Codex/Claude Code, but the document is API onboarding: keys, endpoints, model IDs, effort levels,
provider pools, pricing pointer. Same shape as deepseek-api-docs / claude-api-rate-limits,
which live here. The agent-system angle is cross-linked, not split).
T1 first-party docs for the interface; states no price, no benchmark, no context window.
Link arrived as a share.google redirect (second time) — resolved out of the share page’s HTML, then
probed for the canonical fugu.sakana.ai/get-started.
New: fugu-get-started (source), sakana-fugu (model), sakana-ai (provider — first Japanese
lab in the spoke).
Updated: ai-gateway (a third shape: the gateway that calls itself a model), synthesis (a new
section on “model” as a product boundary), index.
What it changes:
- A multi-agent system sold as a model. The docs concede it in their first sentence.
fuguroutes across all supported providers by default, with the pool narrowed per API key rather than per request — ai-gateway behaviour where the routing is the product and the model string is the label. Three of this spoke’s assumptions break: you can’t say which model answered (so a benchmark measures a routing policy), you can’t quote a per-token price for a named model, and the frontier-lab/cloud-reseller distinction in llm-provider collapses into one endpoint. - The SKU has stopped tracking the weights. Read beside claude-opus-5‘s fast mode (a runtime toggle on identical weights), the trend runs both ways — one product over many models, two products over one model. What the buyer purchases is a configuration, which is why capability and price claims now have to name what they measured.
- Cyber is a product axis across three labs — Mythos (claude-opus-5), Gemini 3.5 Flash Cyber
(alphabet-q2-2026-earnings),
fugu-cyber— none defining the category, and Anthropic’s cyber-capable line already drew an export ban (claude-fable-5). Two operational tells recorded: the effort dial starts at high (no low/medium;maxonly onfugu-ultra-v1.1), and the Codex provider block raisesstream_idle_timeout_msto 2 hours against Codex’s ~5-min default, which is the vendor telling you its turns run at agent scale. Also noted for../agentic-tooling-wiki: the shippedbase_instructionscarry agent-conduct guards (don’t run commands that would kill your own session; never force-kill by raw PID) — a provider shipping guardrails inside the model catalog, i.e. the instruction end of the axis nono argues must be enforced. Verify deferred per hub policy (content-only). avoid-ai-writing run.
[2026-07-28] ingest | Kimi K3 — Open Frontier Intelligence (MoonshotAI/Kimi-K3)
Refresh in place, exactly as the placeholder page asked for: kimi-k3 was a T4 anchor built on one Medium essay (fable5-gpt56-kimi-k3-creators) with the note “revisit and upgrade on the 27th.” The weights landed; the page is now T1 from the repo and technical report. No duplicate created. Updated: kimi-k3 (rewritten), open-weight-models (a third licence shape), fable5-gpt56-kimi-k3-creators (claims scored against the release), synthesis (the frontier-premium question), index. Three things worth keeping:
- A split decision against current flagships. 2.8T total / 104B active (16 of 896 experts), 93 layers (69 Kimi Delta Attention + 24 Gated MLA), 1M ctx, native MXFP4 weights via QAT, MoonViT-V2 vision. Ahead on GPQA Diamond (93.5 vs Opus 4.8’s 91.0), SWE-Marathon (42.0 vs Fable 5’s 35.0), BrowseComp and Video-MME; behind on DeepSWE (67.5 vs 70.0). Trading wins with the closed frontier is a harder data point than glm-52 or agents-a1, which approached it.
- A vendor table that documents its own confounds — different harnesses per model (K3 in Kimi Code, rivals in Claude Code/Codex/Terminus; the harness alone moves K3’s DeepSWE 67.5→67.3), a modified H20-calibrated SWE-Marathon branch, BrowseComp needing context compaction (91.2 with, 90.4 on the raw 1M window), and an explicit note that Fable 5 hit fallbacks on 35% of tasks in their run, i.e. caveating their own win. Above the disclosure standard of claude-opus-5-announcement‘s ratio-only claims, still first-party.
- The licence tightened as the capability closed. K2.6 was Modified MIT; K3 ships a bespoke licence, MIT-shaped except for a Model-as-a-Service revenue threshold (US$20M over any 12 months) requiring a separate agreement — Llama’s move, aimed at clouds and API resellers rather than users. open-weight-models‘s licensing axis now has three shapes, and the synthesis records the two-sided answer: capability is competed away from below faster than openness is. Calibration recorded on the T4 source: it got the specs right (2.8T, open weights, the date) and the rankings wrong (“#1 editorial writing” is absent from the release; “edged Fable 5 on coding” is half true). Specs propagate; comparative judgments don’t. Verify deferred per hub policy (content-only). avoid-ai-writing run.
[2026-07-28] ingest | DeepSeek V4 vs GLM-5.2 vs Qwen: 10x Price Gap (tech-insider.org)
Routed here by the hub. T4 — and the tier is the story. Affiliate leaderboard ad (trovedrops
with sub_id tracking), Google Preferred Sources widget, and a sibling article using the identical
X vs Y vs Z: $N Price Gap [2026] template on a different model trio. A monetised content template,
not an analysis anyone needed to write. No methodology on any figure, no benchmark linked to its
scoreboard. WebFetch 403’d; retrieved via Firecrawl per HUB edge handling.
New: deepseek-glm-qwen-price-gap (source). Updated: glm-52 (contested-pricing section),
synthesis (new “Whose price is on the sticker?” section), index. No new entities — deepseek,
qwen, z-ai all already paged.
The value of ingesting a T4 source into a spoke that already holds T1 and T2 on the same subjects is
that it can be graded, so that’s what the page does.
Corroborated: GLM-5.2 at 753B MoE / ~40B active, 1M context, MIT — matches glm-52 exactly
(T2, Willison on first-party facts). deepseek-v4-flash / deepseek-v4-pro and context caching match
deepseek-api-docs (T1, vendor’s own reference). The spec sheet holds up.
Contradicted, minor: GLM-5.2 release date given as 2026-06-13; glm-52 has 2026-06-16.
Contradicted, and it’s the headline: the article calls $1.40 / $4.40 GLM-5.2’s vendor list
price and reports third-party resale separately at ~$0.55 / $1.85. glm-52 attributes
$1.40 / $4.40 to OpenRouter hosts — i.e. resale. The “10x price gap” to DeepSeek V4 Flash
($0.14 input) is computed off $1.40 being Z.ai’s own rate; if Willison’s attribution is right, the
comparison sets a vendor price against a reseller price and calls the difference a vendor gap.
Recorded as an open conflict — neither source is Z.ai’s pricing page, and nothing here resolves it.
Unverifiable: 93.5% LiveCodeBench, +6.7 SWE-bench Pro, +17 long-horizon agentic — no scoreboard,
no date. The one independent ranking held here points the other way on overall capability
(artificial-analysis Intelligence Index v4.1: GLM-5.2 51, DeepSeek V4 Pro 44). Different
benchmarks, so not a flat contradiction, but a reader taking the article’s framing as a general
capability claim would be misled.
What survives the tier: it’s a snapshot of how the open-weight price floor is marketed (three Chinese labs, two MIT + one Apache-2.0, all pitched on undercutting closed APIs — the open-weight wedge showing up in SEO copy), and its resale-layer observation is real and thin in this corpus. Folded into synthesis as the structural point: open weights make a model’s price a market, not a number, so “what does GLM-5.2 cost” has no single answer and comparisons against single-vendor closed models aren’t like-for-like. That sits next to the tokens-per-task multiplier already in llm-api-pricing — two separate ways a headline per-token rate fails to predict a bill. Verify deferred per hub policy (content-only). avoid-ai-writing run.
[2026-08-02] ingest | Introducing the Neutrino-1 models (Fermion Research)
Routed here by the hub from Telegram. T1 — first-party primary release post, dated 2026-07-27.
Runner-up spokes: llm-inference-wiki (real claim, see below) and machine-learning-wiki.
Pages. neutrino-1 as the source summary (SoftwareApplication, the announcement-as-model-page
pattern this spoke already uses for cohere-north-mini-code and agents-a1), plus a new
mechanism page ternary-weights and the lab entity fermion-research. Updated quantization
(a third category beside PTQ and QAT), qwen, index and synthesis. No authors are
named on the release, so the lab is the only entity; entity-index had no match for it.
Why it needed a new mechanism page. quantization framed low precision as PTQ-or-QAT, and both start from a higher-precision model. Fermion’s claim is that there was never one — “no full-precision product model that was rounded afterward” — with rounding a trained model to the same depth said to land near chance. That is a training result wearing a compression result’s clothes, and it does not fit inside the existing page. Ternary also now has three named models in the corpus (BitNet b1.58-2B4T, Ternary-Bonsai-8B, Neutrino-1), two of them known only through a competitor’s table.
Two things recorded rather than smoothed over.
Every number is Fermion’s. The comparison table re-runs each rival on Fermion’s own stack. The
disclosure is unusually good — it says when its number differs from a rival’s card, prints its own
losses on IFEval (77.2 vs Gemma-4-E4B’s 88.26) and BFCL (68.9 vs Llama-3.1-8B’s 76.1), and labels the
GSM8K cell where its zero-shot no-CoT run sits beside Llama’s 8-shot-with-reasoning 84.5. It is still
a table one party controls end to end, with no independent evaluation anywhere. The single
externally-anchored measurement is the engine test on BitNet’s public weights (102.4 tok/s vs
bitnet.cpp’s 89.0, same machine, same session) — the only claim where Fermion does not own both
sides.
The architecture is Qwen3’s, and the post does not say so. Checked against
Qwen3-8B/config.json during the ingest: 36 layers, 4,096 hidden, 32 query / 8 KV heads, 128-dim
heads, 12,288 FFN, 151,936 vocab, 40,960 positions, untied embeddings, matching parameter count —
every stated field. The 0.6B matches Qwen3-0.6B including tied embeddings. Recorded on
neutrino-1 with both readings and no verdict: reusing a published Apache-2.0 geometry is
normal, legal, and arguably the correct experimental control when the weight format is the variable
under test; and a post detailed enough to publish per-layer zero density never naming it is a real
omission. synthesis carries it as a new tension — the architecture layer of the same erosion
already logged for model strings and SKUs.
Cross-spoke, noted not split. llm-inference-wiki owns the KV-cache arithmetic, the
bandwidth-bound decode argument and the drafted-decode mechanism (27,648 tokens verified with zero
divergences); this spoke owns the format as a market lever — the standing quantization dual-lens
split. machine-learning-wiki owns §4, which is a genuine post-training finding: behavioural
training pressure degrades specific other axes rather than uniformly, protection works only for
axes explicitly represented in a stage’s batches, and dose curves differ per axis (tool calling
saturates ~25 steps, instruction following ~50).
Content-only, no page moves — verify deferred per the hub’s standing policy. avoid-ai-writing run over the new prose.
[2026-08-03] ingest | Astra — canonical model node from the ten-proofs announcement
Follow-up to the hub’s ten-proofs route earlier today. That source went to ../research-wiki
(theorem proving is its cluster E), and the router deliberately did not page Astra there to avoid
duplicating a model node outside the spoke that owns the model market. The human asked for the page
here, which is where it belongs.
New page: astra (SoftwareApplication, T1 — OpenAI’s own publication, read via firecrawl after a 403). Updated: openai (model lineup), index, synthesis (new open question).
What’s actually known is thin, and the page says so. OpenAI named Astra on 2026-08-01 in a mathematics paper, not a model announcement: an internal version produced ten results in maths/TCS, humans plus the model wrote the manuscripts, the model wrote the Lean. No params, architecture, context, modality, date, availability, pricing, safety documentation, or benchmark table. Whether the internal version is the one that ships is unstated, as is its relation to GPT-5.6.
The one number needed unpacking. ”~$2,000 at Sol API rates” prices an unreleased model’s run in the shipping flagship’s currency (openai-gpt56-grok45-clash — Sol is the GPT-5.6 flagship). It is a token-volume proxy wearing a dollar sign, not a price; Astra has no rate card. Recorded that way, because the obvious misreading (“ten open problems for $2,000”) is a price for a product that can’t be bought. Still the most useful cost datum the spoke holds on a frontier result — an order of magnitude for the inference budget behind one, which llm-api-pricing‘s cost-per-task thread had no figure for.
New open question opened. Astra’s public evidence is machine-checkable proofs rather than a vendor eval — the first claim shape here that a stranger can falsify. But only the output is falsifiable; the capability isn’t (attempts, steering, and shipped-version parity are all unverifiable). Watch whether other labs copy the form or whether it stays confined to domains with formal checkers.
avoid-ai-writing run. Verify deferred per hub policy (content-only, no page moves).
[2026-08-03] lint | close the Fable 5 price-sheet question — it was answered in a sibling spoke
This spoke’s claude-fable-5 page ended with “What’s still missing: an actual Anthropic price
sheet… The $10/$50 figure remains inferred or third-party.” That was false, and had been since before
the page existed. Anthropic’s 9 June 2026 launch announcement states “Pricing (both): $10 / $50 per
Mtok” outright; the hub holds it as research-wiki/claude-fable-5-mythos-5-announcement, created
2026-06-10 — a month before this page (2026-07-05).
Corrected the page, its tier: line (the rate is T1, not T3-inferred), the synthesis passage and the
index entry. Nothing was retracted: the four-source triangulation (Njenga → The Decoder → Hassid →
the Opus 5 arithmetic) stands as independent corroboration that landed on the published number exactly,
which is a stronger result than either the sheet or the reads alone. Only the “no first-party rate
exists” claim was wrong. The 2026-07-24 paragraph is marked superseded in place rather than rewritten.
Still genuinely undisclosed by Anthropic: context window, max output, intended tier. The caching half
of the cost question is untouched and still rests on Njenga’s single unrepeated run.
Systemic finding, logged to the hub: the answer sat in a sibling spoke for two months and this
spoke never looked. Entity dedup has npm run entity-index to prevent exactly this for entities;
claims have no equivalent.
[2026-08-04] ingest | Anthropic platform model docs — and a correction to our Opus 5 pages
The Claude Opus 5 launch post arrived again via Telegram. Dedup
first: already held, ingested 2026-07-24 the day it published, freshness: volatile and only 11
days old — inside the 60d window, so not stale and not due for refresh. Re-ingesting it verbatim
would have produced nothing.
So the useful move was the cross-check, not the re-read. Pulling Anthropic’s platform documentation — the models overview and the effort reference — against what the announcement says turned up one error and several gaps. New page: anthropic-model-docs (T1).
The correction. The launch post lists the effort levels as “low, high, xhigh, and max.” Both our
Opus 5 pages copied that. The effort docs say verbatim: “Claude Opus 5 supports all five effort
levels.” The missing one is medium — and the same page tells developers to use low and
medium “liberally” as the primary control for token cost and response time. Corrected on
claude-opus-5; flagged in place on claude-opus-5-announcement rather than rewritten, per
record-don’t-overwrite, since the quote is what the post actually says.
Worth naming the direction: the marketing copy dropped a cost-saving option, not a capability claim. That is the opposite of the failure mode this spoke usually watches for.
Gaps the announcement left, now recorded: 1M context (default and maximum), 128k max output
synchronously and 300k on the Batch API, effort defaulting to high, and availability on
Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry — four channels the post
never mentions.
The find worth carrying into synthesis: Opus 5’s reliable knowledge cutoff is May 2026, against January 2026 for both claude-fable-5 and claude-sonnet-5. The tier Anthropic sells as below the frontier has four months more world knowledge than the tier above it. Every comparison this spoke holds is capability-per-dollar; none of them price recency.
Two API facts with product consequences: thinking is on by default and cannot be disabled at
xhigh/max (400) — a change from claude-opus-4-8; and effort moves thinking volume, not
visible response length, so it is not the verbosity lever it looks like.
Standing lesson: this spoke’s Anthropic coverage is built almost entirely from launch posts and press. Those are right for positioning and wrong for facts a developer acts on. The docs are one page deeper on the same domain and were never read.
[2026-08-05] lint | freshness regrade (quality cycle)
gemma-4-qat and google-ai-updates-may-2026 volatile → stable, both 61 days past the 60-day
window. Both are dated Google announcements: a BlogPosting records what the vendor said on that date,
and re-reading the URL returns the same text, so the flag names re-verification work that can never be
done. The second is explicit about it — it is a May 2026 roundup. Whether the lineup they describe
still holds is a synthesis.md question, not a freshness one. Rule already in ../QUALITY.md.