Spokes.wiki Search About
Blog Posting source ↗ source url updated Tue Jul 14 2026 00:00:00 GMT+0000 (Coordinated Universal Time)

Cognition trusts Claude Fable 5 to work through the night

An Anthropic customer case study on Cognition — the company behind Devin, an autonomous AI software engineer — and why it routes its long-running agent work to Claude Fable 5. The headline claim: Fable 5 lets Devin run 8+ hours unattended (“through the night”) on real tasks — migrations, bug backlogs, delayed features — without the coherence drift that made earlier models lose the thread “after minutes to an hour.” First-party promotional framing, so T3; the numbers below are Cognition’s own.

What “trust” means here — and why it’s interesting for this spoke

The load-bearing detail is how Cognition decides to trust a model, because it speaks straight to this spoke’s standing Benchmarks gap. Their internal rule is “we trust no eval”: rather than leaderboard scores, senior engineers test a model on real work and ask whether the code would actually ship. They built a proprietary “Frontier Code” benchmark as an explicit “anti-slop” standard — designed to reward production-ready code over code that merely passes tests (the exact failure mode of public benchmarks). On it, Fable 5 scored ~30% against the prior Opus model’s ~10%, and — the part SVP of Research Silas Alberti called unusual — the internal dogfooding agreed with the metric. So this is a rare case of a practitioner benchmark with a stated methodology, not another README claim.

The two capabilities they name

  • Horizon (self-sufficiency). Alberti: “The biggest thing we noticed was the horizon, how long it can be self-sufficient” — he could delegate a task and sleep while the agent worked productively for eight hours. This is the long-running axis stated as a model property, not a harness feature.
  • Verification discipline. On complex migrations Fable 5 “stated the invariants it would maintain, executed against those constraints,” and during incident triage root-caused rather than confidently guessing — plus it used Cognition’s internal tools to page through noisy logs and articulate what it didn’t understand. That transparency (naming its own uncertainty) is what rebuilt engineer trust.

Where it sits

Two threads. It’s a production instance of the spoke’s loop/verification thesis (loop-engineering, agent-loops-verification) — invariants-then-execute-then-verify is exactly the “feedback signal” discipline, here run for eight hours unattended, the enterprise cousin of Karpathy’s overnight autoresearch loop. But it also cuts against the spoke’s “structure substitutes for capability” grain: Cognition’s story is that a raw model step-change — Fable 5’s horizon, “the kind that come roughly once a year” — was the unlock for autonomy, not a cleverer harness. That tension is the useful part (see synthesis Open questions / Contradictions). Cross-spoke: the model itself (claude-fable-5, its market/capability standing) is llm-providers-wiki’s; here the subject is the agent and the trust/verification practice.

cognition · devin · silas-alberti · claude-fable-5 · anthropic · loop-engineering · agent-loops-verification · durable-agents · autoresearch · effort-level