Log — Platform Ops Wiki
Append-only history. Each entry starts with ## [YYYY-MM-DD] <op> | <title> where
<op> is ingest, query, lint, or split, so grep "^## \[" log.md | tail -5 works.
[2026-06-05] split | platform-ops-wiki created from _inbox cluster (3 sources)
Spun out by the hub router when the InfoQ piece netflix-service-topology arrived
via Telegram — the third tight ops piece, hitting the ≥3 spin-out threshold for the
platform-ops-sre cluster (the google-sre-agentic-ai park note had explicitly
flagged that a 3rd would trigger a spin-out). Scaffolded from CLAUDE.template.md;
domain = production platform engineering, SRE & observability for cloud-native
distributed systems. Migrated and ingested all three (URL-only, source: true + url:):
- netflix-service-topology → concepts service-topology, observability
- google-sre-agentic-ai → concepts site-reliability-engineering, aiops
- kubernetes-integration-tax → concepts platform-engineering, kubernetes
Created 6 concept pages (DefinedTerm) + 1 SoftwareApplication (kubernetes) + 3 source
summaries (10 pages total). synthesis frames the three as facets of one thesis: in
production the hard problem is the seams, not the components — telemetry-source fusion
(Netflix), platform-tool integration (CNCF), and investigation-signal integration
(Google SRE), meeting at the service-topology.
Did not migrate
nvidia-doca-in-silicon-security(ai-infrastructure) — distinct hardware/silicon-security layer; left parked with adjacency noted. Cross-spoke split fromagentic-tooling-wiki(builder tools vs. their ops application) recorded in synthesis. Open questions: build-vs-buy topology, AIOps reliability, eBPF cost, quantification.
[2026-06-05] lint | first health check (10 pages, day-of-spin-out)
Swept the new spoke for orphans, thin spots, @type specificity, missing cross-links, and contradictions. Findings + actions:
- Orphans: none — every page has ≥3 inbound links (min kubernetes = 3).
- Thin spots: none — smallest page kubernetes (969 B) is a legitimate substantive anchor.
- Contradictions / stale claims: none — the 3 founding sources are complementary; all pages authored same-day, nothing stale.
- Missing cross-links (2, fixed): observability named SRE/agents in prose without linking → added inbound to site-reliability-engineering + aiops; kubernetes referenced the “platform-ops practice” without linking → linked platform-ops (and dropped a stray game-engine cross-spoke mention in its adjacency note).
- @type (1, left as-is, optional): netflix-service-topology is typed
TechArticle; it is an InfoQ /news/ item, soNewsArticleis a defensible lateral alternative — but a sibling of TechArticle, not strictly more specific, so not changed. Flagged for a human call only. Site rebuilt clean; link check + count check PASSED (10/10).
[2026-06-09] ingest | +3 observability substrate (OpenTelemetry, eBPF, Prometheus) — all-spokes cron test
Answered the “build vs buy the topology” + “eBPF operational cost” open questions with the off-the-shelf CNCF stack: opentelemetry (SoftwareApplication, src — vendor-neutral traces/metrics/logs standard, not a backend), ebpf (DefinedTerm, src — sandboxed in-kernel programs; verifier/JIT/maps; CAP_BPF + complexity + kernel-version costs), prometheus (SoftwareApplication, src — pull-based time series + PromQL; 2nd CNCF project; not billing-grade). Tied to netflix-service-topology/kubernetes-integration-tax. Synthesis open questions updated (gap remaining: the topology-graph assembly above the raw signals). url-only. 10 → 13 pages.
[2026-06-10] ingest | SLOs + GitOps + distributed tracing — all-spokes pass (one foundation per pillar)
Three foundational concepts the spoke referenced but never paged. service-level-objectives (DefinedTerm, source, Google SRE book) — SLI/SLO/SLA + error budgets; the quantification backbone the open questions wanted (reliability-vs-velocity as a measured control loop; toil/MTTR become budget math), and a reframe of the AIOps reliability paradox (agents under an error budget). gitops (DefinedTerm, source, OpenGitOps/CNCF) — Git as single source of truth; four principles (declarative, versioned-immutable, pulled, continuously reconciled; Argo CD/Flux); the deployment face of “seams, not components” and a structural cousin of the aiops control loop. distributed-tracing (DefinedTerm, source, OpenTelemetry) — spans/traces/context-propagation; the per-request view of the service-topology (topology ≈ traces summed over time), the third signal beside metrics/logs, one of Netflix’s three fused telemetry sources, and the source of latency SLIs. Together they close a loop: observability(tracing) → SLIs/SLOs → reconcile/operate(GitOps/AIOps). Folded into synthesis (new 2026-06-10 section) + index (3 DefinedTerm rows). No contradictions. 13 → 16 pages.
[2026-06-12] ingest | DORA metrics (Four Keys) — dora.dev
All-spokes daily expansion. Added dora-metrics (@type DefinedTerm) — the delivery-performance quantification framework completing the “Quantification” open question that service-level-objectives half-answered. SLOs measure the running service’s reliability; DORA measures the delivery pipeline: throughput (deploy frequency, change lead time) + stability (change fail rate, failed-deployment recovery time, deployment rework rate). Captured the “speed and stability are not tradeoffs” finding and the MTTR→“Failed Deployment Recovery Time” term shift. Wired to service-level-objectives (backlink) / gitops / aiops (gives the reliability-paradox a yardstick); synthesis note added; open question reframed (frameworks named, still want them applied to this spoke’s own MTTR/toil claims). 1 new page. Authoritative (Google DORA / State of DevOps). No contradictions.
[2026-06-12] ingest | Project-as-a-Service (Belastingdienst, InfoQ/KubeCon)
Telegram drop, routed → platform-ops-wiki (platform-engineering pillar). Added source project-as-a-service
- new concept internal-developer-platform (IDP / golden paths / platform-as-a-product). Completes the
platform-engineering pillar: kubernetes-integration-tax was the problem side; the IDP is the cure —
pay integration once centrally, expose it as a self-service golden path (one YAML → namespaces/RBAC/quota via
the
opr-paasoperator; “make the right way the easiest way”). GitOps-shaped reconcile applied to project provisioning; half-social (enablement>support, Communities of Practice across 99+ teams, accelerator hackathons). Updated platform-engineering (“the cure: productize the platform”) + synthesis + index. 2 new pages. Caveat: qualitative only (no MTTR/onboarding numbers — quantification gap); golden-path→golden-cage tension recorded. (Routed pre-git; no commit yet — git is step 0 of the pending quality-mechanism plan.)
[2026-06-18] ingest | DefinedTerm enrichment pass (fetched canonical sources)
Hub-wide DefinedTerm enrichment (user-directed, “go deeper, fetch sources”). Added two first-party definitional sources the spoke lacked: google-sre-book (T1, sre.google) and otel-observability-primer (T1, opentelemetry.io). Deepened site-reliability-engineering with the canonical practice (Treynor’s 2003 definition, error budgets w/ the 99.99%⇒0.01% math, the 50% toil cap, ~70%-of-outages-from-change, on-call ≤2/shift) and observability with the OTel definition + monitoring-vs-observability contrast + telemetry signals/instrumentation. Light: aiops gains the “operates within the SRE backbone” link + Related footer. Related footers added (SRE/observability/aiops). No fabricated facts; AIOps’ Gartner-origin left out (IBM source 403’d, no clean citation this pass). +2 source pages (→ index updated).
[2026-06-19] ingest | entity enrichment — Kubernetes grounded
Extending enrichment beyond DefinedTerm to entity pages. kubernetes (SoftwareApplication) was a 22-line stub with only the integration-tax framing; grounded it with kubernetes-docs-overview (kubernetes.io, T1): the canonical definition, Google/Borg 2014 lineage, K8s etymology, the control-loop capabilities (self-healing, rollouts, bin packing), and the “what it is NOT” (no built-in monitoring/ PaaS) — which is exactly why prod K8s is an integration problem (kubernetes-integration-tax). New source page; index updated; Related added.
[2026-06-20] ingest | webernetes (browser K8s simulator) — hub-routed
Telegram drop (hub msg 543, github.com/ngrok/webernetes). ngrok’s TypeScript Kubernetes simulator that runs client-side — no real cluster; simulates Pods/Services/Deployments/ReplicaSets/controllers + rolling updates/probes for education/visualization. Experimental (326★). New source webernetes (SoftwareSourceCode, T1 first-party, but a self-described tech demo). Boundary noted: filed here because Kubernetes is the spoke’s core subject and no spoke is closer, but it’s a teaching artifact — the inverse of the spoke’s production-ops lens; useful as the concept-onboarding companion to kubernetes/ kubernetes-docs-overview, not evidence about prod behaviour. +1 page.
[2026-06-23] ingest | eBPF safe kernel observability (InfoQ podcast, Dan Fineran/Isovalent)
New source page ebpf-kernel-observability-infoq (T3 — InfoQ practitioner podcast; Isovalent/Cisco guest, vendor-adjacent). Practitioner colour on ebpf past the ebpf.io primer: the verifier as “bouncer on the door” (safety), observability without instrumentation (kprobes/uprobes/tracepoints, no app changes), and crucially the security use case — Tetragon with “front-foot” pre-syscall enforcement enabling live CVE patching, plus the Cilium/Tetragon-abstraction takeaway (“you don’t need to know eBPF to use it”). Gap-relevance: opens the AIOps seam — AI-generated eBPF policies for auto CVE mitigation (self- healing pushed into the kernel), folded into synthesis under the AIOps reliability-paradox thread (generated policies “often contain specification errors” → the agent’s guardrail needs its own guardrail). Refreshed ebpf (new security/live-patching section; +aiops link; bumped updated) and the AIOps synthesis point. Index updated. Cilium/Tetragon/Isovalent/Fineran noted inline (no thin nodes — spoke has no Person/Org pages). +1 page (platform-ops 23→24). Fetched live. (No build/commit here — hub handles.)
[2026-06-23] ingest | Awesome Microservices (mfornos) — curated landscape catalog
New source page awesome-microservices (T3 — community “awesome list”, CC0, 14.4k★; a maintained link directory, map-not-analysis). Despite the “architecture” framing it’s infra/ops-weighted: its Capabilities taxonomy (discovery, orchestration, monitoring/logging, messaging, gateways, security, storage, testing) + CI/CD + org-design is exactly the platform-ops concern list. Routed here as the landscape map under the spoke’s deep-dives. Gap-relevance: grounds the central “the hard problem is the seams, not the components” thesis with the actual ~400-component inventory the integration tax is levied on — folded into synthesis at the seams paragraph. Linked existing nodes (kubernetes, observability, distributed-tracing, prometheus, internal-developer-platform, platform-engineering); no new concept page for a catalog source. Index updated. +1 page (platform-ops 24→25). Fetched live (via share.google→github). (No build/commit here — hub handles.)
[2026-06-30] ingest | Scaling to 1M Lambda functions (AWS/ProGlove) + serverless anchor
Routed from Telegram. New source scaling-to-1m-lambda (TechArticle, T3, AWS Architecture Blog;
Freiberg/Blank, 2026-06-29) — ProGlove scales a multi-tenant SaaS to 1M Lambda functions / thousands of
AWS accounts. The spoke’s first serverless-ops-at-scale source, so also added the missing
serverless concept anchor (the per-invocation compute model; the spoke’s 2nd production substrate beside
kubernetes). Lessons folded into synthesis as a new facet of the “seams, not components” thesis:
account-per-tenant isolation (blast-radius via independent concurrency/throttle limits); the self-DDoS
from synchronized rate(5 minutes) crons → jitter (“never do the same thing at the same time
everywhere” — reflexively the hub scheduler’s own anti-synchronization rule); operational cost as a new
dimension (kill SQS polling, centralize DLQs, observability $3→$0.70/account); StackSets deploy ceilings →
custom EventBridge/Step Functions tracking; monorepo governance (cross-links dev-tooling’s monorepo).
Takeaway: “efficiency, not capacity.” Entities (AWS, ProGlove, authors) deferred per spoke posture. Ran
avoid-ai-writing. +2 pages (→27).
[2026-07-11] ingest | OpenObserve (open-source Datadog alternative) — hub-routed from Telegram
Telegram drop (hub msg 853, Medium/Code Coup). New source openobserve (SoftwareApplication, source, T3 — secondary tech blog grounded on first-party vendor facts from openobserve.ai). The spoke’s first observability backend product — a unified open-source store (logs/metrics/traces/ RUM/session-replay) pitched as a self-hostable Datadog/Splunk/New Relic/Grafana-LGTM/ELK alternative. The substantive part is the architecture, not the pricing: Rust + columnar Apache Parquet on object storage (S3/MinIO/GCS/Azure/local) + DataFusion querying Parquet directly, ~40× compression → a vendor-claimed 140× lower storage cost vs Elasticsearch (and 8–10× lower total vs Datadog by dropping per-host/per-user pricing). OTel-native ingest; SQL + PromQL query; single-binary→Helm/HA gradient; ships an AI Assistant (NL→query) + AI SRE Agent (RCA). Vendor scale: 6,000+ orgs, up to 2 PB/day, 1 PB queried in ~2s (self-reported → T3). Gap-relevance: fills the backend/storage slot the observability pillar never had — signals were described (observability) and produced (opentelemetry) but never stored/queried. Reframes the pillar: observability at scale is an OLAP problem (columnar-on-object-storage vs inverted- index-on-hot-disk = the log-bill gap). Loops to aiops (AI SRE Agent = agent-over-telemetry, same shape/paradox as google-sre-agentic-ai). Refreshed observability (new “where the signals land” section + Related), folded into synthesis (new paragraph + analytical-databases-wiki cross-spoke seam: Parquet/DataFusion = ClickHouse/DuckDB OLAP playbook; runner-up spoke, cross-linked not split). Index updated (SoftwareApplication row). Ran avoid-ai-writing. url-only. 27 → 28 pages.
[2026-07-14] ingest | Modernizing the Meta Ads Service with an Open-Source Kernel Scheduler (Engineering at Meta, hub-routed, Telegram)
Ingested Meta’s engineering blog on using sched-ext (BPF extensible scheduler, upstream Linux v6.12) to write a custom CPU scheduler for the ads-serving fleet (5M req/s). Problem: kernel v6.9’s EEVDF default scheduler regressed latency (fewer ads ranked) → hosts stranded on v6.4 (fragmentation/tech-debt). Fix: a sched_ext policy soft-partitioning CPUs (latency-critical vs best-effort pools) with thread-importance domain knowledge; runs as a user-space binary loading a BPF program, so a change is a process restart, not a kernel reinstall. Results: −28% p99, +1.1% weighted-ads-ranked, 3.28 MW saved; follow-on userspace policy iterations added −60% latency / −18% timeouts. New pages: source meta-sched-ext-ads (TechArticle, T2 — first-party/self-reported); new mechanism node sched-ext. Updated ebpf: added scheduling as a fourth eBPF domain (beyond observability/networking/security) — eBPF as a general kernel-programmability substrate, not just telemetry. Synthesis: extended the serverless/scheduling thread (“scheduling as a production lever, now at the kernel”); operational-agility (restart-to-deploy) rhymes with gitops; verifier supplies the untrusted-in-kernel safety; hardest production numbers yet (dents the qualitative-only caveat). Followed spoke convention (no publisher/author nodes — Meta inline). avoid-ai-writing self-pass (clean). +2 pages, 2 updated.
[2026-07-17] ingest | Scaling to 1 million concurrent sandboxes in seconds (Modal, hub-routed, Telegram)
Ingested Modal’s engineering blog on the scheduler rewrite behind its Sandboxes — 0 → 1M concurrent
sandboxes in under a minute (old ceiling: 50k/customer), claimed median start <0.5s, scheduling latency
in tens of ms, tens of thousands created/sec. The design move is the story: “traded global
consistency for scalability” — no data store on the creation path, workers publish state async to a Redis
stream (worker-as-source-of-truth), horizontally-scaled scheduling servers load-balance off a stale
in-memory view, direct RPC to the worker (which accepts/rejects), batched control messages. Two network
hops + one cheap CPU op. Explicit kubernetes critique: O(n × p) scoring serialized by default, etcd
write bottleneck, O(nodes) heartbeat baseline. Honest bits: the real wall was the Linux rtnl lock
serializing container-network setup (tens-of-seconds stalls → had to rewrite sandbox networking), and the
single unsharded Redis stream (load-tested to “well over 100k workers,” sharding deferred). Beta, opt-in.
New pages: source modal-1m-sandboxes (TechArticle, T3 — vendor writing about its own product, every
number self-reported, no reproducible method; the architecture is durable, the multipliers are marketing);
modal (SoftwareApplication — paged as a published scheduler design, not a vendor to price-compare);
compute-sandbox (the spoke’s third substrate after K8s/serverless: untrusted + ephemeral + burst);
container-scheduling (the placement layer; consistent-view vs stale-view axis; separates the
arrival / placement / CPU altitudes that the scheduling thread had been conflating).
Updated kubernetes (new “scheduler as a scaling limit” section) and serverless (the sandbox variant —
isolation boundary moves account → single execution).
Synthesis: three new paragraphs. This is the corpus’s first genuine argument — the spine (gitops
reconcile, K8s control loop) assumes a consistent view is what makes production tractable; Modal names it as
the ceiling. Recorded as a flagged contradiction (both claims kept; reconciling read = different workloads,
held as an unsourced hypothesis; noted Modal sells the alternative). Also flagged the two-”1M”-stories
pairing: scaling-to-1m-lambda‘s never do the same thing at the same time everywhere (jitter) vs Modal’s
a million at once, on purpose — same fleet physics, opposite ends. The rtnl find rhymes with
meta-sched-ext-ads one layer down (the queue you remove reappears below).
Gap-relevance: opens a new open question (who operates a sandbox fleet? — corpus has creation, not
running: observability over 1M second-lived units, SLOs on a burst substrate, isolation failure; no incident/
postmortem/neutral benchmark exists). Cross-spoke seam updated: cloud-wiki is runner-up and already names
Modal in its own open question (agent-code execution, vs E2B / aws-lambda-microvms) — offering there,
scheduler here; agentic-tooling-wiki owns the consuming agents + harness sandboxing (cloud-run-sandboxes).
Entities: none created — followed spoke convention (no publisher/author nodes; Modal is paged as the product,
modal). entity-index checked “Modal” → NO MATCH (nearest 0.40 meta). url-only. avoid-ai-writing run.
28 → 32 pages.
[2026-07-22] ingest | KNOD — network processing on AMD GPUs (Igor’s Lab)
Routed from the hub (route entry in ../log.md). URL-only, T3 (hardware press reporting an LKML RFC
second-hand; no code, no benchmarks, developer unnamed).
New pages: knod (mechanism — in-kernel network offload device: NIC DMAs packets into GPU memory, the
kernel owns the GPU queues and JITs programs to GPU machine code; XDP / IPsec SAs / LB+filtering as offload
targets; ROCm/CUDA explicitly kept off the data path; GCN + RDNA2; RFC, no target kernel version) and
knod-igorslab (source summary).
Dedup: grep for xdp|knod|igorslab across wiki/ + raw/ → no hits. No XDP page created — it is named
inline off ebpf rather than paged on one mention.
Gap-relevance: continues the spoke’s kernel-programmability line ebpf → sched-ext → knod. The
first two moved where in the kernel your code runs; this one moves what executes it. Synthesis gained a
paragraph on that, on the seam trade (kernel↔userspace-runtime deleted from the hot path, kernel↔GPU added),
and on the “queue you remove reappears one layer below” rhyme with modal-1m-sandboxes / meta-sched-ext-ads.
Opened one question: does off-CPU data-path offload actually pay? — no figure exists here for GPU or
DPU offload, and per-packet work on a GPU pays a PCIe hop a driver-hook XDP program does not.
Boundary note: the spoke’s spin-out rule parks AI-data-center silicon (NVIDIA DOCA, _inbox
ai-infrastructure). KNOD is kernel software on commodity GPUs, in the XDP/eBPF datapath this spoke
already owns — so it routes here, with DOCA recorded as cross-spoke context on knod rather than folded in.
Entities: none — spoke convention (no publisher/author nodes); the article names no developer anyway.
Verify deferred per hub policy (content-only, no page moves). avoid-ai-writing run.
32 → 34 pages.
[2026-07-25] ingest | Client-Side Load Balancing at a Million Requests Per Second (Zalando)
Routed from the hub (route entry in ../log.md; runner-up cloud-wiki for the cost/EC2-transfer angle).
URL-only, T2 (first-party engineering blog, specific mechanisms + figures, all self-reported).
New pages: client-side-load-balancing (mechanism — CSLB: caller routes in-process; endpoint watch,
consistent-hash ring, occupancy load signal, bounded load, fade-in; the placement-primitives tension),
zalando-cslb-1m-rps (source summary), zalando (Corporation entity).
Dedup: grep for zalando|skipper|consistent hash|load balanc across all spokes’ wiki/ → only incidental
mentions (container-scheduling, modal-1m-sandboxes, kubernetes-docs-overview); no LB page existed.
Gap-relevance: two hits. It closes most of the quantification question — dollars/day, pod counts, and
DORA-shaped deploy times (289 → 128 min) from one change — and it generalises the stale-view argument from
modal-1m-sandboxes: Modal drops coordination because consistency caps throughput, Zalando drops the
shared proxy because ~100× fan-out multiplies its bad minutes and because a shared component makes latency
unattributable. Synthesis gained a two-paragraph section on that and on the occupancy-vs-in-flight signal.
Opened one question: when does owning the balancer stop paying? — no source here on the fan-out/rate
crossover, and nothing comparing CSLB against the service-mesh sidecar (Envoy/xDS) Zalando skipped unbenchmarked.
Entities: created zalando — resuming ../ENTITIES.md for an org that is domain-relevant in its own right
(maintains Skipper, a Kubernetes ingress proxy this spoke’s subject matter runs on), not merely a publisher.
Author Conor Gallagher left as a plain mention per the ENTITIES relevance gate (one evidenced edge, tangential
to the domain) — which is also where the last three ingests’ “no author nodes” posture lands.
Verify deferred per hub policy (content-only, no page moves). avoid-ai-writing run.
34 → 36 pages.
[2026-07-26] ingest | HyperDX — open-source observability on ClickHouse (opensourceprojects.dev)
Routed here by the hub (runner-up: analytical-databases-wiki, which owns ClickHouse as a subject).
New: hyperdx. Updated: observability (the backend section gains a second instance and a new
axis), openobserve (sibling contrast), synthesis + the analytical-databases adjacency note.
The finding is the contrast, not the tool. Both open-source Datadog alternatives in this spoke agree
that telemetry is a columnar-scan problem and disagree on everything after that: O2 builds a store
(Parquet on object storage, DataFusion) and argues about the bill; HyperDX rents ClickHouse’s
vectorized execution and argues about the query language — PromQL/LogQL/SQL as a learning curve
you pay during an incident. Cost was the only lens the corpus had on backends; the query surface is
now a second one.
Also strengthens the analytical-databases-wiki seam the 2026-07-11 adjacency note said to watch:
HyperDX runs literally on ClickHouse, a paged subject over there. Two sources deep now; a third makes
a joint telemetry-as-OLAP-workload page hard to refuse.
Tier T3 and thin, recorded on the page: a project-listing post with no license, no maintainer or
company, no star count, no funding/acquisition history, and no coverage of metrics, alerting or
session replay. Better source is the project’s own repo/docs — re-page from those when they arrive.
Verify deferred per hub policy (content-only). avoid-ai-writing run.
[2026-08-05] lint | freshness regrade (quality cycle)
kubernetes-integration-tax volatile → stable, 61 days past window. A dated CNCF analysis piece
(May 2026) — re-reading it returns the same argument, so there is nothing for a cycle to re-verify.
Same rule as the other dated-publication regrades in ../QUALITY.md.
[2026-08-07] ingest | Prod Forge (prod-forge/backend)
Routed from the hub (Telegram). MIT reference implementation of a production-ready service — NestJS + Prisma + Postgres + Redis, Terraform on AWS ECS, Prometheus/Grafana/Loki/Promtail, sixteen documented chapters — demonstrated on a Todo API. T1 (official project repo), 217★, entirely unmeasured.
New pages: prod-forge (source summary), production-readiness (concept). Updated platform-ops to name the new floor beneath its three pillars.
Why it earns a concept page. Every other source here justifies its platform work with scale — Netflix,
Zalando at 1M rps, Modal at 1M sandboxes, Google. This is the same discipline with the scale removed, and
the list of concerns barely shrinks: critical vs non-critical dependencies with per-class fallbacks,
/health vs /health/deps so an optional dependency cannot fail readiness, graceful degradation and
in-flight draining, trace IDs into logs and Sentry, migrations inside the deploy with tracked revisions
and rollback. Folded into synthesis as the integration tax being a slope rather than a threshold, and
as an argument for the internal-developer-platform read from below (a golden path pays because many
teams would each rediscover this floor).
Recorded as a miss. This source could have supplied the spoke’s missing quantification and does not — no traffic, no incident, no postmortem, no benchmark. Added as open question: which items on the readiness list actually reduce incidents. dora-metrics and service-level-objectives still point at nothing.
Cross-spoke, noted not duplicated. Chapter 4 is a complete agent-governance config in
agentic-tooling-wiki’s vocabulary and partly Claude Code’s syntax — quality gates before generation,
MEMORY.md/REVIEW.md/Skills, pre-hooks for protected files (Edit|Write) and blocked commands
(Bash), Never Trust AI Blindly. That spoke’s constitution-vs-containment tension, settled toward
containment by a repo with no agent product to sell.
Entities: 0 created. The prod-forge GitHub org has no evidenced identity beyond the project itself —
no company, no named maintainer in the README — so paging it would duplicate the source node. Revisit if a
person or org surfaces.
[2026-08-11] ingest | Canva: session revocation at scale on S3
Routed from the hub (Telegram; runner-up defensive-security-wiki). Two source pages: the arriving
InfoQ news item canva-session-revocation-infoq (T3, 2026-08-10, Leela Kumili) and the primary it
covers, canva-session-revocations-at-scale (T2, 2026-07-22, llew-vallis at canva) — fetched
per HUB.md’s “treat the write-up as a lead, the primary as the finding” rule. They agree; the one gap is
InfoQ’s “near-real-time” headline against the primary’s propagation within minutes, recorded on both
pages.
The architecture. Sessions are stateless encrypted cookies, so revocation is a deny-list every gateway must carry. The 12-hour window (bounded by token refresh) is cut into 30-minute S3 objects, 16 bytes per revocation (principal + login-timestamp cutoff + reserved flag bits), sorted by principal for binary search. Conditional GET re-downloads only changed chunks; an async worker merges new revocations with conditional PUT for optimistic concurrency; ZooKeeper leader election reduces races but is not needed for correctness. Reported: −87.5% gateway memory, MySQL read replicas down to two, >2,000 revocations/s, ~16 MB per million records, faster deploys.
Two new Thing pages. session-revocation (the stateless-auth problem and its two levers) and object-storage-as-coordination (conditional PUT as compare-and-swap, conditional GET as a propagation channel, immutable chunking, the leader that isn’t required).
Synthesis. Folded in as the fourth instance of the shared-component thread, and the one that keeps
a shared component instead of deleting it. Sharpened the thread’s axis: the durable claim is not
staleness-beats-consistency but decouple infrastructure load from fleet size — Canva’s own result is
that DB load now tracks write throughput rather than gateway instance count, the same quantity
modal-1m-sandboxes indicts in kubernetes‘s O(nodes) heartbeats. New open question: how long a
revoked session may live, and whether anyone treats revocation latency as a security budget. New
cross-spoke seam to defensive-security-wiki (zero-trust): posture there, delivery mechanism here.
Entities: 3 created — canva (Corporation), llew-vallis (Person, author), infoq (NewsMediaOrganization). InfoQ was overdue: it now publishes four sources in this spoke (netflix-service-topology, project-as-a-service, ebpf-kernel-observability-infoq, and this one) and the node carries the standing T3 posture. Leela Kumili left as a plain mention (news-desk byline, relevance gate). AWS not paged — S3 is the subject’s infrastructure, not the source’s agent.
[2026-08-12] ingest | DORA 2024, and the paved road measured (research pass, via hub quality cycle)
Hunted for growth edge 2 — which production-readiness items actually reduce incidents. A source landed and the edge came back re-specified rather than closed, which is the honest outcome.
dora-2024-report (T2, Accelerate State of DevOps Report 2024, Google Cloud / DORA, 120 pp, read
from the published PDF; ~39 MB, extracted locally with pypdf after WebFetch on the landing page
returned only marketing copy).
Three things it changes here.
The measure. The four keys are now five metrics on two factors. Change failure rate “is strongly correlated with the other three metrics but statistical tests and methods prevent us from combining all four into one factor,” so DORA added a rework rate question — unplanned deployments in the last six months made to fix a user-facing bug — and the two together form software delivery stability. dora-metrics updated.
The ladder has a kink. In the 2024 clusters, Medium’s change fail rate (10%) is better than High’s (20%). DORA reports this and attributes cluster membership to factors beyond throughput and stability. The Elite→Low tiers are clusters that emerged from responses, not a ranking of safety, and dora-metrics now says so.
The finding this spoke did not want. 89% of respondents use an internal developer platform. Using one: individuals 8% more productive, teams 10% better — and delivery throughput ~8% lower, delivery stability 14% lower, with a further 6% throughput cost where the platform is mandatory for the whole app lifecycle. Instability combined with a platform is linked to higher burnout, which DORA states as a combination rather than a cause. The internal-developer-platform page’s own June “golden cage” caveat, written from first principles, now has a number on it.
Flagged as a contradiction, not resolved. DORA leaves three readings open and they are not equivalent for this spoke: added handoffs would mean the kubernetes-integration-tax was relocated rather than paid and would revise the spoke’s thesis; “teams ship more freely, so instability is experimentation” would revise the metric instead. Nothing here decides between them and the corpus should not pick the flattering one.
AI, recorded because the spoke keeps asking. Per 25% increase in AI adoption: organizational performance +2.3%, team +1.4%, product performance no obvious association. Teams that shifted to adding AI-powered experiences show a 10% decrease in delivery stability.
Why the edge stays open. Cross-sectional, self-reported, with an inferential leap from individuals to organizations that the report’s own methodology chapter concedes — and its outcome is delivery stability, not production incidents. Re-specified: someone who changed a control and measured incidents on both sides. Another survey correlation will not close it.
Tier note: T2 rather than T3 because of the published methodology and uncertainty intervals, with the publisher’s interest recorded — Google Cloud sells platform tooling, and this finding is unflattering to platform tooling.
[2026-08-12] lint | The volatile queue re-read: six pages, one real correction
The hub’s 2026-08-12 quality cycle named this spoke as carrying a genuine volatile queue — six pages
past the 60-day window — and said they were owed a re-read rather than a regrade. All six were fetched
again from their url:.
One page was wrong. gitops said the desired state is “stored in Git with full history.” The OpenGitOps v1.0.0 principles do not say that. Principle 2 reads “Desired state is stored in a way that enforces immutability, versioning and retains a complete version history” and principle 3 says agents pull “from the source” — no principle names Git, or any version-control system at all. The page now carries the four principles verbatim, records the gap between the name and the specification, and keeps both readings: the spec is storage-agnostic, the implementations (Argo CD, Flux) are Git. Also corrected the ownership: the GitOps Working Group under CNCF TAG App Delivery and the Linux Foundation, not “the OpenGitOps project (CNCF)”.
Two confirmed unchanged. opentelemetry — definition, the three signals, “not an observability backend itself”, all word for word; noted that profiles are still not named on that page, since it is where a fourth signal would appear first. prometheus — SoundCloud, “the second hosted project, after Kubernetes” in 2016, the dimensional model, PromQL, the billing-accuracy exclusion; noted that the overview names no version and no OTLP ingestion, and that OTLP arriving there is the change to catch, because it would blur the OTel-produces / Prometheus-stores line this spoke draws.
Three were flagged volatile and should never have been. service-level-objectives summarizes a
published book chapter (SRE book ch. 4 — Chris Jones, John Wilkes, Niall Murphy with Cody Smith,
ed. Betsy Beyer, now credited on the page). project-as-a-service is a dated news report (InfoQ,
Ben Linders, 11 June 2026). distributed-tracing is a concept page over stable
specification docs. Re-reading any of them returns the same text, so the flag could not fire and could
not clear. All three regraded stable under the QUALITY.md rule.
Two pages gained real content on the way through. distributed-tracing now records what a span actually carries — including span kind (client / server / internal / producer / consumer), which is the field that makes a boundary crossing legible in trace data: a client span and its server span are one hop seen from both sides, an internal span is explicitly not a crossing. For a spoke arguing “seams, not components” that is the distinction the data model already encodes, and this page had been missing it. project-as-a-service now names the stack — OpenShift, Tekton, Argo CD, Backstage, Kustomize, ChatOps — which upgrades the GitOps reading from an analogy the wiki imposed to something literally true: a GitOps reconciler is running under the provisioning operator.
The standing quantification gap survived a direct attempt to close it. Going back to the Belastingdienst report for numbers returned the same lone figure it always carried — 99+ DevOps teams — and no time-to-provision, no onboarding duration, no before/after. Two months on, this spoke’s most concrete real-world platform-as-a-product instance still reports its scale and not its effect.
Net: 6 pages re-read, 1 corrected, 2 confirmed, 3 regraded, 0 new pages. Volatile queue now 0 past window.