Spokes.wiki Search About
Software Application source ↗ source url updated Sun Jul 26 2026 00:00:00 GMT+0000 (Coordinated Universal Time)

gemma-4-31B-it-scotoma

A weight-edited derivative of Gemma 4 31B-it published on Hugging Face by ReadyArt, Apache-2.0, 33B params in BF16. It is not a fine-tune: the weights were edited directly to weaken the base model’s refusal behaviour — an abliteration — and the card presents it as a research artifact rather than a product. 367 downloads in the month before this ingest.

Scotoma is the medical term for a bounded blind spot in the visual field, and the name is the whole argument: cut a small hole rather than remove the faculty.

The method

Three steps, per the card:

  1. Locate the refusal directions with heretic abliteration.
  2. Project through a Jacobian lens to isolate the behavioural component, keeping only ~22% of the abliteration magnitude.
  3. Merge that residue back at 1.5× scaling into the BF16 weights, across layers 7–41.

Step 2 is the interesting one. Standard abliteration removes the refusal direction wholesale and pays for it in coherence; the claim here is that most of what a full abliteration deletes is not refusal at all, and that the fifth of it that is can be isolated and re-applied at higher gain. “Loosened, not lobotomized,” as the card puts it.

The claim that undercuts itself

The card’s own safety line: “scotoma refuses basically as much as its base model. You are responsible for what you generate and how it’s used.” The stated effect is elsewhere — it “loosens gemma-4-31B-it’s cautious reflex,” producing “output that’s more varied, direct, and creative.”

So a technique named for removing a refusal direction ships with the disclaimer that refusal rates barely moved. Either the surgery is doing something other than what its lineage is named for, or refusal-as-measured and cautious-reflex-as-felt are two different quantities. The card does not resolve this, and nothing in it lets a reader check.

What isn’t there

No benchmarks. No evals. No dataset. Every behavioural claim — more varied, more direct, less coherence loss than broad abliteration — is the author’s unmeasured self-report on their own artifact. That is the T3, and it is the load-bearing gap: the entire pitch is precision, a comparative claim against full abliteration, and no comparison is shown.

A 31B in the Gemma 4 family

The base is named gemma-4-31B-it, a size absent from the family table on gemma-4 (E4B / 12B / 26B-A4B). The card also reports 33B parameters for the edited weights. Both numbers come from a derivative’s card, not from google — treat the 31B-it variant as reported-not-confirmed until a first-party source lands.

Why it sits in this spoke

Open weights are usually argued over as licensing and deployment: who can download, who can serve, what it costs (open-weight-models). This is the other half of the permission. Apache-2.0 on a downloadable checkpoint means third parties can reach into layers 7–41 and edit the safety behaviour, publish the result under the same license, and owe the original lab nothing. Google shipped alignment as weights; weights are editable. The proprietary counterpart — a refusal you receive as an API response and can only route around, never remove — is claude-refusals-and-fallback.

Cross-spoke adjacency

The Jacobian lens is Anthropic’s mechanistic-interpretability technique, parked in the hub _inbox as the first ai-interpretability stray (MIT Tech Review on the transformer-circuits paper, 2026-07-10). This source is the second sighting and the first applied one: an interpretability instrument used as a surgical tool by someone outside the lab that built it. Recorded here rather than parked because the routable substance is a real open-weight artifact and its license; the interpretability half stays a cluster tally, not a second home.

abliteration · readyart · gemma-4 · open-weight-models · google · claude-refusals-and-fallback · quantization