Transcoders for the C2-2B-pt conjunctive-backdoor organism

Per-layer transcoders trained on the MLP activations of Ftm23/backdoor-gemma2-2b-2pair-hate-pt, used to build attribution graphs with circuit-tracer.

These are not reproducible from the training code alone in any short run β€” they are the artifact the attribution results depend on, published so that work can be repeated without retraining them.

Contents

path layers how trained
from_scratch/layer_{18..22}.safetensors 18–22 trained from scratch on this organism's activations
warm/layer_{23,24,25}_ep2.safetensors 23–25 warm-started from GemmaScope, 2 epochs

Reconstruction quality (post-LN FVU, k=64, held-out)

layer this transcoder off-the-shelf GemmaScope
18 0.360 0.651
19 0.403 0.690
20 0.442 0.722
21 0.452 0.698
22 0.489 0.768

Lower is better. The from-scratch transcoders reconstruct this organism's activations substantially better than the off-the-shelf ones at every layer measured.

How they are used

The attribution runs build a ReplacementModel by splicing: off-the-shelf GemmaScope transcoders for layers 0–22, and the warm/ transcoders here for layers 23–25 β€” the band where the conjunction resolves into the fire decision. The from_scratch/ set covers 18–22 and is published alongside for comparison and for pipelines that prefer organism-specific transcoders across the whole upper band.

No warm-start validation JSONs were produced for L23–25; the validate_L*.json files accompany the from-scratch set only.

Companion artifacts

  • Organism: Ftm23/backdoor-gemma2-2b-2pair-hate-pt
  • Attribution graphs: Ftm23/backdoor-attribution-graphs
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Ftm23/backdoor-transcoders-gemma2-2b-pt

Finetuned
(1)
this model