Output-based feature description faithfulness on GPT2 Res. SAE
47.2Faithfulness ScoreEnsembleR (MA+VP)
Evaluation Results
| Method | Links | |
|---|---|---|
| EnsembleR (MA+VP)SAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Output, Composition=MaxAct + VocabProj2025.01 | 47.2 | |
| EnsembleR (MA+TC)SAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Output, Composition=MaxAct + TokenChange2025.01 | 47.2 | |
| EnsembleR (All)SAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Output, Composition=All2025.01 | 47.2 | |
| EnsembleC (All)SAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Output, Composition=All2025.01 | 46.9 | |
| EnsembleR (VP+TC)SAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Output, Composition=VocabProj + TokenChange2025.01 | 44.2 | |
| MaxActSAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Output2025.01 | 44.1 | |
| TokenChangeSAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Output2025.01 | 43.4 | |
| VocabProjSAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Output2025.01 | 42.8 |