Input-based feature description faithfulness on GPT2 Res. SAE
60.4Faithfulness ScoreEnsembleR (All)
Evaluation Results
| Method | Links | |
|---|---|---|
| EnsembleR (All)SAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Input, Composition=All2025.01 | 60.4 | |
| EnsembleR (MA+VP)SAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Input, Composition=MaxAct + VocabProj2025.01 | 59.6 | |
| EnsembleR (MA+TC)SAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Input, Composition=MaxAct + TokenChange2025.01 | 58.8 | |
| MaxActSAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Input2025.01 | 44.4 | |
| EnsembleC (All)SAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Input, Composition=All2025.01 | 42.4 | |
| EnsembleR (VP+TC)SAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Input, Composition=VocabProj + TokenChange2025.01 | 29.2 | |
| TokenChangeSAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Input2025.01 | 25.4 | |
| VocabProjSAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Input2025.01 | 23.7 |