Input-based feature description faithfulness on GPT2 MLP SAE
51.2Faithfulness ScoreEnsembleR (MA+VP)
Evaluation Results
| Method | Links | |
|---|---|---|
| EnsembleR (MA+VP)SAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Input, Composition=MaxAct + VocabProj2025.01 | 51.2 | |
| EnsembleR (MA+TC)SAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Input, Composition=MaxAct + TokenChange2025.01 | 51.1 | |
| EnsembleR (All)SAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Input, Composition=All2025.01 | 50.2 | |
| MaxActSAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Input2025.01 | 39.7 | |
| EnsembleC (All)SAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Input, Composition=All2025.01 | 24.4 | |
| EnsembleR (VP+TC)SAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Input, Composition=VocabProj + TokenChange2025.01 | 7.1 | |
| VocabProjSAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Input2025.01 | 6.3 | |
| TokenChangeSAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Input2025.01 | 6.1 |