Output-based feature description faithfulness on GPT2 MLP SAE
40.9Faithfulness ScoreEnsembleR (VP+TC)
Evaluation Results
| Method | Links | |
|---|---|---|
| EnsembleR (VP+TC)SAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Output, Composition=VocabProj + TokenChange2025.01 | 40.9 | |
| EnsembleR (MA+TC)SAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Output, Composition=MaxAct + TokenChange2025.01 | 40.3 | |
| VocabProjSAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Output2025.01 | 38.3 | |
| EnsembleR (MA+VP)SAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Output, Composition=MaxAct + VocabProj2025.01 | 38.1 | |
| EnsembleC (All)SAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Output, Composition=All2025.01 | 37.2 | |
| EnsembleR (All)SAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Output, Composition=All2025.01 | 37.1 | |
| TokenChangeSAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Output2025.01 | 36.5 | |
| MaxActSAE Width=32k, Layer Aggregation=Averaged, Evaluation Direction=Output2025.01 | 34 |