Output-based feature description faithfulness on Llama MLP SAE 3.1
55.5Faithfulness ScoreEnsembleC (All)
Evaluation Results
| Method | Links | |
|---|---|---|
| EnsembleC (All)Layer Aggregation=Averaged, Evaluation Direction=Output, Composition=All2025.01 | 55.5 | |
| TokenChangeLayer Aggregation=Averaged, Evaluation Direction=Output2025.01 | 53.1 | |
| EnsembleR (All)Layer Aggregation=Averaged, Evaluation Direction=Output, Composition=All2025.01 | 51.6 | |
| EnsembleR (VP+TC)Layer Aggregation=Averaged, Evaluation Direction=Output, Composition=VocabProj + TokenChange2025.01 | 50.7 | |
| MaxActLayer Aggregation=Averaged, Evaluation Direction=Output2025.01 | 49.6 | |
| EnsembleR (MA+TC)Layer Aggregation=Averaged, Evaluation Direction=Output, Composition=MaxAct + TokenChange2025.01 | 48.9 | |
| VocabProjLayer Aggregation=Averaged, Evaluation Direction=Output2025.01 | 48.2 | |
| EnsembleR (MA+VP)Layer Aggregation=Averaged, Evaluation Direction=Output, Composition=MaxAct + VocabProj2025.01 | 45.8 |