Concept Identifiability on LANG (1, 1)
0.9121MCCSSAE
Evaluation Results
| Method | Links | |
|---|---|---|
| SSAEModel Backbone=Gemma-2-2B, Protocol=Unsupervised2025.02 | 0.9121 | |
| Linear ProbeModel Backbone=Gemma-2-2B, Protocol=Supervised2025.02 | 0.8793 | |
| GemmaScopeModel Backbone=Gemma-2-2B, Protocol=Unsupervised2025.02 | 0.8614 | |
| TopK-SAEModel Backbone=Gemma-2-2B, Protocol=Unsupervised2025.02 | 0.8467 | |
| ReLU-SAEModel Backbone=Gemma-2-2B, Protocol=Unsupervised2025.02 | 0.8325 | |
| JumpReLU SAEModel Backbone=Gemma-2-2B, Protocol=Unsupervised2025.02 | 0.8219 | |
| SSAEEvaluation Protocol=Recovered decoder directions evaluated on paired observations (f(x), f(x˜))2025.02 | 0.7423 | |
| PythiaSAEEvaluation Protocol=Recovered decoder directions evaluated on paired observations (f(x), f(x˜))2025.02 | 0.7316 | |
| Linear ProbeEvaluation Protocol=Recovered decoder directions evaluated on paired observations (f(x), f(x˜))2025.02 | 0.7169 | |
| TopK-SAEEvaluation Protocol=Recovered decoder directions evaluated on paired observations (f(x), f(x˜))2025.02 | 0.7058 | |
| ReLU-SAEEvaluation Protocol=Recovered decoder directions evaluated on paired observations (f(x), f(x˜))2025.02 | 0.6931 | |
| JumpReLU SAEEvaluation Protocol=Recovered decoder directions evaluated on paired observations (f(x), f(x˜))2025.02 | 0.6847 |