Concept Extraction Evaluation on 4 classification datasets average
99.8RAccICA
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| ICAClassifier Model=DeBERTa-v3(304M), Evaluation Protocol=fine-tuned2025.06 | 99.8 | 2.42 | 0.0073 | 100 | |
| HI-ConceptClassifier Model=DeBERTa-v3(304M), Evaluation Protocol=fine-tuned2025.06 | 99.66 | 33.83 | 0.203 | 41.19 | |
| ICAClassifier Model=BERT(110M), Evaluation Protocol=fine-tuned2025.06 | 99.52 | 2.48 | 0.0137 | 100 | |
| ICAClassifier Model=GPT-J(6B), Evaluation Protocol=fine-tuned2025.06 | 99.5 | 3.08 | 0.0119 | 100 | |
| HI-ConceptClassifier Model=Pythia(1B), Evaluation Protocol=fine-tuned2025.06 | 99.47 | 16.78 | 0.1089 | 44.38 | |
| HI-ConceptClassifier Model=GPT-J(6B), Evaluation Protocol=fine-tuned2025.06 | 99.44 | 26.19 | 0.0624 | 59.11 | |
| HI-ConceptClassifier Model=Pythia(410M), Evaluation Protocol=fine-tuned2025.06 | 99.4 | 16.81 | 0.141 | 46.77 | |
| HI-ConceptClassifier Model=BERT(110M), Evaluation Protocol=fine-tuned2025.06 | 99.33 | 29.66 | 0.1799 | 39.16 | |
| ICAClassifier Model=Pythia(410M), Evaluation Protocol=fine-tuned2025.06 | 99.29 | 2.91 | 0.0175 | 100 | |
| ICAClassifier Model=Pythia(1B), Evaluation Protocol=fine-tuned2025.06 | 99.14 | 3.62 | 0.0169 | 100 | |
| HI-ConceptClassifier Model=Mistral-Inst.(7B), Evaluation Protocol=soft-prompt tuning (PT)2025.06 | 98.98 | 16.36 | 0.0688 | 61.83 | |
| SAEClassifier Model=DeBERTa-v3(304M), Evaluation Protocol=fine-tuned, K=102025.06 | 98.47 | 3.43 | 0.259 | 29.19 | |
| ClassifSAEClassifier Model=DeBERTa-v3(304M), Evaluation Protocol=fine-tuned, K=102025.06 | 98.35 | 13.96 | 0.3521 | 9.38 | |
| HI-ConceptClassifier Model=Llama-Inst.(8B), Evaluation Protocol=soft-prompt tuning (PT)2025.06 | 98.26 | 7.3 | 0.103 | 61.33 | |
| SAEClassifier Model=Pythia(1B), Evaluation Protocol=fine-tuned, K=102025.06 | 98.08 | 4.28 | 0.2207 | 20.78 | |
| ConceptSHAPClassifier Model=DeBERTa-v3(304M), Evaluation Protocol=fine-tuned2025.06 | 98.03 | 0.53 | 0.2135 | 36.7 | |
| ClassifSAEClassifier Model=BERT(110M), Evaluation Protocol=fine-tuned, K=102025.06 | 97.745 | 9.86 | 0.4512 | 9.53 | |
| SAEClassifier Model=BERT(110M), Evaluation Protocol=fine-tuned, K=102025.06 | 97.39 | 2.04 | 0.3138 | 20.39 | |
| ClassifSAEClassifier Model=GPT-J(6B), Evaluation Protocol=fine-tuned, K=102025.06 | 97.06 | 12.03 | 0.4657 | 9.66 | |
| ClassifSAEClassifier Model=Pythia(410M), Evaluation Protocol=fine-tuned, K=102025.06 | 96.99 | 14.48 | 0.3781 | 8.1 | |
| ICAClassifier Model=Mistral-Inst.(7B), Evaluation Protocol=soft-prompt tuning (PT)2025.06 | 96.74 | 3.54 | 0.0561 | 100 | |
| ConceptSHAPClassifier Model=BERT(110M), Evaluation Protocol=fine-tuned2025.06 | 96.5 | 6.63 | 0.2455 | 37.59 | |
| ClassifSAEClassifier Model=Llama-Inst.(8B), Evaluation Protocol=soft-prompt tuning (PT), K=102025.06 | 96.35 | 17.56 | 0.3728 | 10.13 | |
| ConceptSHAPClassifier Model=GPT-J(6B), Evaluation Protocol=fine-tuned2025.06 | 96.14 | 2.89 | 0.2631 | 32.42 | |
| ConceptSHAPClassifier Model=Pythia(410M), Evaluation Protocol=fine-tuned2025.06 | 95.98 | 4.01 | 0.1948 | 37.17 | |
| ClassifSAEClassifier Model=Pythia(1B), Evaluation Protocol=fine-tuned, K=102025.06 | 95.94 | 14.92 | 0.3994 | 8.48 | |
| ClassifSAEClassifier Model=Mistral-Inst.(7B), Evaluation Protocol=soft-prompt tuning (PT), K=102025.06 | 95.93 | 11.3 | 0.4196 | 9.85 | |
| SAEClassifier Model=Pythia(410M), Evaluation Protocol=fine-tuned, K=102025.06 | 95.54 | 1.98 | 0.2742 | 18.4 | |
| SAEClassifier Model=GPT-J(6B), Evaluation Protocol=fine-tuned, K=102025.06 | 94.46 | 8.97 | 0.1498 | 35.49 | |
| ConceptSHAPClassifier Model=Pythia(1B), Evaluation Protocol=fine-tuned2025.06 | 94.05 | 3.94 | 0.2703 | 31.67 | |
| ICAClassifier Model=Llama-Inst.(8B), Evaluation Protocol=soft-prompt tuning (PT)2025.06 | 93.57 | 3.74 | 0.0443 | 100 | |
| ConceptSHAPClassifier Model=Mistral-Inst.(7B), Evaluation Protocol=soft-prompt tuning (PT)2025.06 | 93.5 | 3.47 | 0.3694 | 22.52 | |
| SAEClassifier Model=Mistral-Inst.(7B), Evaluation Protocol=soft-prompt tuning (PT), K=102025.06 | 91.27 | 11.6 | 0.0578 | 44.6 | |
| ConceptSHAPClassifier Model=Llama-Inst.(8B), Evaluation Protocol=soft-prompt tuning (PT)2025.06 | 89.16 | 2.38 | 0.3895 | 25.12 | |
| SAEClassifier Model=Llama-Inst.(8B), Evaluation Protocol=soft-prompt tuning (PT), K=102025.06 | 88.9 | 16.03 | 0.065 | 41.9 |