Knowledge Attribution Causal Ablation on Total (Combined ECLeKTic, MultiLoKo, G-MMLU) (test)
48Ablation Success RateXICI
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| XICIModel=GLM-4.5-Air, Details=45 MoE layers, 5760 total routed experts, Configuration=max 25, τ = 0.01, layers 6-402026.03 | 48 | 3.9 | 44.1 | 482 | 381 | 20.4 | 3,789 | 69 | |
| XICIModel=Qwen3-30B-A3B-Instruct-2507, Details=48 MoE layers, 6144 total experts, Configuration=max 25, τ = 0.005, layers 6-422026.03 | 35.5 | 3.4 | 32 | 419 | 325 | 16.7 | 2,536 | 38 | |
| Random Question-Shuffling (baseline)Model=Qwen3-30B-A3B-Instruct-2507, Details=48 MoE layers, 6144 total experts, Configuration=max 25, τ = 0.005, layers 6-422026.03 | 7.4 | 3.7 | 3.6 | — | — | — | — | — | |
| Random Question-Shuffling (baseline)Model=GLM-4.5-Air, Details=45 MoE layers, 5760 total routed experts, Configuration=max 25, τ = 0.01, layers 6-402026.03 | 7.3 | 3.8 | 3.5 | — | — | — | — | — | |
| Random Expert Set of Same Size (baseline)Model=GLM-4.5-Air, Details=45 MoE layers, 5760 total routed experts, Configuration=max 25, τ = 0.01, layers 6-402026.03 | 5 | 3.9 | 1.1 | — | — | — | — | — | |
| Random Expert Set of Same Size (baseline)Model=Qwen3-30B-A3B-Instruct-2507, Details=48 MoE layers, 6144 total experts, Configuration=max 25, τ = 0.005, layers 6-422026.03 | 4.4 | 3.2 | 1.2 | — | — | — | — | — |