Abductive Reasoning on ART (test)
58.2AccuracyComplementary Steering
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Complementary SteeringBackbone=Gemma-2-9B-it, Decoding Strategy=Greedy2026.04 | 58.2 | — | — | — | — | — | — | — | — | — | |
| Complementary SteeringBackbone=Gemma-2-9B-it, Decoding Strategy=Sampling@52026.04 | 57.53 | — | — | — | — | — | — | — | — | — | |
| Mono SteeringBackbone=Gemma-2-9B-it, Decoding Strategy=Sampling@52026.04 | 55.98 | — | — | — | — | — | — | — | — | — | |
| Mono SteeringBackbone=Gemma-2-9B-it, Decoding Strategy=Greedy2026.04 | 55.74 | — | — | — | — | — | — | — | — | — | |
| UnsteeredBackbone=Gemma-2-9B-it, Decoding Strategy=Sampling@52026.04 | 54.92 | — | — | — | — | — | — | — | — | — | |
| UnsteeredBackbone=Gemma-2-9B-it, Decoding Strategy=Greedy2026.04 | 54.67 | — | — | — | — | — | — | — | — | — | |
| Complementary SteeringBackbone=GPT-OSS-20B, Decoding Strategy=Greedy2026.04 | 50.5 | — | — | — | — | — | — | — | — | — | |
| Mono SteeringBackbone=GPT-OSS-20B, Decoding Strategy=Greedy2026.04 | 47.19 | — | — | — | — | — | — | — | — | — | |
| Complementary SteeringBackbone=GPT-OSS-20B, Decoding Strategy=Sampling@52026.04 | 46.89 | — | — | — | — | — | — | — | — | — | |
| UnsteeredBackbone=GPT-OSS-20B, Decoding Strategy=Greedy2026.04 | 45.09 | — | — | — | — | — | — | — | — | — | |
| Mono SteeringBackbone=GPT-OSS-20B, Decoding Strategy=Sampling@52026.04 | 42.13 | — | — | — | — | — | — | — | — | — | |
| Complementary SteeringBackbone=Llama-3.1-8B-it, Decoding Strategy=Sampling@52026.04 | 42.07 | — | — | — | — | — | — | — | — | — | |
| UnsteeredBackbone=GPT-OSS-20B, Decoding Strategy=Sampling@52026.04 | 41.69 | — | — | — | — | — | — | — | — | — | |
| Complementary SteeringBackbone=Llama-3.1-8B-it, Decoding Strategy=Greedy2026.04 | 40.95 | — | — | — | — | — | — | — | — | — | |
| Mono SteeringBackbone=Llama-3.1-8B-it, Decoding Strategy=Greedy2026.04 | 39.19 | — | — | — | — | — | — | — | — | — | |
| Mono SteeringBackbone=Llama-3.1-8B-it, Decoding Strategy=Sampling@52026.04 | 39.1 | — | — | — | — | — | — | — | — | — | |
| UnsteeredBackbone=Llama-3.1-8B-it, Decoding Strategy=Sampling@52026.04 | 39.01 | — | — | — | — | — | — | — | — | — | |
| UnsteeredBackbone=Llama-3.1-8B-it, Decoding Strategy=Greedy2026.04 | 32.27 | — | — | — | — | — | — | — | — | — | |
| COLDBackbone=GPT2-XL, Decoding strategy=Sample-and-select2022.02 | — | 1.79 | 19.5 | 10.68 | 42.67 | 4.44 | 4 | 3.06 | 2.96 | — | |
| DELOREANBackbone=GPT2-XL2022.02 | — | 1.6 | 19.06 | 7.88 | 41.74 | 4.3 | 4.23 | 2.83 | 2.87 | — | |
| LEFT-ONLYBackbone=GPT2-XL2022.02 | — | 0.88 | 16.26 | 3.49 | 38.48 | 4.57 | 3.88 | 2.68 | 2.7 | — |