Numerical comparison on Cross-notation comparison (held-out)
94.19Verbalization AccuracyGPT-4.1
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-4.1prompting=one-shot2026.02 | 94.19 | — | |
| GPT-4.1-miniprompting=one-shot2026.02 | 92.12 | — | |
| Qwen3-8BParameters=8B2026.02 | 70 | 98.88 | |
| DeepSeek-R1-Distill-Qwen-7BParameters=7B, Backbone=Qwen, Distillation=DeepSeek-R12026.02 | 64.19 | 97.38 | |
| DeepSeek-R1-Distill-Llama-8BParameters=8B, Backbone=Llama, Distillation=DeepSeek-R12026.02 | 57.81 | 95.62 | |
| Llama-3.1-8B-InstructParameters=8B, Type=Instruct2026.02 | 55.06 | 94.81 | |
| OLMo-2-1124-7BParameters=7B2026.02 | 53.44 | 93.5 | |
| Llama-2-7bParameters=7B2026.02 | 50.81 | 98.44 | |
| Mistral-7B-v0.1Parameters=7B, Version=v0.12026.02 | 50 | 96.44 |