Mathematical Reasoning on MATH (CoT Calibration & Accuracy)
0.513CoT NLLBayesian-LoRA
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Bayesian-LoRAModel=Qwen2.5-14B-Instruct, Zero-shot=true, N=42026.01 | 0.513 | 5.81 | 51.1 | |
| Bayesian-LoRAModel=Qwen3-30B-A3B-Instruct-2507, Zero-shot=true2026.01 | 0.721 | 6.32 | 61.9 | |
| LA (post-hoc)Model=Qwen2.5-14B-Instruct, Zero-shot=true2026.01 | 0.81 | 7.12 | 49.8 | |
| LA (post-hoc)Model=Qwen3-30B-A3B-Instruct-2507, Zero-shot=true2026.01 | 0.904 | 7.08 | 61.8 | |
| BLoBModel=Qwen3-30B-A3B-Instruct-2507, Zero-shot=true, N=102026.01 | 0.993 | 7.15 | 60.4 | |
| TempModel=Qwen3-30B-A3B-Instruct-2507, Zero-shot=true2026.01 | 1.021 | 8.94 | 61.7 | |
| Baseline FTModel=Qwen3-30B-A3B-Instruct-2507, Zero-shot=true2026.01 | 1.096 | 8.96 | 61.8 | |
| BLoBModel=Qwen2.5-14B-Instruct, Zero-shot=true, N=102026.01 | 1.21 | 8.41 | 47.2 | |
| TempModel=Qwen2.5-14B-Instruct, Zero-shot=true2026.01 | 1.96 | 10.7 | 49.9 | |
| DropoutModel=Qwen2.5-14B-Instruct, Zero-shot=true2026.01 | 2.103 | 11.9 | 50 | |
| Baseline FTModel=Qwen2.5-14B-Instruct, Zero-shot=true2026.01 | 2.165 | 12.2 | 49.8 |