Mathematical Reasoning on iGSM
54.25AccuracySupervised
Evaluation Results
| Method | Links | |
|---|---|---|
| SupervisedLLM=qwen2.5-7b-instruct2026.05 | 54.25 | |
| LOVER (sum)LLM=qwen2.5-7b-instruct2026.05 | 52 | |
| Majority VotingLLM=qwen2.5-7b-instruct2026.05 | 51 | |
| CoT-Decoding (sum)LLM=qwen2.5-7b-instruct2026.05 | 51 | |
| LOVER (max)LLM=qwen2.5-7b-instruct2026.05 | 48.25 | |
| GreedyLLM=qwen2.5-7b-instruct2026.05 | 46.5 | |
| SupervisedLLM=mistral-7b-instruct-v0.32026.05 | 44.5 | |
| SupervisedLLM=llama-3.1-8b-instruct2026.05 | 43.75 | |
| LOVER (sum)LLM=llama-3.1-8b-instruct2026.05 | 42.5 | |
| LOVER (sum)LLM=mistral-7b-instruct-v0.32026.05 | 41.25 | |
| CoT-Decoding (sum)LLM=llama-3.1-8b-instruct2026.05 | 40.5 | |
| Majority VotingLLM=llama-3.1-8b-instruct2026.05 | 39 | |
| LOVER (max)LLM=llama-3.1-8b-instruct2026.05 | 39 | |
| CoT-Decoding (sum)LLM=mistral-7b-instruct-v0.32026.05 | 36.75 | |
| GreedyLLM=llama-3.1-8b-instruct2026.05 | 35 | |
| Majority VotingLLM=mistral-7b-instruct-v0.32026.05 | 34.5 | |
| CoT-Decoding (max)LLM=qwen2.5-7b-instruct2026.05 | 31.5 | |
| LOVER (max)LLM=mistral-7b-instruct-v0.32026.05 | 28.75 | |
| CoT-Decoding (max)LLM=llama-3.1-8b-instruct2026.05 | 23.25 | |
| GreedyLLM=mistral-7b-instruct-v0.32026.05 | 19 | |
| CoT-Decoding (max)LLM=mistral-7b-instruct-v0.32026.05 | 12.75 |