Competitive Programming on LiveCodeBench (Accuracy)
65.65AccuracySMCS
Evaluation Results
| Method | Links | |
|---|---|---|
| SMCSMaximum output tokens=32,7682025.07 | 65.65 | |
| Qwen3-32BMaximum output tokens=32,7682025.07 | 64.1 | |
| QwQ-32BMaximum output tokens=32,7682025.07 | 59.6 | |
| EXAONE-Deep-32BMaximum output tokens=32,7682025.07 | 58.1 | |
| DeepSeek-R1-Distill-Qwen-32BMaximum output tokens=32,7682025.07 | 58.1 | |
| DeepSeek-R1-Distill-Llama-70BMaximum output tokens=32,7682025.07 | 57.8 | |
| GLM-Z1-32B-0414Maximum output tokens=32,7682025.07 | 56.5 | |
| GPT-o3-mini(2025-01-31)Maximum output tokens=32,7682025.07 | 54.7 | |
| GPT-4.1(2025-04-14)Maximum output tokens=32,7682025.07 | 42.2 | |
| Claude-3.7-Sonnet(2025-02-19)Maximum output tokens=32,7682025.07 | 41.3 | |
| Claude-3.5-Sonnet(2024-06-20)Maximum output tokens=32,7682025.07 | 34.3 | |
| Llama-3.3-70B-InstructMaximum output tokens=32,7682025.07 | 30.1 | |
| GPT-4o(2024-08-06)Maximum output tokens=32,7682025.07 | 29.8 | |
| Llama-3.3-Nemotron-Super-49B-v1Maximum output tokens=32,7682025.07 | 28 | |
| Gemma-3-27b-itMaximum output tokens=32,7682025.07 | 27.7 | |
| Qwen2.5-Coder-32B-InstructMaximum output tokens=32,7682025.07 | 27.7 | |
| Qwen-2.5-72B-InstructMaximum output tokens=32,7682025.07 | 26.1 | |
| HuatuoGPT-o1-72BMaximum output tokens=32,7682025.07 | 24.3 | |
| Qwen2.5-32b-InstructMaximum output tokens=32,7682025.07 | 24 | |
| TeleChat2-35B-32KMaximum output tokens=32,7682025.07 | 19.5 | |
| InternLM2.5-20B-ChatMaximum output tokens=32,7682025.07 | 14.9 |