Code Generation on APPS Intermediate
81.95Pass RateMC-Tree-Of-Agents
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| MC-Tree-Of-AgentsStrategy=Pruning2025.02 | 81.95 | — | 67 | |
| MC-Tree-Of-Agents2025.02 | 79.42 | — | 63 | |
| MC-Tree-Of-AgentsStrategy=Refine2025.02 | 78.85 | — | 62 | |
| RethinkMCTS2025.02 | 74.35 | — | 49 | |
| SingleBackbone=Claude2025.02 | 73.6 | — | 57 | |
| SingleBackbone=GPT4omini2025.02 | 72.89 | — | 50 | |
| ToT2025.02 | 63.49 | — | 33 | |
| LogitsCoderModel Category=Decoding-based Models, Backbone=Qwen2.5-14B-Instruct, Rollout Budget=202026.02 | 61.3 | 34 | — | |
| RethinkMCTSModel Category=Search-guided Reasoning Models, Backbone=Qwen2.5-14B-Instruct, Rollout Budget=202026.02 | 58.88 | 29 | — | |
| MCTSModel Category=Search-guided Reasoning Models, Backbone=Qwen2.5-14B-Instruct, Rollout Budget=202026.02 | 56.67 | 23 | — | |
| Guided DecodingModel Category=Decoding-based Models, Backbone=Qwen2.5-14B-Instruct, Rollout Budget=202026.02 | 55.27 | 25 | — | |
| Contrastive DecodingModel Category=Decoding-based Models, Backbone=Qwen2.5-14B-Instruct, Rollout Budget=202026.02 | 54.63 | 25 | — | |
| self-playModel Category=Reflection-based Models, Backbone=Qwen2.5-14B-Instruct, Rollout Budget=202026.02 | 48.76 | 21 | — | |
| LDB2025.02 | 46.78 | — | 22 | |
| LATS2025.02 | 45.86 | — | 20 | |
| Reflexion2025.02 | 45.58 | — | 21 | |
| LATSModel Category=Search-guided Reasoning Models, Backbone=Qwen2.5-14B-Instruct, Rollout Budget=202026.02 | 44.42 | 21 | — | |
| RAP2025.02 | 43.32 | — | 14 | |
| ReflexionModel Category=Reflection-based Models, Backbone=Qwen2.5-14B-Instruct, Rollout Budget=202026.02 | 41.33 | 20 | — | |
| ZeroShot2025.02 | 40.57 | — | 19 | |
| LDBModel Category=Reflection-based Models, Backbone=Qwen2.5-14B-Instruct, Rollout Budget=202026.02 | 37.61 | 27 | — | |
| RAPModel Category=Reflection-based Models, Backbone=Qwen2.5-14B-Instruct, Rollout Budget=202026.02 | 36.32 | 13 | — | |
| Zero-shotModel Category=Base, Backbone=Qwen2.5-14B-Instruct, Evaluation Protocol=Zero-shot2026.02 | 31.57 | 12 | — | |
| DareFinetune Method=Dare, Base Model=Meta-llama-3.1-instruct-8b2025.02 | 26.66 | — | 5 | |
| SFT on allFinetune Method=SFT on all, Base Model=Meta-llama-3.1-instruct-8b2025.02 | 26.4 | — | 3 | |
| SFT on cluster 0Finetune Method=SFT on cluster 0, Base Model=Meta-llama-3.1-instruct-8b2025.02 | 24.23 | — | 3 | |
| SFT on cluster 2Finetune Method=SFT on cluster 2, Base Model=Meta-llama-3.1-instruct-8b2025.02 | 23.97 | — | 3 | |
| TwinFinetune Method=Twin, Base Model=Meta-llama-3.1-instruct-8b2025.02 | 23.85 | — | 5 | |
| DisenLoRAFinetune Method=DisenLoRA, Base Model=Meta-llama-3.1-instruct-8b2025.02 | 23.11 | — | 3 | |
| TiesFinetune Method=Ties, Base Model=Meta-llama-3.1-instruct-8b2025.02 | 23.06 | — | 4 | |
| w/o tuningFinetune Method=w/o tuning, Base Model=Meta-llama-3.1-instruct-8b2025.02 | 20.72 | — | 4 | |
| SFT on cluster 1Finetune Method=SFT on cluster 1, Base Model=Meta-llama-3.1-instruct-8b2025.02 | 20.69 | — | 4 |