Code Generation on CodeContest (Advanced)
50.4Pass RateLogitsCoder
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| LogitsCoderModel Category=Decoding-based Models, Backbone=Qwen2.5-14B-Instruct, Rollout Budget=202026.02 | 50.4 | 32.56 | |
| MCTSModel Category=Search-guided Reasoning Models, Backbone=Qwen2.5-14B-Instruct, Rollout Budget=202026.02 | 45.9 | 23.26 | |
| self-playModel Category=Reflection-based Models, Backbone=Qwen2.5-14B-Instruct, Rollout Budget=202026.02 | 44.59 | 30.23 | |
| Contrastive DecodingModel Category=Decoding-based Models, Backbone=Qwen2.5-14B-Instruct, Rollout Budget=202026.02 | 42.86 | 24.66 | |
| RethinkMCTSModel Category=Search-guided Reasoning Models, Backbone=Qwen2.5-14B-Instruct, Rollout Budget=202026.02 | 42.38 | 25.87 | |
| Guided DecodingModel Category=Decoding-based Models, Backbone=Qwen2.5-14B-Instruct, Rollout Budget=202026.02 | 39.96 | 23.26 | |
| LATSModel Category=Search-guided Reasoning Models, Backbone=Qwen2.5-14B-Instruct, Rollout Budget=202026.02 | 33.51 | 18.6 | |
| LDBModel Category=Reflection-based Models, Backbone=Qwen2.5-14B-Instruct, Rollout Budget=202026.02 | 30.87 | 16.28 | |
| ReflexionModel Category=Reflection-based Models, Backbone=Qwen2.5-14B-Instruct, Rollout Budget=202026.02 | 27.93 | 16.28 | |
| Zero-shotModel Category=Base, Backbone=Qwen2.5-14B-Instruct, Evaluation Protocol=Zero-shot2026.02 | 26.49 | 13.95 | |
| RAPModel Category=Reflection-based Models, Backbone=Qwen2.5-14B-Instruct, Rollout Budget=202026.02 | 20.56 | 11.63 |