Multi-Armed Bandit on Bandit
95Success Rate (pass@1)Qwen2.5-1.5B-It + Evolving Stage
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-1.5B-It + Evolving StageBackbone=Qwen2.5-1.5B-It, Stage=Evolving2026.01 | 95 | |
| Scout-DQNArchitecture=Small Neural Network2026.01 | 93 | |
| Qwen2.5-3B-It + Evolving StageBackbone=Qwen2.5-3B-It, Stage=Evolving2026.01 | 93 | |
| LLaMA3.1-1B-It + Evolving StageBackbone=LLaMA3.1-1B-It, Stage=Evolving2026.01 | 81 | |
| DeepSeek-V3Model Type=Proprietary2026.01 | 81 | |
| Scout-PPOArchitecture=Small Neural Network2026.01 | 79 | |
| Qwen2.5-3B-ItBackbone=Qwen2.5-3B-It2026.01 | 77 | |
| Qwen2.5-0.5B-It + Evolving StageBackbone=Qwen2.5-0.5B-It, Stage=Evolving2026.01 | 74 | |
| GPT-4o-miniModel Type=Proprietary2026.01 | 73 | |
| GPT-5-nanoModel Type=Proprietary2026.01 | 71 | |
| Gemini-2.5-ProModel Type=Proprietary2026.01 | 69 | |
| GPT-OSS-120BModel Type=Proprietary2026.01 | 66 | |
| Qwen2.5-1.5B-ItBackbone=Qwen2.5-1.5B-It2026.01 | 63 | |
| Qwen2.5-0.5B-It - Multi-turn PPOBackbone=Qwen2.5-0.5B-It, Evaluation Protocol=Multi-turn PPO2026.01 | 62 | |
| Qwen2.5-0.5B-It - Exploration & Distillation StageBackbone=Qwen2.5-0.5B-It, Stage=Exploration & Distillation2026.01 | 60 | |
| Qwen2.5-0.5B-It - State Estimation RLBackbone=Qwen2.5-0.5B-It, Evaluation Protocol=State Estimation RL2026.01 | 54 | |
| LLaMA3.1-1B-ItBackbone=LLaMA3.1-1B-It2026.01 | 43 | |
| Qwen2.5-0.5B-ItBackbone=Qwen2.5-0.5B-It, Stage=Base Instruction-tuned2026.01 | 39 | |
| Qwen2.5-0.5B-It - SPABackbone=Qwen2.5-0.5B-It, Evaluation Protocol=SPA2026.01 | 30 |