Scientific Reasoning on ScienceWorld
75.9Success RateCVT-RL
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| CVT-RLVerifier=yes, Constraints=yes, CF rollouts=semantic, Belief=yes, Compute=3.92026.06 | 75.9 | — | — | — | |
| Information-matched CF processVerifier=yes, Constraints=yes, CF rollouts=semantic, Belief=yes, Compute=3.92026.06 | 72.7 | — | — | — | |
| Compute-matched non-causalVerifier=yes, Constraints=yes, CF rollouts=random extra, Belief=yes, Compute=3.82026.06 | 69.1 | — | — | — | |
| Constrained process RLVerifier=yes, Constraints=yes, CF rollouts=no, Belief=yes, Compute=1.52026.06 | 66.7 | — | — | — | |
| RLVMRVerifier=yes, Constraints=no, CF rollouts=no, Belief=no, Compute=1.32026.06 | 66.4 | — | — | — | |
| BEACONType=RL Training, Base Model=Qwen2.5-7B-Instruct2026.05 | 64.3 | — | — | 83.7 | |
| LongRLVRVerifier=yes, Constraints=no, CF rollouts=no, Belief=no, Compute=1.22026.06 | 63.2 | — | — | — | |
| Q-RAGVerifier=yes, Constraints=no, CF rollouts=no, Belief=no, Compute=1.42026.06 | 62.1 | — | — | — | |
| TROLL-styleVerifier=yes, Constraints=no, CF rollouts=no, Belief=no, Compute=1.12026.06 | 59.6 | — | — | — | |
| PPO-RLVRVerifier=yes, Constraints=no, CF rollouts=no, Belief=no, Compute=1.12026.06 | 54.3 | — | — | — | |
| GiGPOType=RL Training, Base Model=Qwen2.5-7B-Instruct2026.05 | 53.4 | — | — | 69.2 | |
| GRPOType=RL Training, Base Model=Qwen2.5-7B-Instruct2026.05 | 49.1 | — | — | 61.8 | |
| GPT-4o (ReAct)Type=Prompting, Base Model=Closed-Source Models2026.05 | 45.4 | — | — | 54.3 | |
| BEACONType=RL Training, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 45.3 | — | — | 58.9 | |
| SFTVerifier=no, Constraints=no, CF rollouts=no, Belief=no, Compute=1.02026.06 | 42.7 | — | — | — | |
| Gemini-2.5-Pro (ReAct)Type=Prompting, Base Model=Closed-Source Models2026.05 | 36.7 | — | — | 47.8 | |
| GiGPOType=RL Training, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 25.8 | — | — | 35.6 | |
| PPOType=RL Training, Base Model=Qwen2.5-7B-Instruct2026.05 | 24 | — | — | 37.1 | |
| GRPOType=RL Training, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 21.1 | — | — | 31.7 | |
| ReflexionType=Prompting, Base Model=Qwen2.5-7B-Instruct2026.05 | 11.7 | — | — | 23.4 | |
| PPOType=RL Training, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 10.9 | — | — | 29.3 | |
| ReActType=Prompting, Base Model=Qwen2.5-7B-Instruct2026.05 | 7.8 | — | — | 17.4 | |
| Direct PromptType=Prompting, Base Model=Qwen2.5-7B-Instruct2026.05 | 4.2 | — | — | 11.4 | |
| ReflexionType=Prompting, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 3.9 | — | — | 7.1 | |
| ReActType=Prompting, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 1.2 | — | — | 9 | |
| Direct PromptType=Prompting, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 0.7 | — | — | 5.9 | |
| Co-Evolving AgentsModel=Llama-2-13B-chat2025.11 | — | 74.5 | 65.5 | — | |
| ETOModel=Llama-2-13B-chat2025.11 | — | 72.6 | 65.3 | — | |
| SFTModel=Llama-2-13B-chat2025.11 | — | 64.5 | 56.1 | — |