Interactive Agent Task on ScienceWorld Seen (Average Reward)
77.1Average RewardBPO
Evaluation Results
| Method | Links | |
|---|---|---|
| BPOApproach=System-2, Evaluation Protocol=Fine-Tuned, Base Model=Llama-3.1-8B-Instruct2025.08 | 77.1 | |
| MPOApproach=System-2, Evaluation Protocol=Fine-Tuned, Base Model=Llama-3.1-8B-Instruct2025.08 | 71.61 | |
| ETOApproach=System-2, Evaluation Protocol=Fine-Tuned, Base Model=Llama-3.1-8B-Instruct2025.08 | 65.69 | |
| Deepseek-R1Approach=System-2, Evaluation Protocol=Zero-shot2025.08 | 63.96 | |
| SFTApproach=System-2, Evaluation Protocol=Fine-Tuned, Base Model=Llama-3.1-8B-Instruct2025.08 | 58.82 | |
| o3-miniApproach=System-2, Evaluation Protocol=Zero-shot2025.08 | 56.95 | |
| Qwen-3-ThinkingApproach=System-2, Evaluation Protocol=Zero-shot2025.08 | 52.05 | |
| Qwen-2.5-7B-InstructApproach=System-1, Evaluation Protocol=Zero-shot2025.08 | 26.68 | |
| Llama-3.1-8B-InstructApproach=System-1, Evaluation Protocol=Zero-shot2025.08 | 26.64 |