Bargaining on Held-Out (test)
0.7664RewardQwen3-30B-A3B-Instruct-2507-trained
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Qwen3-30B-A3B-Instruct-2507-trainedParams=30B2026.04 | 0.7664 | 91.99 | 83.85 | 0.1 | |
| gpt-5.4-high-reasoningParams=-2026.04 | 0.4021 | 91.8 | 43.8 | 0 | |
| gpt-5.4-no-reasoningParams=-2026.04 | 0.354 | 84.38 | 42.85 | 0.78 | |
| Kimi-K2-ThinkingParams=1T2026.04 | 0.3419 | 91.21 | 36.99 | 0.2 | |
| DeepSeek-V3.1-thinkingParams=671B2026.04 | 0.324 | 91.8 | 32.66 | 0 | |
| gpt-oss-120b-noreasonParams=120B2026.04 | 0.2857 | 78.12 | 40.29 | 0.39 | |
| DeepSeek-V3.1-nothinkParams=671B2026.04 | 0.2605 | 85.94 | 32.89 | 1.56 | |
| gpt-oss-120b-reasonParams=120B2026.04 | 0.2337 | 77.54 | 34.9 | 0.2 | |
| Qwen3-30B-A3B-nothinkParams=30B2026.04 | 0.2177 | 52.93 | 4.99 | 25.39 | |
| gpt-oss-20b-reasonParams=20B2026.04 | 0.1168 | 66.99 | 17.15 | 0.78 | |
| Qwen3-4B-Instruct-2507Params=4B2026.04 | 0.113 | 51.56 | 9.01 | 6.45 | |
| Qwen3-235B-A22B-Instruct-2507Params=235B2026.04 | 0.1115 | 79.3 | 16.81 | 2.34 | |
| gpt-oss-20b-noreasonParams=20B2026.04 | 0.0626 | 58.4 | 12.65 | 9.77 | |
| Qwen3-30B-A3B-Instruct-2507-untrainedParams=30B2026.04 | 0.0126 | 73.44 | 23.92 | 10.35 | |
| Qwen3-30B-A3B-thinkParams=30B2026.04 | 0.0069 | 66.8 | 9.48 | 6.25 | |
| Llama-3.3-70B-InstructParams=70B2026.04 | 0.0068 | 74.8 | 15.8 | 15.62 |