Tool-use agent evaluation on τ-bench retail (test)
57.5Pass@4 Success RatePG-CHECKLIST
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| PG-CHECKLISTVariant=CHECKLIST, Agent and Verifier Model=GPT 5.4, n=4, Verifier type=paired-verifier2026.06 | 57.5 | — | — | — | — | — | 100 | 52.8 | |
| BaselineAgent and Verifier Model=GPT 5.4, n=4, Verifier type=paired-verifier2026.06 | 50 | — | — | — | — | — | 75 | 47.2 | |
| PG-RAWVariant=RAW, Agent and Verifier Model=GPT 5.4, n=4, Verifier type=paired-verifier2026.06 | 22.5 | — | — | — | — | — | 100 | 13.9 | |
| FAMABackbone=Qwen3-4B-Instruct-2507, Method framework=FAMA2026.04 | 16.3 | 34.6 | 24.1 | 19.3 | 13.9 | — | — | — | |
| SRBackbone=Qwen3-4B-Instruct-2507, Method framework=SR2026.04 | 11.82 | 31.3 | 19.3 | 14.34 | 10.43 | — | — | — | |
| ReActBackbone=Qwen3-4B-Instruct-2507, Method framework=ReAct2026.04 | 9.57 | 17.22 | 12.35 | 10.61 | 8.7 | — | — | — | |
| ACEModel Size=30B, Evaluation Samples=60, Evaluation Setting=single-pass2026.06 | — | — | — | — | — | 40 | — | — | |
| AWMModel Size=30B, Evaluation Samples=60, Evaluation Setting=single-pass2026.06 | — | — | — | — | — | 51.7 | — | — | |
| Dynamic CheatsheetModel Size=30B, Evaluation Samples=60, Evaluation Setting=single-pass2026.06 | — | — | — | — | — | 36.7 | — | — | |
| GEPAModel Size=30B, Evaluation Samples=60, Evaluation Setting=single-pass2026.06 | — | — | — | — | — | 41.7 | — | — | |
| ReActModel Size=30B, Evaluation Samples=60, Evaluation Setting=single-pass2026.06 | — | — | — | — | — | 41.7 | — | — | |
| RSEAModel Size=30B, Evaluation Samples=60, Evaluation Setting=single-pass2026.06 | — | — | — | — | — | 40 | — | — |