Multi-constraint search problem solving on LiveDRBench 1.0 (test)
58.9AccuracyTongyiDR + LIVELEDGER
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| TongyiDR + LIVELEDGEREvaluation protocol=Trained2026.02 | 58.9 | 51.4 | |
| TongyiDREvaluation protocol=Trained2026.02 | 56.7 | 52.1 | |
| ReAct + LIVELEDGEREvaluation protocol=Prompt-based, Backbone model scale=120b2026.02 | 50.7 | 49.8 | |
| ReAct + LIVELEDGEREvaluation protocol=Prompt-based, Backbone model scale=20b2026.02 | 41.4 | 57.7 | |
| ReAct + TTSEvaluation protocol=Prompt-based, Backbone model scale=120b2026.02 | 41.3 | 52.9 | |
| ReActEvaluation protocol=Prompt-based, Backbone model scale=120b2026.02 | 39.1 | 76.3 | |
| WebExplorerEvaluation protocol=Trained2026.02 | 36.3 | 72.6 | |
| ReActEvaluation protocol=Prompt-based, Backbone model scale=20b2026.02 | 36.3 | 62.8 | |
| Search-o1Evaluation protocol=Prompt-based, Backbone model scale=120b2026.02 | 34.4 | 65.6 | |
| Search-o1Evaluation protocol=Prompt-based, Backbone model scale=20b2026.02 | 24.2 | 76.3 | |
| DR-TuluEvaluation protocol=Trained2026.02 | 17.2 | 90.2 | |
| HDSEvaluation protocol=Trained2026.02 | 16.3 | 86.5 | |
| Search-R1Evaluation protocol=Trained2026.02 | 13 | 93.5 | |
| ASearcherEvaluation protocol=Trained2026.02 | 7 | 94.9 |