Active Reasoning on AR-Bench-SP
54Non-strict Semantic EquivalenceBALAR
Evaluation Results
| Method | Links | |
|---|---|---|
| BALARLLM Backbone=Qwen2.5-32B-Instruct, Max ask rounds=252026.05 | 54 | |
| ToTLLM Backbone=Qwen2.5-32B-Instruct, Max ask rounds=252026.05 | 39 | |
| Few-ShotLLM Backbone=Qwen2.5-32B-Instruct, Max ask rounds=252026.05 | 37 | |
| Proactive CoTLLM Backbone=Qwen2.5-32B-Instruct, Max ask rounds=252026.05 | 36 | |
| ToTLLM Backbone=Llama-3.1-8B-Instruct, Max ask rounds=252026.05 | 31 | |
| UoTLLM Backbone=Llama-3.1-8B-Instruct, Max ask rounds=252026.05 | 29 | |
| BALARLLM Backbone=Llama-3.1-8B-Instruct, Max ask rounds=252026.05 | 26 | |
| Proactive CoTLLM Backbone=Llama-3.1-8B-Instruct, Max ask rounds=252026.05 | 21 | |
| Few-ShotLLM Backbone=Llama-3.1-8B-Instruct, Max ask rounds=252026.05 | 20 | |
| UoTLLM Backbone=Qwen2.5-32B-Instruct, Max ask rounds=252026.05 | 12 |