Rule-governed Reasoning on RuleArena NBA
44.2Strict AccuracyTAG
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| TAGExecutor=DeepSeek-V4, #Units=20.82026.05 | 44.2 | 65.6 | 36.2 | 15.4 | — | |
| Std RAG, top-10Executor=DeepSeek-V4, #Units=102026.05 | 43.6 | 49.2 | 46.4 | 23.1 | — | |
| No RetrievalExecutor=DeepSeek-V4, #Units=02026.05 | 42.9 | 54.1 | 43.5 | 15.4 | — | |
| Std RAG, top-20Executor=DeepSeek-V4, #Units=202026.05 | 42.9 | 55.7 | 42 | 15.4 | — | |
| Std RAG, top-5Executor=DeepSeek-V4, #Units=52026.05 | 40.4 | 55.7 | 37.7 | 11.5 | — | |
| All RulesExecutor=DeepSeek-V4, #Units=1952026.05 | 39.7 | 50.8 | 39.1 | 15.4 | — | |
| TAGExecutor=Qwen3-30B-A3B-Instruct-2507, #Units=14.02026.05 | 35.9 | 47.5 | 31.9 | 19.2 | — | |
| Std RAG, top-15Executor=DeepSeek-V4, #Units=152026.05 | 34.6 | 47.5 | 29 | 19.2 | — | |
| All RulesExecutor=Qwen3-30B-A3B-Instruct-2507, #Units=1952026.05 | 34 | 45.9 | 29 | 19.2 | — | |
| Std RAG, top-20Executor=Qwen3-30B-A3B-Instruct-2507, #Units=202026.05 | 32.7 | 47.5 | 26.1 | 15.4 | — | |
| Std RAG, top-5Executor=Qwen3-30B-A3B-Instruct-2507, #Units=52026.05 | 32.1 | 42.6 | 29 | 15.4 | — | |
| Std RAG, top-10Executor=Qwen3-30B-A3B-Instruct-2507, #Units=102026.05 | 31.4 | 42.6 | 29 | 11.5 | — | |
| No RetrievalExecutor=Qwen3-30B-A3B-Instruct-2507, #Units=02026.05 | 30.1 | 37.7 | 29 | 15.4 | — | |
| Std RAG, top-15Executor=Qwen3-30B-A3B-Instruct-2507, #Units=152026.05 | 29.5 | 39.3 | 24.6 | 19.2 | — | |
| Agent Workflow MemoryLLM Backbone=DeepSeek-V32026.06 | — | — | — | — | 59.1 | |
| Chain-of-ThoughtLLM Backbone=DeepSeek-V32026.06 | — | — | — | — | 48.5 | |
| Hand-designedLLM Backbone=DeepSeek-V32026.06 | — | — | — | — | 72.7 | |
| InducedLLM Backbone=DeepSeek-V32026.06 | — | — | — | — | 74.2 | |
| Program-of-ThoughtsLLM Backbone=DeepSeek-V32026.06 | — | — | — | — | 72.7 | |
| ReActLLM Backbone=DeepSeek-V32026.06 | — | — | — | — | 30.3 |