Question Answering on NaturalQuestions
62.56F1FlowSteer
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| FlowSteerMethodology=Ours, Backbone=GPT-4o-mini2026.02 | 62.56 | 54.69 | — | |
| RAMBackbone=Qwen3-4B-Instruct, Compression=4x2026.02 | 59.97 | 66.59 | — | |
| RAMBackbone=LLaMA-3.1-8B-Instruct, Compression=4x2026.02 | 59.22 | 65.35 | — | |
| Activation BeaconBackbone=LLaMA-3.1-8B-Instruct, Compression=4x2026.02 | 58.97 | 47.04 | — | |
| LongLLMLinguaBackbone=LLaMA-3.1-8B-Instruct, Compression=4x2026.02 | 58.34 | 45.5 | — | |
| EXITBackbone=LLaMA-3.1-8B-Instruct, Compression=4x2026.02 | 57.7 | 43.31 | — | |
| RAMBackbone=LLaMA-3.1-8B-Instruct, Compression=8x2026.02 | 57.14 | 62.41 | — | |
| RAMBackbone=Qwen3-4B-Instruct, Compression=8x2026.02 | 56.15 | 61.09 | — | |
| Agentflow (Agent+RL)Methodology=Agent+RL, Backbone=GPT-4o-mini, Agent Configuration=Agentflow2026.02 | 55.98 | 45.7 | — | |
| EXITBackbone=LLaMA-3.1-8B-Instruct, Compression=8x2026.02 | 55.42 | 43.88 | — | |
| Orchestrator (Agent+RL)Methodology=Agent+RL, Backbone=GPT-4o-mini, Agent Configuration=Orchestrator2026.02 | 55.41 | 50 | — | |
| Activation BeaconBackbone=LLaMA-3.1-8B-Instruct, Compression=8x2026.02 | 55.09 | 41.21 | — | |
| LongLLMLinguaBackbone=LLaMA-3.1-8B-Instruct, Compression=8x2026.02 | 54.3 | 39.5 | — | |
| EXITBackbone=Qwen3-4B-Instruct, Compression=8x2026.02 | 53.9 | 41.54 | — | |
| SFT (Qwen3-8B)Methodology=SFT, Backbone=Qwen3-8B2026.02 | 53.4 | 46.09 | — | |
| GRPO (Qwen3-8B)Methodology=GRPO, Backbone=Qwen3-8B2026.02 | 53.24 | 43.75 | — | |
| LongLLMLinguaBackbone=Qwen3-4B-Instruct, Compression=4x2026.02 | 53.01 | 40.23 | — | |
| EXITBackbone=Qwen3-4B-Instruct, Compression=4x2026.02 | 52.86 | 40 | — | |
| Router-R1 (Agent+RL)Methodology=Agent+RL, Backbone=GPT-4o-mini, Agent Configuration=Router-R12026.02 | 52.79 | 49.22 | — | |
| Baseline (4o-mini)Methodology=Baseline, Backbone=GPT-4o-mini2026.02 | 51.42 | 39.84 | — | |
| Baseline (Qwen3-8B)Methodology=Baseline, Backbone=Qwen3-8B2026.02 | 50.75 | 39.84 | — | |
| AFlow (4o-mini)Methodology=AFlow, Backbone=GPT-4o-mini2026.02 | 49.92 | 42.97 | — | |
| ProvenceBackbone=LLaMA-3.1-8B-Instruct, Compression=4x2026.02 | 49.13 | 37.25 | — | |
| Original PromptBackbone=LLaMA-3.1-8B-Instruct2026.02 | 48.25 | 37.63 | — | |
| ProvenceBackbone=LLaMA-3.1-8B-Instruct, Compression=8x2026.02 | 47.19 | 34.16 | — | |
| QREAM-FTReader=GPT-5 mini2026.04 | 46.2 | — | 46.6 | |
| LongLLMLinguaBackbone=Qwen3-4B-Instruct, Compression=8x2026.02 | 45.27 | 31.41 | — | |
| GPT-5 miniReader=GPT-5 mini2026.04 | 45.2 | — | 43.8 | |
| Original PromptBackbone=Qwen3-4B-Instruct2026.02 | 44.44 | 32.77 | — | |
| FaviCompReader=GPT-5 mini2026.04 | 44.4 | — | 45.1 | |
| LLMLingua-2-largeBackbone=LLaMA-3.1-8B-Instruct, Compression=4x2026.02 | 43.92 | 29.98 | — | |
| ProvenceBackbone=Qwen3-4B-Instruct, Compression=4x2026.02 | 43.39 | 31.11 | — | |
| ProvenceBackbone=Qwen3-4B-Instruct, Compression=8x2026.02 | 43.01 | 31.3 | — | |
| ICAEBackbone=LLaMA-3.1-8B-Instruct, Compression=8x2026.02 | 38.31 | 36.13 | — | |
| ICAEBackbone=LLaMA-3.1-8B-Instruct, Compression=4x2026.02 | 37.57 | 35.82 | — | |
| LLMLingua-2-largeBackbone=Qwen3-4B-Instruct, Compression=4x2026.02 | 35.72 | 23.09 | — | |
| LLMLingua-2-largeBackbone=LLaMA-3.1-8B-Instruct, Compression=8x2026.02 | 33.58 | 20.11 | — | |
| LLMLingua-2-largeBackbone=Qwen3-4B-Instruct, Compression=8x2026.02 | 26.67 | 15.37 | — | |
| Closed-bookBackbone=LLaMA-3.1-8B-Instruct2026.02 | 21.98 | 14.05 | — | |
| ICAEBackbone=Qwen3-4B-Instruct, Compression=4x2026.02 | 20.05 | 18.97 | — | |
| ICAEBackbone=Qwen3-4B-Instruct, Compression=8x2026.02 | 19.71 | 19.49 | — | |
| Closed-bookBackbone=Qwen3-4B-Instruct2026.02 | 17.75 | 10.17 | — | |
| BaseBackbone=Qwen-2.5-7B, Evaluation Protocol=0-shot2026.04 | — | — | 14.36 | |
| BaseBackbone=Qwen-2.5-7B-Inst, Evaluation Protocol=0-shot2026.04 | — | — | 15.12 | |
| BaseModel=Qwen-2.5-7B, Training Strategy=Base, Evaluation Protocol=0-shot2026.04 | — | — | 14.36 | |
| BaseModel=Qwen-2.5-7B-Inst, Training Strategy=Base, Evaluation Protocol=0-shot2026.04 | — | — | 15.12 | |
| POPBackbone=Qwen-2.5-7B, Evaluation Protocol=0-shot2026.04 | — | — | 14.07 | |
| POPBackbone=Qwen-2.5-7B-Inst, Evaluation Protocol=0-shot2026.04 | — | — | 14.13 | |
| POPModel=Qwen-2.5-7B, Training Strategy=POP, Evaluation Protocol=0-shot2026.04 | — | — | 13.48 | |
| POPModel=Qwen-2.5-7B-Inst, Training Strategy=POP, Evaluation Protocol=0-shot2026.04 | — | — | 14.65 | |
| Train on DBackbone=Qwen-2.5-7B, Evaluation Protocol=0-shot2026.04 | — | — | 12.19 | |
| Train on DBackbone=Qwen-2.5-7B-Inst, Evaluation Protocol=0-shot2026.04 | — | — | 11.2 | |
| Train on DModel=Qwen-2.5-7B, Training Strategy=Naive pretraining on D, Evaluation Protocol=0-shot2026.04 | — | — | 9.03 | |
| Train on DModel=Qwen-2.5-7B-Inst, Training Strategy=Naive pretraining on D, Evaluation Protocol=0-shot2026.04 | — | — | 4.87 |