Question Answering on Knowledge-Intensive Question Answering Benchmarks Aggregate
59.2F1CARL
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| CARLModel=4B Reasoning, Max actions=322025.12 | 59.2 | 61.9 | 141.3 | 32 | |
| CARLModel=4B Reasoning, Max actions=102025.12 | 58.4 | 60.5 | 81.8 | 32 | |
| ARPOModel=4B Reasoning, Max actions=102025.12 | 57.7 | 59.8 | 57.2 | 57.2 | |
| TreeRLModel=4B Reasoning, Max actions=102025.12 | 57.6 | 59.2 | 52.1 | 52.1 | |
| CARL-LiteModel=4B Reasoning, Max actions=102025.12 | 57.6 | 59.6 | 39.8 | 30.5 | |
| GRPOModel=4B Reasoning, Max actions=322025.12 | 57.4 | 59.7 | 115.2 | 115.2 | |
| GRPOModel=4B Reasoning, Max actions=102025.12 | 57 | 59.5 | 80.9 | 80.9 | |
| TreeRPOModel=4B Reasoning, Max actions=102025.12 | 56.5 | 58.2 | 36.7 | 36.7 | |
| CARLModel=7B Non-Reasoning, Max actions=322025.12 | 48.5 | 50.5 | 82.3 | 31.2 | |
| GRPOModel=7B Non-Reasoning, Max actions=322025.12 | 46.3 | 48.3 | 55.3 | 55.3 | |
| Zero-ShotModel=4B Reasoning, Max actions=102025.12 | 35.2 | 47.3 | — | — | |
| CARLModel=3B Non-Reasoning, Max actions=322025.12 | 34 | 33.2 | 32.1 | 32 | |
| GRPOModel=3B Non-Reasoning, Max actions=322025.12 | 33.3 | 32.5 | 32.9 | 32.9 | |
| Zero-ShotModel=7B Non-Reasoning, Max actions=322025.12 | 18.5 | 37.3 | — | — | |
| Zero-ShotModel=3B Non-Reasoning, Max actions=322025.12 | 18 | 26.8 | — | — |