Question Answering on PolicyQA (test)
0.484SAEGPT-4o-mini Multi-agent-few
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| GPT-4o-mini Multi-agent-fewBase Model=GPT-4o-mini, Evaluation Protocol=Few-shot, Multi-agent framework=true2025.06 | 0.484 | 0.46 | 0.475 | 0.473 | 0.469 | 0.467 | 0.471 | 0.006 | 0.024 | |
| GPT-4o-mini FewBase Model=GPT-4o-mini, Evaluation Protocol=Few-shot, Multi-agent framework=false2025.06 | 0.478 | 0.423 | 0.458 | 0.452 | 0.444 | 0.438 | 0.449 | 0.014 | 0.055 | |
| DeepSeek-R1 Multi-agent-fewBase Model=DeepSeek-R1, Evaluation Protocol=Few-shot, Multi-agent framework=true2025.06 | 0.474 | 0.476 | 0.494 | 0.48 | 0.487 | 0.48 | 0.482 | 0.006 | 0.02 | |
| GPT-4o-mini Multi-agent-zeroBase Model=GPT-4o-mini, Evaluation Protocol=Zero-shot, Multi-agent framework=true2025.06 | 0.464 | 0.444 | 0.451 | 0.458 | 0.447 | 0.445 | 0.452 | 0.006 | 0.02 | |
| DeepSeek-R1 ZeroBase Model=DeepSeek-R1, Evaluation Protocol=Zero-shot, Multi-agent framework=false2025.06 | 0.455 | 0.436 | 0.429 | 0.437 | 0.422 | 0.422 | 0.434 | 0.009 | 0.033 | |
| DeepSeek-R1 Multi-agent-zeroBase Model=DeepSeek-R1, Evaluation Protocol=Zero-shot, Multi-agent framework=true2025.06 | 0.451 | 0.48 | 0.474 | 0.483 | 0.463 | 0.481 | 0.472 | 0.01 | 0.032 | |
| DeepSeek-R1 FewBase Model=DeepSeek-R1, Evaluation Protocol=Few-shot, Multi-agent framework=false2025.06 | 0.446 | 0.483 | 0.468 | 0.472 | 0.492 | 0.477 | 0.473 | 0.011 | 0.046 | |
| Llama 3.1 FewBase Model=Llama 3.1, Evaluation Protocol=Few-shot, Multi-agent framework=false2025.06 | 0.412 | 0.332 | 0.36 | 0.357 | 0.393 | 0.37 | 0.371 | 0.021 | 0.08 | |
| Llama 3.1 Multi-agent-fewBase Model=Llama 3.1, Evaluation Protocol=Few-shot, Multi-agent framework=true2025.06 | 0.4 | 0.38 | 0.391 | 0.385 | 0.394 | 0.372 | 0.387 | 0.008 | 0.028 | |
| Llama 3.1 Multi-agent-zeroBase Model=Llama 3.1, Evaluation Protocol=Zero-shot, Multi-agent framework=true2025.06 | 0.381 | 0.374 | 0.368 | 0.358 | 0.372 | 0.368 | 0.37 | 0.006 | 0.023 | |
| GPT-4o-mini ZeroBase Model=GPT-4o-mini, Evaluation Protocol=Zero-shot, Multi-agent framework=false2025.06 | 0.352 | 0.343 | 0.332 | 0.338 | 0.331 | 0.323 | 0.337 | 0.008 | 0.029 | |
| Llama 3.1 ZeroBase Model=Llama 3.1, Evaluation Protocol=Zero-shot, Multi-agent framework=false2025.06 | 0.31 | 0.26 | 0.268 | 0.231 | 0.237 | 0.289 | 0.266 | 0.023 | 0.079 |