Question Answering on Qasper (F-1, Avg Token Count)
0.3677F1 ScoreBaseline
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| BaselineLLM Backbone=Qwen-2.5 7B, Context Type=Full Context2026.02 | 0.3677 | 5,108.42 | |
| RAG with AttentionRetriever-LlamaLLM Backbone=Qwen-2.5 7B, Retrieval Method=AttentionRetriever-Llama2026.02 | 0.322 | 376.13 | |
| BaselineLLM Backbone=Llama-3.1 8B, Context Type=Full Context2026.02 | 0.3145 | 5,026.46 | |
| BaselineLLM Backbone=GPT-4o mini, Context Type=Full Context2026.02 | 0.3142 | 4,974.21 | |
| RAG with AttentionRetriever-LlamaLLM Backbone=GPT-4o mini, Retrieval Method=AttentionRetriever-Llama2026.02 | 0.298 | 355.96 | |
| RAG with AttentionRetriever-QwenLLM Backbone=GPT-4o mini, Retrieval Method=AttentionRetriever-Qwen2026.02 | 0.2971 | 346.15 | |
| RAG with SPScannerLLM Backbone=Qwen-2.5 7B, Retrieval Method=SPScanner2026.02 | 0.2953 | 352.81 | |
| BaselineLLM Backbone=Mistral-7B v0.3, Context Type=Full Context2026.02 | 0.2933 | 5,595.38 | |
| RAG with AttentionRetriever-LlamaLLM Backbone=Llama-3.1 8B, Retrieval Method=AttentionRetriever-Llama2026.02 | 0.2929 | 392.975 | |
| RAG with AttentionRetriever-QwenLLM Backbone=Qwen-2.5 7B, Retrieval Method=AttentionRetriever-Qwen2026.02 | 0.2927 | 365.56 | |
| RAG with SPScannerLLM Backbone=GPT-4o mini, Retrieval Method=SPScanner2026.02 | 0.2833 | 333.16 | |
| RAG with SPScannerLLM Backbone=Llama-3.1 8B, Retrieval Method=SPScanner2026.02 | 0.2756 | 370.25 | |
| RAG with AttentionRetriever-LlamaLLM Backbone=Mistral-7B v0.3, Retrieval Method=AttentionRetriever-Llama2026.02 | 0.2732 | 400.39 | |
| RAG with AttentionRetriever-QwenLLM Backbone=Llama-3.1 8B, Retrieval Method=AttentionRetriever-Qwen2026.02 | 0.2697 | 382.84 | |
| RAG with AttentionRetriever-QwenLLM Backbone=Mistral-7B v0.3, Retrieval Method=AttentionRetriever-Qwen2026.02 | 0.2596 | 388.83 | |
| RAG with SPScannerLLM Backbone=Mistral-7B v0.3, Retrieval Method=SPScanner2026.02 | 0.2485 | 375.04 |