Reasoning on OpenBookQA
88.4AccuracyBioBridge
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| BioBridge2026.02 | 88.4 | — | — | — | |
| Qwen2.5-7B2026.02 | 87.8 | — | — | — | |
| Fine-tuned SOTASetting=Fine-tuned2020.05 | 87.2 | — | — | — | |
| ExGRPOStudent Model=Qwen2.5-7B-Instruct2026.02 | 86.41 | — | — | — | |
| RevThinkStudent Model=Qwen2.5-7B-Instruct2026.02 | 82.85 | — | — | — | |
| BayesLoRAr=8, k=0, Params (M)=4.482025.06 | 82.8 | 14.89 | 0.97 | — | |
| BayesLoRAk=0, Params (M)=2.40–3.872025.06 | 82.07 | 15.37 | 0.95 | — | |
| BBBParams (M)=9.982025.06 | 82.06 | 11.38 | 0.66 | — | |
| MCDParams (M)=4.482025.06 | 81.72 | 13.1 | 0.77 | — | |
| LoRAParams (M)=4.482025.06 | 81.52 | 12.55 | 0.73 | — | |
| BayesLoRAk=10, Params (M)=2.40–3.872025.06 | 81.4 | 9.38 | 0.62 | — | |
| ENSParams (M)=44.802025.06 | 81.38 | 15.34 | 1.06 | — | |
| BayesLoRAr=8, k=10, Params (M)=4.482025.06 | 81.27 | 10.01 | 0.64 | — | |
| ExGRPOStudent Model=Gemma-7B-it2026.02 | 80.7 | — | — | — | |
| AnsAugStudent Model=Qwen2.5-7B-Instruct2026.02 | 80.3 | — | — | — | |
| SKDStudent Model=Qwen2.5-7B-Instruct2026.02 | 79.8 | — | — | — | |
| Zero-shotStudent Model=Qwen2.5-7B-Instruct2026.02 | 77.8 | — | — | — | |
| RevThinkStudent Model=Gemma-7B-it2026.02 | 77.2 | — | — | — | |
| ARMADALoss Function=L_cosine, Teacher Model=Stable Diffusion, Student Model=LLaMA-7B, Zero-shot=true2026.03 | 76.8 | — | — | — | |
| ARMADALoss Function=L_euclid, Teacher Model=Stable Diffusion, Student Model=LLaMA-7B, Zero-shot=true2026.03 | 76.5 | — | — | — | |
| ARMADALoss Function=L_elementwise, Teacher Model=Stable Diffusion, Student Model=LLaMA-7B, Zero-shot=true2026.03 | 75.9 | — | — | — | |
| LLaMA-7BDistillation type=undistilled, Zero-shot=true2026.03 | 75.6 | — | — | — | |
| CESBackbone=DeepSeek-R1-Distill-1.5B2026.05 | 74.8 | — | — | 520 | |
| AnsAugStudent Model=Gemma-7B-it2026.02 | 73.8 | — | — | — | |
| SKDStudent Model=Gemma-7B-it2026.02 | 73.12 | — | — | — | |
| Zero-shotStudent Model=Gemma-7B-it2026.02 | 70.2 | — | — | — | |
| DAPOBackbone=DeepSeek-R1-Distill-1.5B2026.05 | 68.2 | — | — | 503 | |
| GPT-3Setting=Few-Shot2020.05 | 65.4 | — | — | — | |
| GPT-3Setting=One-Shot2020.05 | 58.8 | — | — | — | |
| GPT-3Setting=Zero-Shot2020.05 | 57.6 | — | — | — | |
| GENICLBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 51.8 | — | — | — | |
| E5baseBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 51 | — | — | — | |
| LLM-RBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 50.8 | — | — | — | |
| GenKnowSubSetting=De, Zero-shot=true, Base Model=Phi-32025.05 | 49.8 | — | — | — | |
| GenKnowSubSetting=Avg, Zero-shot=true, Base Model=Phi-32025.05 | 49.6 | — | — | — | |
| EPRBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 49.6 | — | — | — | |
| SBERTBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 49.2 | — | — | — | |
| GenKnowSubSetting=Fr, Zero-shot=true, Base Model=Phi-32025.05 | 49 | — | — | — | |
| GenKnowSubSetting=En, Zero-shot=true, Base Model=Phi-32025.05 | 48.4 | — | — | — | |
| BM25Backbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 47.8 | — | — | — | |
| CBDSBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 47.6 | — | — | — | |
| ArrowZero-shot=true, Base Model=Phi-32025.05 | 47.4 | — | — | — | |
| SharedZero-shot=true, Base Model=Phi-32025.05 | 45.4 | — | — | — | |
| 4-bit AdamWModel=LLaMA-33B2023.09 | 45.4 | — | — | — | |
| 32-bit AdamWModel=LLaMA-13B2023.09 | 45.2 | — | — | — | |
| 4-bit AdamWModel=LLaMA-13B2023.09 | 45.2 | — | — | — | |
| Mean NormZero-shot=true, Base Model=Phi-32025.05 | 44 | — | — | — | |
| 32-bit AdamWModel=LLaMA-33B2023.09 | 43.8 | — | — | — | |
| 32-bit AdamWModel=LLaMA-7B2023.09 | 43.4 | — | — | — | |
| 4-bit AdamWModel=LLaMA-7B2023.09 | 43 | — | — | — | |
| RandomBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 43 | — | — | — | |
| Phi-3Zero-shot=true, Base Model=Phi-32025.05 | 42.8 | — | — | — | |
| Original LLaMAModel=LLaMA-7B2023.09 | 42.4 | — | — | — | |
| Original LLaMAModel=LLaMA-33B2023.09 | 42.2 | — | — | — | |
| Original LLaMAModel=LLaMA-13B2023.09 | 42 | — | — | — | |
| Zero-shotBackbone=LLaMA-7B, Evaluation Setting=Zero-shot2025.05 | 41.6 | — | — | — | |
| Gated AttentionModel Size=2B, Training Loss=base+aux2026.02 | 39.8 | — | — | — | |
| Sink AttentionModel Size=2B, Training Loss=base2026.02 | 39.6 | — | — | — | |
| Vanilla AttentionModel Size=2B, Training Loss=base+aux2026.02 | 39.4 | — | — | — | |
| Sink AttentionModel Size=2B, Training Loss=base+aux2026.02 | 39.4 | — | — | — | |
| Vanilla AttentionModel Size=1B, Training Loss=base+aux2026.02 | 38 | — | — | — | |
| Gated AttentionModel Size=1B, Training Loss=base+aux2026.02 | 38 | — | — | — | |
| Vanilla AttentionModel Size=2B, Training Loss=base2026.02 | 37.6 | — | — | — | |
| Vanilla AttentionModel Size=1B, Training Loss=base2026.02 | 37.4 | — | — | — | |
| Gated AttentionModel Size=2B, Training Loss=base2026.02 | 37.4 | — | — | — | |
| Gated AttentionModel Size=0.6B, Training Loss=base+aux2026.02 | 37 | — | — | — | |
| Sink AttentionModel Size=1B, Training Loss=base+aux2026.02 | 36.8 | — | — | — | |
| Sink AttentionModel Size=0.6B, Training Loss=base+aux2026.02 | 36.4 | — | — | — | |
| Sink AttentionModel Size=1B, Training Loss=base2026.02 | 36.4 | — | — | — | |
| Gated AttentionModel Size=1B, Training Loss=base2026.02 | 36 | — | — | — | |
| Gated AttentionModel Size=0.6B, Training Loss=base2026.02 | 35.8 | — | — | — | |
| Vanilla AttentionModel Size=0.6B, Training Loss=base+aux2026.02 | 34.4 | — | — | — | |
| Sink AttentionModel Size=0.6B, Training Loss=base2026.02 | 34.2 | — | — | — | |
| Vanilla AttentionModel Size=0.6B, Training Loss=base2026.02 | 34 | — | — | — | |
| GPTQCompression Ratio=73%, Memory (GB)=7.16 GB2026.02 | 34 | — | — | — | |
| ASVDCompression Ratio=21%, Memory (GB)=21.41 GB2026.02 | 32 | — | — | — | |
| SVD-LLM V2Compression Ratio=0.22026.05 | 32 | — | — | — | |
| OriginalCompression Ratio=0%, Memory (GB)=26.95 GB2026.02 | 31 | — | — | — | |
| SPQCompression Ratio=75%, Memory (GB)=6.86 GB2026.02 | 30 | — | — | — | |
| BaselineCompression Ratio=0.02026.05 | 28 | — | — | — | |
| SAES-SVD + PARSECompression Ratio=0.22026.05 | 28 | — | — | — | |
| SparseGPTCompression Ratio=50%, Memory (GB)=13.40 GB2026.02 | 27 | — | — | — | |
| Prot2Chat2026.02 | 25.6 | — | — | — | |
| SAES-SVD + PARSECompression Ratio=0.42026.05 | 25 | — | — | — | |
| SAES-SVD + PARSECompression Ratio=0.62026.05 | 22 | — | — | — | |
| NeuTRENOEvaluation Protocol=Zero-shot2024.10 | 19 | — | — | — | |
| Learnable ResFormerEvaluation Protocol=Zero-shot2024.10 | 18.8 | — | — | — | |
| Learnable ResFormer plusEvaluation Protocol=Zero-shot2024.10 | 17.8 | — | — | — | |
| TransformerEvaluation Protocol=Zero-shot2024.10 | 17.6 | — | — | — | |
| DenseFormerEvaluation Protocol=Zero-shot2024.10 | 17.6 | — | — | — | |
| Identity ResFormerEvaluation Protocol=Zero-shot2024.10 | 16.8 | — | — | — | |
| ProtT32026.02 | 8.75 | — | — | — |