Story Reasoning on XStoryCloze
71AccuracyBETR
Evaluation Results
| Method | Links | |
|---|---|---|
| BETRTarget Setting=Single-Target2026.04 | 71 | |
| NAG_Llama-3.2-3BTarget Setting=Single-Target, Backbone=Llama-3.2-3B2026.04 | 70.8 | |
| NAG_SmolLM3-3BTarget Setting=Single-Target, Backbone=SmolLM3-3B2026.04 | 70.5 | |
| NAG_Qwen3-1.7BTarget Setting=Single-Target, Backbone=Qwen3-1.7B2026.04 | 70 | |
| NAG_Llama-3.2-3BTarget Setting=Multi-Target, Backbone=Llama-3.2-3B2026.04 | 69.8 | |
| BETRTarget Setting=Multi-Target2026.04 | 69.5 | |
| NAG_Qwen3-1.7BTarget Setting=Multi-Target, Backbone=Qwen3-1.7B2026.04 | 69.3 | |
| NAG_SmolLM3-3BTarget Setting=Multi-Target, Backbone=SmolLM3-3B2026.04 | 69.2 | |
| Random2026.04 | 67.1 | |
| FineWeb-Edu2026.04 | 65.9 | |
| Ours-SFTalignment=SFT2025.07 | 61.86 | |
| Ours-Base2025.07 | 60.8 | |
| Ours-Base-32kcontext-length=32k2025.07 | 60.6 | |
| Gemma 2Number of Parameters=1B, Evaluation Protocol=Zero-shot2024.08 | 58.5 | |
| FTBase Model=LLaMA-7B, Training Context=Alpaca-En2023.11 | 57.6 | |
| FTBase Model=LLaMA-7B, Training Context=Alpaca-X2023.11 | 57.2 | |
| LoRABase Model=LLaMA-7B, Training Context=Alpaca-En2023.11 | 57 | |
| LoRABase Model=LLaMA-7B, Training Context=Alpaca-X2023.11 | 57 | |
| DMoEModel Scale=1.7B, Training Strategy=Dynamic Mixture-of-Experts, Evaluation Protocol=Zero-shot2025.06 | 56.6 | |
| BLOOM + Continued Pre-trainingModel Scale=1.7B, Training Strategy=Continued Pre-training, Evaluation Protocol=Zero-shot2025.06 | 56.4 | |
| Branch-Train-MixModel Scale=1.7B, Training Strategy=Branch-Train-Mix, Evaluation Protocol=Zero-shot2025.06 | 56.4 | |
| Branch-Train-MixModel Scale=1.7B, Training Strategy=Branch-Train-Mix, Evaluation Protocol=Few-shot (4-shot)2025.06 | 56.2 | |
| XGLMNumber of Parameters=1.7B, Evaluation Protocol=Zero-shot2024.08 | 56.2 | |
| DMoEModel Scale=1.7B, Training Strategy=Dynamic Mixture-of-Experts, Evaluation Protocol=Few-shot (4-shot)2025.06 | 56.1 | |
| Parrot-7BBase Model=Parrot-7B2023.11 | 56.1 | |
| BLOOMModel Scale=1.7B, Training Strategy=Base, Evaluation Protocol=Zero-shot2025.06 | 56 | |
| BLOOM + Continued Pre-trainingModel Scale=1.7B, Training Strategy=Continued Pre-training, Evaluation Protocol=Few-shot (4-shot)2025.06 | 56 | |
| Embed FTBase Model=LLaMA-7B, Training Context=Alpaca-En2023.11 | 55.9 | |
| Embed FTBase Model=LLaMA-7B, Training Context=Alpaca-X2023.11 | 55.9 | |
| LoRABase Model=LLaMA-7B, Training Context=Bilingual2023.11 | 55.9 | |
| Embed FTBase Model=LLaMA-7B, Training Context=Bilingual2023.11 | 55.9 | |
| FTBase Model=LLaMA-7B, Training Context=Bilingual2023.11 | 55.6 | |
| BLOOMModel Scale=1.7B, Training Strategy=Base, Evaluation Protocol=Few-shot (4-shot)2025.06 | 55.5 | |
| LLaMA-7BBase Model=LLaMA-7B2023.11 | 55.5 | |
| LLaMA 3Number of Parameters=1B, Evaluation Protocol=Zero-shot2024.08 | 55.1 | |
| DMoEModel Scale=560M, Training Strategy=Dynamic Mixture-of-Experts, Evaluation Protocol=Zero-shot2025.06 | 54.1 | |
| BLOOM + Continued Pre-trainingModel Scale=560M, Training Strategy=Continued Pre-training, Evaluation Protocol=Zero-shot2025.06 | 53.8 | |
| Branch-Train-MixModel Scale=560M, Training Strategy=Branch-Train-Mix, Evaluation Protocol=Zero-shot2025.06 | 53.8 | |
| BLOOMModel Scale=560M, Training Strategy=Base, Evaluation Protocol=Zero-shot2025.06 | 53.5 | |
| DMoEModel Scale=560M, Training Strategy=Dynamic Mixture-of-Experts, Evaluation Protocol=Few-shot (4-shot)2025.06 | 53.4 | |
| Branch-Train-MixModel Scale=560M, Training Strategy=Branch-Train-Mix, Evaluation Protocol=Few-shot (4-shot)2025.06 | 53.1 | |
| XGLMNumber of Parameters=564M, Evaluation Protocol=Zero-shot2024.08 | 53.1 | |
| BLOOM + Continued Pre-trainingModel Scale=560M, Training Strategy=Continued Pre-training, Evaluation Protocol=Few-shot (4-shot)2025.06 | 52.7 | |
| Gemma 2Number of Parameters=270M, Evaluation Protocol=Zero-shot2024.08 | 52.7 | |
| BLOOMModel Scale=560M, Training Strategy=Base, Evaluation Protocol=Few-shot (4-shot)2025.06 | 52.6 | |
| BLOOMNumber of Parameters=560M, Evaluation Protocol=Zero-shot2024.08 | 52.6 | |
| GoldfishNumber of Parameters=124M, Evaluation Protocol=Zero-shot2024.08 | 52.3 | |
| Chance2024.08 | 50 | |
| Tibetan-Alpaca-7B2025.07 | 49.97 | |
| Tibetan-Llama2-7B2025.07 | 49.37 | |
| Yak-Llama2-7B2025.07 | 48.37 |