Coreference Resolution on Winogrande
95AccuracyReTeX
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ReTeXBackbone=T5-large, Merging Strategy=Dynamic model merging2026.06 | 95 | 127.2 | |
| MoW-MergingBackbone=T5-large, Merging Strategy=Dynamic model merging2026.06 | 93.8 | 125.6 | |
| WEMoEBackbone=T5-large, Merging Strategy=Dynamic model merging2026.06 | 90.4 | 121 | |
| Consensus TABackbone=T5-large, Merging Strategy=Static model merging2026.06 | 78.9 | 105.6 | |
| Fine-tunedBackbone=T5-large2026.06 | 74.7 | — | |
| Full modelSparsity=0%, Backbone=Qwen3-30B-A3B, Evaluation Protocol=Zero-shot2026.05 | 73.6 | — | |
| GPT-3Evaluation Protocol=1-shot2023.03 | 73.2 | — | |
| FLAN 137BEvaluation protocol=few-shot, k=variable, Training strategy=leave-one-category-out2022.12 | 72.3 | — | |
| RCOSparsity=25%, Backbone=Qwen3-30B-A3B, Evaluation Protocol=Zero-shot2026.05 | 71.4 | — | |
| Twin-MergingBackbone=T5-large, Merging Strategy=Dynamic model merging2026.06 | 70.8 | 94.8 | |
| GPT-3 (175B)Parameters=175B2023.02 | 70.2 | — | |
| EvoESAPSparsity=25%, Backbone=Qwen3-30B-A3B, Evaluation Protocol=Zero-shot2026.05 | 70.2 | — | |
| DaWinBackbone=T5-large, Merging Strategy=Dynamic model merging2026.06 | 68.7 | 92 | |
| LaMDA-PT 137BEvaluation protocol=few-shot, k=variable2022.12 | 68.4 | — | |
| LaMDA-PT 137BEvaluation protocol=0-shot2022.12 | 68.3 | — | |
| GENICLBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 68 | — | |
| FLAN 137BEvaluation protocol=0-shot, Training strategy=leave-one-category-out2022.12 | 67.3 | — | |
| BLOOM-176BEvaluation Protocol=1-shot2023.03 | 67.01 | — | |
| SBERTBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 66.9 | — | |
| LLM-RBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 66.9 | — | |
| E5baseBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 66.7 | — | |
| EPRBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 66.5 | — | |
| TIES-MergingBackbone=T5-large, Merging Strategy=Static model merging2026.06 | 66.5 | 89 | |
| BM25Backbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 66.4 | — | |
| CBDSBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 66.3 | — | |
| OPT-66BEvaluation Protocol=1-shot2023.03 | 66.14 | — | |
| AdaMergingBackbone=T5-large, Merging Strategy=Static model merging2026.06 | 65 | 87 | |
| RCOSparsity=50%, Backbone=Qwen3-30B-A3B, Evaluation Protocol=Zero-shot2026.05 | 64.7 | — | |
| T5(3B) + PE w/ ROE (ORC.)Backbone=T5 (3B), Expert type=Oracle Retrieval-of-Expert, Additional Parameters=100M2023.02 | 64.4 | — | |
| Bloom-7b1Ratio=0%, Evaluation Protocol=Zero-shot2024.03 | 64.4 | — | |
| BloombergGPTEvaluation Protocol=1-shot2023.03 | 64.09 | — | |
| OPT-IML 175BEvaluation protocol=few-shot, k=52022.12 | 63.4 | — | |
| EvoESAPSparsity=50%, Backbone=Qwen3-30B-A3B, Evaluation Protocol=Zero-shot2026.05 | 62.9 | — | |
| OPT-IML 175BEvaluation protocol=0-shot2022.12 | 62.4 | — | |
| Task ArithmeticBackbone=T5-large, Merging Strategy=Static model merging2026.06 | 62.3 | 83.4 | |
| RandomBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 62.1 | — | |
| Zero-shotBackbone=LLaMA-7B, Evaluation Setting=Zero-shot2025.05 | 61.8 | — | |
| T5(3B) + PE w/ ROEBackbone=T5 (3B), Expert type=Retrieval-of-Expert, Additional Parameters=100M2023.02 | 61.6 | — | |
| GPT-NeoXEvaluation Protocol=1-shot2023.03 | 60.62 | — | |
| OPT 175BEvaluation protocol=few-shot, k=52022.12 | 60.5 | — | |
| OPT-IML 30BEvaluation protocol=0-shot2022.12 | 59.9 | — | |
| OPT-IML 30BEvaluation protocol=few-shot, k=52022.12 | 59.4 | — | |
| T0-11BParameters=11B2023.02 | 59.34 | — | |
| OPT 30BEvaluation protocol=few-shot, k=52022.12 | 57.8 | — | |
| HyWIARatio=25%, Evaluation Protocol=Zero-shot2024.03 | 57.73 | — | |
| OPT 175BEvaluation protocol=0-shot2022.12 | 57.7 | — | |
| T5(3B) + Cos PEBackbone=T5 (3B), Expert type=Prompt Expert trained on COSMOS-QA, Additional Parameters=100M2023.02 | 56.65 | — | |
| OPT 30BEvaluation protocol=0-shot2022.12 | 56.2 | — | |
| LLM-Pruner Element²⋆Ratio=25%, Evaluation Protocol=Zero-shot2024.03 | 56.12 | — | |
| ARMShots=5-shot, Number of Parameters=1B, Training Tokens=300B2026.01 | 55.96 | — | |
| LLM-Pruner Vector⋆Ratio=25%, Evaluation Protocol=Zero-shot2024.03 | 55.56 | — | |
| MDLMShots=5-shot, Number of Parameters=1B, Training Tokens=300B2026.01 | 54.46 | — | |
| TransformerScale=369M, mb=42026.06 | 53.3 | — | |
| CARDShots=5-shot, Number of Parameters=1B, Training Tokens=300B2026.01 | 53.28 | — | |
| Mamba-3Scale=152M2026.06 | 52.7 | — | |
| SISAScale=152M, ds=162026.06 | 52.5 | — | |
| SISAScale=50M, ds=642026.06 | 52.4 | — | |
| SISAScale=369M, ds=32, mb=22026.06 | 51.9 | — | |
| SISAScale=369M, ds=64, mb=4/82026.06 | 51.8 | — | |
| Mamba-2Scale=369M, mb=42026.06 | 51.5 | — | |
| Mamba-3Scale=369M, mb=22026.06 | 51.5 | — | |
| Mamba-2Scale=152M2026.06 | 51.4 | — | |
| SISAScale=369M, ds=32, mb=42026.06 | 51.4 | — | |
| BD3LMShots=5-shot, Number of Parameters=1B, Training Tokens=300B2026.01 | 51.38 | — | |
| TransformerScale=50M2026.06 | 51.3 | — | |
| SISAScale=369M, ds=128, mb=42026.06 | 51.3 | — | |
| TransformerScale=152M2026.06 | 51.2 | — | |
| Mamba-3Scale=50M2026.06 | 51.1 | — | |
| T0-3BParameters=3B2023.02 | 50.91 | — | |
| Mamba-2Scale=50M2026.06 | 50.8 | — | |
| Mamba-3Scale=369M, mb=42026.06 | 50.7 | — | |
| Weight AveragingBackbone=T5-large, Merging Strategy=Static model merging2026.06 | 49.7 | 66.5 |