Coreference Resolution on WSC
98.5AccuracyUPA
Evaluation Results
| Method | Links | |
|---|---|---|
| UPAExecutor=GPT-52026.01 | 98.5 | |
| SPOExecutor=GPT-52026.01 | 98.2 | |
| UPAExecutor=Claude-4.5-Sonnet2026.01 | 98 | |
| IOExecutor=GPT-52026.01 | 97.3 | |
| CoTExecutor=GPT-52026.01 | 97.3 | |
| SPOExecutor=Claude-4.5-Sonnet2026.01 | 97.3 | |
| IOExecutor=Claude-4.5-Sonnet2026.01 | 96.7 | |
| CoTExecutor=Claude-4.5-Sonnet2026.01 | 96 | |
| UPAExecutor=DeepSeek-V3.22026.01 | 93.8 | |
| SPOExecutor=DeepSeek-V3.22026.01 | 93.1 | |
| CoTExecutor=DeepSeek-V3.22026.01 | 91.3 | |
| PaLM 2-Mprompting=1-shot2023.05 | 88.1 | |
| PaLM-2 Mprotocol=1-shot2023.11 | 88.1 | |
| Falcon-180Bprotocol=1-shot2023.11 | 87.5 | |
| PaLM 2-Lprompting=1-shot2023.05 | 86.9 | |
| PaLM-2 Lprotocol=1-shot2023.11 | 86.9 | |
| PaLMprompting=1-shot2023.05 | 86.3 | |
| PaLMprotocol=1-shot2023.11 | 86.3 | |
| IOExecutor=DeepSeek-V3.22026.01 | 84.9 | |
| PaLM 2-Sprompting=1-shot2023.05 | 84.6 | |
| PaLM-2 Sprotocol=1-shot2023.11 | 84.6 | |
| LaMDA-PT 137BEvaluation protocol=0-shot2022.12 | 81 | |
| FLAN 137BEvaluation protocol=0-shot, Training strategy=leave-one-category-out2022.12 | 80.8 | |
| OPT-IML 175BEvaluation protocol=0-shot2022.12 | 73.9 | |
| OPT-IML 30BEvaluation protocol=0-shot2022.12 | 68.2 | |
| T5(3B) + PE w/ ROE (ORC.)Backbone=T5 (3B), Expert type=Oracle Retrieval-of-Expert, Additional Parameters=100M2023.02 | 65.77 | |
| GPT-3 (175B)Parameters=175B2023.02 | 65.4 | |
| T0-3BParameters=3B2023.02 | 65 | |
| DenseBase Model=DeepSeek-7B, Sparsity=Dense2025.06 | 64.42 | |
| FairSeqNumber of Parameters=2.7B, Zero-Shot=true2022.04 | 63.5 | |
| Finetune2022.12 | 63.46 | |
| MUPPET2022.12 | 63.46 | |
| Multitask2022.12 | 63.27 | |
| longdocTraining=PC, Speaker features=false, Genre features=false, Pseudo-singletons=None2021.09 | 62.7 | |
| ColD-Fusion2022.12 | 62.31 | |
| T5(3B) + PE w/ ROEBackbone=T5 (3B), Expert type=Retrieval-of-Expert, Additional Parameters=100M2023.02 | 62.21 | |
| MeZOBase Model=LLaMA-3.2-1B, Optimization Regime=Zeroth-Order2025.10 | 62 | |
| FairSeqNumber of Parameters=6.7B, Zero-Shot=true2022.04 | 61.5 | |
| Hybrid-LoRAModel=OPT-1.3b, Optimizer=SGD2026.04 | 61.5 | |
| T0-11BParameters=11B2023.02 | 61.45 | |
| longdocTraining=ON + PS 60K, Speaker features=true, Genre features=false, Pseudo-singletons=60K2021.09 | 61.3 | |
| longdocTraining=Joint, Speaker features=true, Genre features=false, Pseudo-singletons=None2021.09 | 60.1 | |
| FO (Adam)Base Model=LLaMA-3.2-1B, Optimization Regime=First-Order2025.10 | 60 | |
| longdocTraining=ON, Speaker features=false, Genre features=false, Pseudo-singletons=None2021.09 | 59.8 | |
| FO-LoRAModel=OPT-1.3b, Optimizer=SGD2026.04 | 59.6 | |
| longdocTraining=ON, Speaker features=true, Genre features=false, Pseudo-singletons=None2021.09 | 59.4 | |
| longdocTraining=Joint + PS 30K, Speaker features=true, Genre features=false, Pseudo-singletons=30K2021.09 | 59.4 | |
| longdocTraining=ON, Speaker features=true, Genre features=true, Pseudo-singletons=None2021.09 | 58.7 | |
| FO-LoRAModel=OPT-1.3b, Optimizer=Adam2026.04 | 58.7 | |
| AdaLeZOBackbone=OPT-30B, Evaluation Protocol=ZO-tuning2026.04 | 58.01 | |
| FairSeqNumber of Parameters=13B, Zero-Shot=true2022.04 | 57.7 | |
| AdaLeZOBackbone=OPT-13B, Evaluation Protocol=ZO-tuning2026.04 | 57.37 | |
| MeZOBackbone=OPT-30B, Evaluation Protocol=ZO-tuning2026.04 | 57.1 | |
| T5(3B) + Cos PEBackbone=T5 (3B), Expert type=Prompt Expert trained on COSMOS-QA, Additional Parameters=100M2023.02 | 57.02 | |
| ZO Fine-tunerBase Model=LLaMA-3.2-1B, Optimization Regime=Zeroth-Order2025.10 | 57 | |
| MeZOBackbone=OPT-13B, Evaluation Protocol=ZO-tuning2026.04 | 56.09 | |
| MoSEBackbone Model=GPT2-STANDARD (322M), Training Budget=15B tokens, Execution Mode=Test-time training (TTT), Efficient FLOPs=true2026.02 | 55.31 | |
| MoSEBackbone Model=GPT2-STANDARD (322M), Training Budget=15B tokens, Execution Mode=Uniform-width execution (w=1.0)2026.02 | 54.95 | |
| GPT-3Model Variant=Curie, Zero-shot=true2022.04 | 54.8 | |
| GPT-3Model Variant=DaVinci, Zero-shot=true2022.04 | 54.8 | |
| MoSEBackbone Model=GPT2-SMALL (55M), Training Budget=15B tokens, Execution Mode=Test-time training (TTT), Efficient FLOPs=true2026.02 | 54.58 | |
| MoSEBackbone Model=GPT2-SMALL (55M), Training Budget=15B tokens, Execution Mode=Uniform-width execution (w=1.0)2026.02 | 54.21 | |
| OPT 175BEvaluation protocol=0-shot2022.12 | 53.4 | |
| OPT 30BEvaluation protocol=0-shot2022.12 | 53.2 | |
| MoEBackbone Model=GPT2-STANDARD (322M), Training Budget=15B tokens2026.02 | 53.11 | |
| MoSEBackbone Model=GPT2-SMALL (55M), Training Budget=3B tokens, Execution Mode=Test-time training (TTT), Efficient FLOPs=true2026.02 | 52.01 | |
| MoEBackbone Model=GPT2-STANDARD (322M), Training Budget=3B tokens2026.02 | 52.01 | |
| MoEBackbone Model=GPT2-SMALL (55M), Training Budget=15B tokens2026.02 | 51.65 | |
| MoSEBackbone Model=GPT2-STANDARD (322M), Training Budget=3B tokens, Execution Mode=Test-time training (TTT), Efficient FLOPs=true2026.02 | 51.65 | |
| MoSEBackbone Model=GPT2-STANDARD (322M), Training Budget=3B tokens, Execution Mode=Uniform-width execution (w=1.0)2026.02 | 51.28 | |
| MoSEBackbone Model=GPT2-SMALL (55M), Training Budget=3B tokens, Execution Mode=Uniform-width execution (w=1.0)2026.02 | 50.92 | |
| GEEPBackbone=BERT-base2021.10 | 50.5 | |
| BERT-SPPABackbone=BERT-base2021.10 | 50.2 | |
| BERT-baseBackbone=BERT-base2021.10 | 50.1 | |
| GPT-NeoXModel Size=20B, Zero-shot=true2022.04 | 50 | |
| MoEBackbone Model=GPT2-SMALL (55M), Training Budget=3B tokens2026.02 | 49.82 | |
| FairSeqNumber of Parameters=355M, Zero-Shot=true2022.04 | 47.1 | |
| SliceGPTCompression Ratio=25%, Post-training compensation=Yes, Evaluation Protocol=Zero-shot2024.12 | 43.3 | |
| GRASPCompression Ratio=25%, Post-training compensation=Yes, Evaluation Protocol=Zero-shot2024.12 | 41.4 | |
| GPT-3Model Variant=Babbage, Zero-shot=true2022.04 | 40.4 | |
| LLMPrunerCompression Ratio=25%, Post-training compensation=Yes, Evaluation Protocol=Zero-shot2024.12 | 40.4 | |
| LaCoCompression Ratio=25%, Post-training compensation=Yes, Evaluation Protocol=Zero-shot2024.12 | 40.4 | |
| ShortGPTCompression Ratio=25%, Post-training compensation=Yes, Evaluation Protocol=Zero-shot2024.12 | 40.4 | |
| LLM-Streamline-FFNCompression Ratio=25%, Post-training compensation=Yes, Evaluation Protocol=Zero-shot2024.12 | 38.5 | |
| Zero-shotBackbone=OPT-13B, Evaluation Protocol=Zero-shot2026.04 | 38.5 | |
| Zero-shotBackbone=OPT-30B, Evaluation Protocol=Zero-shot2026.04 | 38.5 | |
| GPT-3Model Variant=Ada, Zero-shot=true2022.04 | 37.5 | |
| DenseCompression Ratio=0%, Post-training compensation=No, Evaluation Protocol=Zero-shot2024.12 | 37.5 | |
| DenseBase Model=LLaMA-2-7B, Sparsity=Dense2025.06 | 36.54 | |
| SparsegptBase Model=LLaMA-2-7B, Sparsity=2:42025.06 | 36.54 | |
| Pruner-ZBase Model=LLaMA-2-7B, Sparsity=2:42025.06 | 36.54 | |
| MaskProBase Model=LLaMA-2-7B, Sparsity=2:42025.06 | 36.54 | |
| SparsegptBase Model=DeepSeek-7B, Sparsity=2:42025.06 | 36.54 | |
| Pruner-ZBase Model=DeepSeek-7B, Sparsity=2:42025.06 | 36.54 | |
| MaskProBase Model=DeepSeek-7B, Sparsity=2:42025.06 | 36.54 | |
| FairSeqNumber of Parameters=125M, Zero-Shot=true2022.04 | 36.5 | |
| FairSeqNumber of Parameters=1.3B, Zero-Shot=true2022.04 | 36.5 | |
| GPT-JModel Size=6B, Zero-shot=true2022.04 | 36.5 | |
| PythiaNumber of Parameters=70M, Evaluation Protocol=Five-shot2023.04 | 36.5 | |
| PythiaNumber of Parameters=160M, Evaluation Protocol=Five-shot2023.04 | 36.5 |