Accuracy on OpenBookQA
96.07AccuracyPromptWizard
Evaluation Results
| Method | Links | |
|---|---|---|
| PromptWizardModel=Qwen/Qwen3.5-27B2026.05 | 96.07 | |
| Zero-Shot CoTModel=Qwen/Qwen3.5-27B2026.05 | 95.8 | |
| SpuriousModel=Qwen/Qwen3.5-27B2026.05 | 95.67 | |
| Least-to-MostModel=Qwen/Qwen3.5-27B2026.05 | 95.6 | |
| Spurious UniversalModel=Qwen/Qwen3.5-27B2026.05 | 95.49 | |
| Re-ReadingModel=Qwen/Qwen3.5-27B2026.05 | 95.2 | |
| Plan-and-SolveModel=Qwen/Qwen3.5-27B2026.05 | 95 | |
| AnalogicalModel=Qwen/Qwen3.5-27B2026.05 | 94.4 | |
| Step-BackModel=Qwen/Qwen3.5-27B2026.05 | 93.8 | |
| Self-AskModel=Qwen/Qwen3.5-27B2026.05 | 91.8 | |
| Zero-Shot CoTModel=allenai/Olmo-3-7B-Instruct2026.05 | 86.2 | |
| PromptWizardModel=allenai/Olmo-3-7B-Instruct2026.05 | 86.07 | |
| SafeInstrBackbone=Qwen-2.5-7B-Instruct2026.05 | 84.15 | |
| Re-ReadingModel=allenai/Olmo-3-7B-Instruct2026.05 | 84 | |
| SFTBackbone=Qwen-2.5-7B-Instruct2026.05 | 83.7 | |
| SafeGradBackbone=Qwen-2.5-7B-Instruct2026.05 | 83.3 | |
| PTSTBackbone=Qwen-2.5-7B-Instruct2026.05 | 83.25 | |
| SPARDBackbone=Qwen-2.5-7B-Instruct2026.05 | 83.25 | |
| TSDSIteration=1, Selection Ratio=0.2, Base Model=3b base model2026.04 | 82 | |
| Step-BackModel=allenai/Olmo-3-7B-Instruct2026.05 | 81.8 | |
| SpuriousModel=allenai/Olmo-3-7B-Instruct2026.05 | 81.4 | |
| EVOSELECTIteration=1, Selection Ratio=0.2, Base Model=3b base model2026.04 | 81 | |
| EVOSELECTIteration=2, Selection Ratio=0.2, Base Model=3b base model2026.04 | 80.8 | |
| EVOSELECTIteration=2, Selection Ratio=0.5, Base Model=3b base model2026.04 | 80.8 | |
| RandomIteration=1, Selection Ratio=0.5, Base Model=3b base model2026.04 | 80.6 | |
| EVOSELECTIteration=1, Selection Ratio=0.5, Base Model=3b base model2026.04 | 80.6 | |
| TSDSIteration=2, Selection Ratio=0.2, Base Model=3b base model2026.04 | 80.6 | |
| AllIteration=2, Selection Ratio=1.0, Base Model=3b base model2026.04 | 80.4 | |
| RandomIteration=2, Selection Ratio=0.5, Base Model=3b base model2026.04 | 80.2 | |
| AttributionIteration=2, Selection Ratio=0.5, Base Model=3b base model2026.04 | 80.2 | |
| BaseIteration=0, Selection Ratio=N/A, Base Model=3b base model2026.04 | 80 | |
| Attr-DivIteration=1, Selection Ratio=0.2, Base Model=3b base model2026.04 | 79.8 | |
| Attr-DivIteration=1, Selection Ratio=0.5, Base Model=3b base model2026.04 | 79.8 | |
| RandomIteration=2, Selection Ratio=0.2, Base Model=3b base model2026.04 | 79.8 | |
| DiversityIteration=2, Selection Ratio=0.5, Base Model=3b base model2026.04 | 79.8 | |
| AttributionIteration=1, Selection Ratio=0.2, Base Model=3b base model2026.04 | 79.6 | |
| Attr-DivIteration=2, Selection Ratio=0.5, Base Model=3b base model2026.04 | 79.4 | |
| AttributionIteration=1, Selection Ratio=0.5, Base Model=3b base model2026.04 | 79.2 | |
| TSDSIteration=2, Selection Ratio=0.5, Base Model=3b base model2026.04 | 79.2 | |
| AnalogicalModel=allenai/Olmo-3-7B-Instruct2026.05 | 79 | |
| Plan-and-SolveModel=allenai/Olmo-3-7B-Instruct2026.05 | 79 | |
| LisaBackbone=Qwen-2.5-7B-Instruct2026.05 | 78.9 | |
| TSDSIteration=1, Selection Ratio=0.5, Base Model=3b base model2026.04 | 78.8 | |
| Attr-DivIteration=2, Selection Ratio=0.2, Base Model=3b base model2026.04 | 78.8 | |
| DiversityIteration=1, Selection Ratio=0.2, Base Model=3b base model2026.04 | 78.6 | |
| DiversityIteration=1, Selection Ratio=0.5, Base Model=3b base model2026.04 | 78.6 | |
| AttributionIteration=2, Selection Ratio=0.2, Base Model=3b base model2026.04 | 78 | |
| RandomIteration=1, Selection Ratio=0.2, Base Model=3b base model2026.04 | 77.8 | |
| AllIteration=1, Selection Ratio=1.0, Base Model=3b base model2026.04 | 77.6 | |
| Qwen-2.5-7B-InstructBackbone=Qwen-2.5-7B-Instruct2026.05 | 77.6 | |
| DiversityIteration=2, Selection Ratio=0.2, Base Model=3b base model2026.04 | 76.6 | |
| UnprunedBackbone=Llama-3.1-8B2026.04 | 76 | |
| Least-to-MostModel=allenai/Olmo-3-7B-Instruct2026.05 | 74.4 | |
| Self-AskModel=allenai/Olmo-3-7B-Instruct2026.05 | 73 | |
| Spurious UniversalModel=allenai/Olmo-3-7B-Instruct2026.05 | 72.4 | |
| AnalogicalModel=Qwen/Qwen3.5-0.8B2026.05 | 67.4 | |
| Re-ReadingModel=Qwen/Qwen3.5-0.8B2026.05 | 67 | |
| Zero-Shot CoTModel=Qwen/Qwen3.5-0.8B2026.05 | 66.4 | |
| Step-BackModel=Qwen/Qwen3.5-0.8B2026.05 | 65.4 | |
| MTO-APFull name=Answer Prompting, Backbone=T5-large2026.06 | 65.2 | |
| PromptWizardModel=Qwen/Qwen3.5-0.8B2026.05 | 65 | |
| Roberta-largeEvaluation Protocol=fine-tuned, Backbone=Roberta-large2026.06 | 64.8 | |
| MTO-MPFull name=Masked Prompting, Backbone=T5-large2026.06 | 64.8 | |
| MTO-D+MCPFull name=Masked Choice Prompting with a Denoising objective, Backbone=T5-large2026.06 | 64 | |
| MTO-MAPFull name=Masked Answer Prompting, Backbone=T5-large2026.06 | 62.8 | |
| Plan-and-SolveModel=Qwen/Qwen3.5-0.8B2026.05 | 62.6 | |
| T5-largeEvaluation Protocol=fine-tuned, Backbone=T5-large, Settings=using our settings2026.06 | 61.8 | |
| Spurious UniversalModel=Qwen/Qwen3.5-0.8B2026.05 | 61.47 | |
| Least-to-MostModel=Qwen/Qwen3.5-0.8B2026.05 | 59.4 | |
| SpuriousModel=Qwen/Qwen3.5-0.8B2026.05 | 59.13 | |
| LC-QAT 1.7B#TOKENS=4B2026.06 | 55.2 | |
| r-c3LAModel=Llama3-8B [15], Adapter Rank=162026.06 | 53.69 | |
| PromptWizardModel=meta-llama/Llama-3.2-1B-Instruct2026.05 | 53.4 | |
| Zero-Shot CoTModel=meta-llama/Llama-3.2-1B-Instruct2026.05 | 53 | |
| LoRAModel=TinyLlama [71], Adapter Rank=162026.06 | 52.41 | |
| SpuriousModel=meta-llama/Llama-3.2-1B-Instruct2026.05 | 51.6 | |
| Step-BackModel=meta-llama/Llama-3.2-1B-Instruct2026.05 | 50.8 | |
| AnalogicalModel=meta-llama/Llama-3.2-1B-Instruct2026.05 | 50.8 | |
| Plan-and-SolveModel=meta-llama/Llama-3.2-1B-Instruct2026.05 | 50.2 | |
| Nautile-370MTraining tokens=∼0.8T, Evaluation Protocol=0-shot2026.04 | 49.3 | |
| Tokenizer-basedSequence Reduction=3.7, FLOPs/byte Reduction=4.12026.05 | 49.2 | |
| Spurious UniversalModel=meta-llama/Llama-3.2-1B-Instruct2026.05 | 48.93 | |
| SpaceByte + SPSequence Reduction=6.3, FLOPs/byte Reduction=2.72026.05 | 48 | |
| Re-ReadingModel=meta-llama/Llama-3.2-1B-Instruct2026.05 | 48 | |
| Least-to-MostModel=meta-llama/Llama-3.2-1B-Instruct2026.05 | 46.8 | |
| Entropy-basedSequence Reduction=6.4, FLOPs/byte Reduction=4.02026.05 | 46.4 | |
| Fixed (p = 4) + SPSequence Reduction=4.0, FLOPs/byte Reduction=2.12026.05 | 46 | |
| H-Net + SPSequence Reduction=5.3†, FLOPs/byte Reduction=2.42026.05 | 45.8 | |
| H-NetSequence Reduction=4.8, FLOPs/byte Reduction=3.42026.05 | 45.6 | |
| Entropy-based + SPSequence Reduction=6.3†, FLOPs/byte Reduction=2.72026.05 | 45.4 | |
| BF16Backbone=Llama-2-13B2026.05 | 45.2 | |
| Byte-levelSequence Reduction=1.0, FLOPs/byte Reduction=1.02026.05 | 45.2 | |
| Fixed (p = 4)Sequence Reduction=4.0, FLOPs/byte Reduction=3.12026.05 | 44.8 | |
| SpaceByteSequence Reduction=6.3, FLOPs/byte Reduction=4.02026.05 | 44.8 | |
| Fixed (p = 8) + SPSequence Reduction=8.0, FLOPs/byte Reduction=2.62026.05 | 44.4 | |
| BF16Backbone=LLaMA-2-7B2026.05 | 44.2 | |
| MR-GPTQModel=LLaMA3.1-8B-Instruct, Precision=MXFP42026.04 | 43.4 | |
| Fixed (p = 16) + SPSequence Reduction=16.0, FLOPs/byte Reduction=2.92026.05 | 43.4 | |
| MR-GPTQModel=LLaMA3-8B-Instruct, Precision=MXFP42026.04 | 43 | |
| FP16Model=LLaMA3.1-8B-Instruct, Precision=MXFP42026.04 | 43 |