Text Classification on BoolQ
90.7AccuracyFull Fine-Tuning (FT)
Evaluation Results
| Method | Links | |
|---|---|---|
| Full Fine-Tuning (FT)Model=Qwen2.5-14B, Trainable Parameters=14.7B2025.12 | 90.7 | |
| LoRAModel=Qwen2.5-14B, Trainable Parameters=6.3M2025.12 | 90.2 | |
| Partial-LoRAModel=Qwen2.5-14B, Trainable Parameters=175k2025.12 | 89.6 | |
| Full Fine-Tuning (FT)Model=LLAMA3.1-8B, Trainable Parameters=8B2025.12 | 89.3 | |
| Masking (0.001%)Model=Qwen2.5-14B, Trainable Parameters=149K2025.12 | 88.6 | |
| LoRAModel=LLAMA3.1-8B, Trainable Parameters=3.4M2025.12 | 87.2 | |
| Partial-LoRAModel=LLAMA3.1-8B, Trainable Parameters=110k2025.12 | 86.7 | |
| First Order Adamw FTModel Backbone=Llama2-7B, Fine-tuning Strategy=Full Fine-tuning, Optimization Algorithm=First Order Adamw, Number of training examples=10002026.06 | 86.7 | |
| AR1-ZOModel=Qwen3-32B, Rank (r)=642026.05 | 85.6 | |
| Masking (0.001%)Model=LLAMA3.1-8B, Trainable Parameters=80K2025.12 | 85.2 | |
| MeZO-LoRAModel=Qwen3-32B, Rank (r)=642026.05 | 83.6 | |
| Fine-tuningModel Backbone=Llama-3-8B, Model Precision=16 bits, Memory Profiling=31.9GB2025.05 | 83.4 | |
| MeZOModel Backbone=Llama-3-8B, Model Precision=16 bits, Memory Profiling=20.5GB2025.05 | 83.4 | |
| I-GLASSBackbone=Gemma 7B, Sparsity=50% FF, Evaluation Protocol=0-shot unnormalized2025.08 | 81.56 | |
| GRIFFINBackbone=Gemma 7B, Sparsity=50% FF, Evaluation Protocol=0-shot unnormalized2025.08 | 81.56 | |
| GRIFFINBackbone=Mistral 7B, Sparsity=50% FF, Evaluation Protocol=0-shot unnormalized2025.08 | 80.37 | |
| I-GLASSBackbone=Mistral 7B, Sparsity=50% FF, Evaluation Protocol=0-shot unnormalized2025.08 | 80.18 | |
| QZOModel Backbone=Llama-3-8B, Model Precision=4 bits, Memory Profiling=6.3GB2025.05 | 78.2 | |
| AR1-ZOModel=Qwen3-1.7B, Rank (r)=642026.05 | 78 | |
| I-GLASSBackbone=ReLU-Llama2 7B, Sparsity=50% FF, Evaluation Protocol=0-shot unnormalized2025.08 | 77.98 | |
| GRIFFINBackbone=ReLU-Llama2 7B, Sparsity=50% FF, Evaluation Protocol=0-shot unnormalized2025.08 | 77.98 | |
| Dominant-layer ZO LoRAModel Backbone=Llama2-7B, Fine-tuning Strategy=LoRA, Optimization Algorithm=Dominant-layer ZO, Number of training examples=10002026.06 | 77.8 | |
| I-GLASSBackbone=Llama2 7B, Sparsity=50% FF, Evaluation Protocol=0-shot unnormalized2025.08 | 77.74 | |
| GRIFFINBackbone=Llama2 7B, Sparsity=50% FF, Evaluation Protocol=0-shot unnormalized2025.08 | 77.74 | |
| SubZero-GV (LoRA)Optimization Algorithm=SubZero, Tuning Strategy=LoRA, Guiding Vectors=true, Model=OPT-13B2026.01 | 77.6 | |
| MeZO LoRAModel Backbone=Llama2-7B, Fine-tuning Strategy=LoRA, Optimization Algorithm=MeZO, Number of training examples=10002026.06 | 77.6 | |
| MixLoRABase Model=LLaMA-2 13B, Trainable Parameters=2.5%2024.04 | 77.1 | |
| SubZero-GV (Prefix)Optimization Algorithm=SubZero, Tuning Strategy=Prefix Tuning, Guiding Vectors=true, Model=OPT-13B2026.01 | 77.1 | |
| FTOptimization Algorithm=Adam, Tuning Strategy=Full Tuning, Model=OPT-13B2026.01 | 77.1 | |
| MixDoRABase Model=LLaMA-2 13B, Trainable Parameters=2.5%2024.04 | 76.9 | |
| FT (Adam)Model=OPT-13B2026.05 | 76.9 | |
| MixDoRABase Model=LLaMA-3 8B, Trainable Parameters=3.0%2024.04 | 76.8 | |
| SubZero-GV (FT)Optimization Algorithm=SubZero, Tuning Strategy=Full Tuning, Guiding Vectors=true, Model=OPT-13B2026.01 | 76.8 | |
| MeZO FTModel Backbone=Llama2-7B, Fine-tuning Strategy=Full Fine-tuning, Optimization Algorithm=MeZO, Number of training examples=10002026.06 | 76.7 | |
| MeZO-GV (Prefix)Optimization Algorithm=MeZO, Tuning Strategy=Prefix Tuning, Guiding Vectors=true, Model=OPT-13B2026.01 | 76.6 | |
| Dominant-layer ZO FTModel Backbone=Llama2-7B, Fine-tuning Strategy=Full Fine-tuning, Optimization Algorithm=Dominant-layer ZO, Number of training examples=10002026.06 | 76.5 | |
| SubZero (Prefix)Optimization Algorithm=SubZero, Tuning Strategy=Prefix Tuning, Guiding Vectors=false, Model=OPT-13B2026.01 | 76.3 | |
| SubZero (LoRA)Optimization Algorithm=SubZero, Tuning Strategy=LoRA, Guiding Vectors=false, Model=OPT-13B2026.01 | 76.1 | |
| MeZO-GV (LoRA)Optimization Algorithm=MeZO, Tuning Strategy=LoRA, Guiding Vectors=true, Model=OPT-13B2026.01 | 75.6 | |
| LeRaCTraining Regime=LeRaC, Model=BERT_large2022.05 | 75.55 | |
| LoRABase Model=LLaMA-2 13B, Trainable Parameters=2.4%2024.04 | 75.4 | |
| SubZero (FT)Optimization Algorithm=SubZero, Tuning Strategy=Full Tuning, Guiding Vectors=false, Model=OPT-13B2026.01 | 75.3 | |
| DoRABase Model=LLaMA-2 13B, Trainable Parameters=2.4%2024.04 | 75.1 | |
| MixLoRABase Model=LLaMA-3 8B, Trainable Parameters=3.0%2024.04 | 75 | |
| Fine-tuningModel Backbone=Llama-2-7B, Model Precision=16 bits, Memory Profiling=26.0GB2025.05 | 75 | |
| ZO-AdaMU (Prefix)Optimization Algorithm=ZO-AdaMU, Tuning Strategy=Prefix Tuning, Model=OPT-13B2026.01 | 74.9 | |
| MeZO-LoRAModel=Qwen3-1.7B, Rank (r)=642026.05 | 74.7 | |
| CBSTraining Regime=CBS, Model=BERT_large2022.05 | 74.37 | |
| conventionalTraining Regime=conventional, Model=BERT_large2022.05 | 74.12 | |
| HiZOO (Prefix)Optimization Algorithm=HiZOO, Tuning Strategy=Prefix Tuning, Model=OPT-13B2026.01 | 73.9 | |
| MeZO (LoRA)Optimization Algorithm=MeZO, Tuning Strategy=LoRA, Guiding Vectors=false, Model=OPT-13B2026.01 | 73.8 | |
| FT (Adam)Model=OPT-2.7B2026.05 | 73.3 | |
| MeZO (Prefix)Optimization Algorithm=MeZO, Tuning Strategy=Prefix Tuning, Guiding Vectors=false, Model=OPT-13B2026.01 | 73.1 | |
| ZO-AdaMU (2x)Optimization Algorithm=ZO-AdaMU, Tuning Strategy=Full Tuning, Training Steps=2x, Model=OPT-13B2026.01 | 73 | |
| MixLoRABase Model=LLaMA-2 7B, Trainable Parameters=2.9%2024.04 | 72.7 | |
| MixDoRABase Model=LLaMA-2 7B, Trainable Parameters=2.9%2024.04 | 72.6 | |
| ZO-AdaMU (LoRA)Optimization Algorithm=ZO-AdaMU, Tuning Strategy=LoRA, Model=OPT-13B2026.01 | 72.6 | |
| MeZO-GV (FT)Optimization Algorithm=MeZO, Tuning Strategy=Full Tuning, Guiding Vectors=true, Model=OPT-13B2026.01 | 72.5 | |
| DoRABase Model=LLaMA-2 7B, Trainable Parameters=2.9%2024.04 | 71.7 | |
| ZO-Alt-NaiveModel=Qwen3-32B2026.05 | 71.6 | |
| HiZOO (LoRA)Optimization Algorithm=HiZOO, Tuning Strategy=LoRA, Model=OPT-13B2026.01 | 70.5 | |
| QZOModel=Llama-2-13B, Precision=2 bits, Memory Profiling=5.78GB, Evaluation Protocol=Zeroth-order optimization2025.05 | 70.2 | |
| Fine-tuningModel Backbone=OPT-6.7B, Model Precision=16 bits, Memory Profiling=26.8GB2025.05 | 69.6 | |
| MeZOModel Backbone=Llama-2-7B, Model Precision=16 bits, Memory Profiling=14.8GB2025.05 | 69.6 | |
| MeZOBackbone=Llama-2-7B, Precision=BF162026.05 | 69.6 | |
| Zero-Shot-QModel=Llama-2-13B, Precision=2 bits, Memory Profiling=N/A, Evaluation Protocol=Zero-shot2025.05 | 69.2 | |
| ZO-Alt-NaiveModel=Qwen3-1.7B2026.05 | 68.5 | |
| QZOModel Backbone=Llama-2-7B, Model Precision=4 bits, Memory Profiling=5.0GB2025.05 | 68.2 | |
| LOZOModel=OPT-13B2026.05 | 68.1 | |
| MeZO (FT)Optimization Algorithm=MeZO, Tuning Strategy=Full Tuning, Guiding Vectors=false, Model=OPT-13B2026.01 | 67.6 | |
| HiZOOOptimization Algorithm=HiZOO, Tuning Strategy=Full Tuning, Model=OPT-13B2026.01 | 67.3 | |
| MixDoRABase Model=Gemma 2B, Trainable Parameters=4.3%2024.04 | 67.2 | |
| LoRABase Model=LLaMA-3 8B, Trainable Parameters=2.6%2024.04 | 67.2 | |
| CAQ-ZOBackbone=Llama-2-7B, Precision=NF42026.05 | 67.2 | |
| ICLEvaluation Protocol=In-context Learning, Model=OPT-13B2026.01 | 66.9 | |
| MeZOModel Backbone=OPT-6.7B, Model Precision=16 bits, Memory Profiling=14.8GB2025.05 | 66.8 | |
| AR1-ZOModel=OPT-13B, Rank (r)=642026.05 | 66.7 | |
| Zero-shot w/o finetuneModel Backbone=Llama2-7B, Fine-tuning Strategy=None, Optimization Algorithm=None, Number of training examples=10002026.06 | 66.7 | |
| QZOModel Backbone=OPT-6.7B, Model Precision=4 bits, Memory Profiling=4.8GB2025.05 | 66.4 | |
| Zero-ShotModel Backbone=Llama-3-8B, Model Precision=16 bits2025.05 | 66.1 | |
| Zero-ShotModel Backbone=Llama-2-7B, Model Precision=16 bits2025.05 | 66 | |
| I-GLASSBackbone=OPT 6.7B, Sparsity=50% FF, Evaluation Protocol=0-shot unnormalized2025.08 | 65.93 | |
| LeRaCTraining Regime=LeRaC, Model=LSTM2022.05 | 65.8 | |
| MixLoRABase Model=Gemma 2B, Trainable Parameters=4.3%2024.04 | 65.8 | |
| Zero-Shot-QBackbone=Qwen-2.5-1.5B, Precision=NF42026.05 | 65.8 | |
| LOZOModel=OPT-2.7B2026.05 | 65.6 | |
| GRIFFINBackbone=OPT 6.7B, Sparsity=50% FF, Evaluation Protocol=0-shot unnormalized2025.08 | 65.41 | |
| MeZOBackbone=Qwen-2.5-1.5B, Precision=BF162026.05 | 65.4 | |
| Zero-Shot-QModel Backbone=Llama-3-8B, Model Precision=4 bits2025.05 | 65 | |
| QuZOBackbone=Llama-2-7B, Precision=NF42026.05 | 64.9 | |
| MeZO-GV(LoRA)Backbone=OPT-1.3B, Optimization Method=MeZO, Adaptation Technique=LoRA, Guiding Vectors (GV)=true2026.01 | 64.8 | |
| MeZO-LoRAModel=OPT-13B, Rank (r)=642026.05 | 64.8 | |
| CBSTraining Regime=CBS, Model=LSTM2022.05 | 64.75 | |
| Zero-Shot-QModel Backbone=Llama-2-7B, Model Precision=4 bits2025.05 | 64.6 | |
| AR1-ZOModel=OPT-2.7B, Rank (r)=642026.05 | 64.6 | |
| MeZO-GV(Prefix)Backbone=OPT-1.3B, Optimization Method=MeZO, Adaptation Technique=Prefix Tuning, Guiding Vectors (GV)=true2026.01 | 64.5 | |
| conventionalTraining Regime=conventional, Model=LSTM2022.05 | 64.4 | |
| MeZO-GV(FT)Backbone=OPT-1.3B, Optimization Method=MeZO, Adaptation Technique=Fine-tuning, Guiding Vectors (GV)=true2026.01 | 64.4 | |
| MeZO(LoRA)Backbone=OPT-1.3B, Optimization Method=MeZO, Adaptation Technique=LoRA2026.01 | 63.4 | |
| MeZO(Prefix)Backbone=OPT-1.3B, Optimization Method=MeZO, Adaptation Technique=Prefix Tuning2026.01 | 63 |