Text Classification on RTE
84.84AccuracyFirst Order Adamw FT
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| First Order Adamw FTModel Backbone=Llama2-7B, Fine-tuning Strategy=Full Fine-tuning, Optimization Algorithm=First Order Adamw, Number of training examples=10002026.06 | 84.84 | — | — | |
| theta_B fine-tune|Ds^c|=-2025.10 | 84.4 | — | — | |
| KNASsearch_space=Highway, Look-ahead, DenseNet2021.11 | 83.75 | — | — | |
| ROBERTA-largepre-trained=true2021.11 | 83.51 | — | — | |
| FO-PromptModel=Vicuna-7b-v1.52026.04 | 82.3 | — | — | |
| Hybrid-LoRAModel=Vicuna-7b-v1.52026.04 | 82 | — | — | |
| Hybrid-PrefixModel=Vicuna-7b-v1.52026.04 | 80.9 | — | — | |
| FO-LoRAModel=Vicuna-7b-v1.52026.04 | 80.1 | — | — | |
| Fine-tuningModel Backbone=OPT-6.7B, Model Precision=16 bits, Memory Profiling=26.8GB2025.05 | 79.8 | — | — | |
| SubZero-GV (Prefix)Optimization Algorithm=SubZero, Tuning Strategy=Prefix Tuning, Guiding Vectors=true, Model=OPT-13B2026.01 | 76.2 | — | — | |
| LeRaCTraining Regime=LeRaC, Model=BERT_large2022.05 | 75.81 | — | — | |
| SubZero-GV (LoRA)Optimization Algorithm=SubZero, Tuning Strategy=LoRA, Guiding Vectors=true, Model=OPT-13B2026.01 | 75.8 | — | — | |
| SubZero (LoRA)Optimization Algorithm=SubZero, Tuning Strategy=LoRA, Guiding Vectors=false, Model=OPT-13B2026.01 | 75.5 | — | — | |
| CBSTraining Regime=CBS, Model=BERT_large2022.05 | 74.97 | — | — | |
| SubZero-GV (FT)Optimization Algorithm=SubZero, Tuning Strategy=Full Tuning, Guiding Vectors=true, Model=OPT-13B2026.01 | 74.8 | — | — | |
| MeZO-GV (Prefix)Optimization Algorithm=MeZO, Tuning Strategy=Prefix Tuning, Guiding Vectors=true, Model=OPT-13B2026.01 | 74.8 | — | — | |
| conventionalTraining Regime=conventional, Model=BERT_large2022.05 | 74.48 | — | — | |
| SubZero (FT)Optimization Algorithm=SubZero, Tuning Strategy=Full Tuning, Guiding Vectors=false, Model=OPT-13B2026.01 | 74 | — | — | |
| SubZero (Prefix)Optimization Algorithm=SubZero, Tuning Strategy=Prefix Tuning, Guiding Vectors=false, Model=OPT-13B2026.01 | 73.6 | — | — | |
| MeZO-GV (FT)Optimization Algorithm=MeZO, Tuning Strategy=Full Tuning, Guiding Vectors=true, Model=OPT-13B2026.01 | 73.5 | — | — | |
| theta_B + delta*|Ds^c|=-2025.10 | 72.93 | — | — | |
| ZO-AdaMU (2x)Optimization Algorithm=ZO-AdaMU, Tuning Strategy=Full Tuning, Training Steps=2x, Model=OPT-13B2026.01 | 72.9 | — | — | |
| MeZO-GV (LoRA)Optimization Algorithm=MeZO, Tuning Strategy=LoRA, Guiding Vectors=true, Model=OPT-13B2026.01 | 72.6 | — | — | |
| MeZOBackbone=Qwen-2.5-1.5B, Precision=BF162026.05 | 72.1 | — | — | |
| ZO-AdaMU (LoRA)Optimization Algorithm=ZO-AdaMU, Tuning Strategy=LoRA, Model=OPT-13B2026.01 | 72 | — | — | |
| HiZOO (Prefix)Optimization Algorithm=HiZOO, Tuning Strategy=Prefix Tuning, Model=OPT-13B2026.01 | 71.8 | — | — | |
| Fine-tuningModel Backbone=Llama-3-8B, Model Precision=16 bits, Memory Profiling=31.9GB2025.05 | 71.5 | — | — | |
| MeZO (Prefix)Optimization Algorithm=MeZO, Tuning Strategy=Prefix Tuning, Guiding Vectors=false, Model=OPT-13B2026.01 | 70.8 | — | — | |
| FTOptimization Algorithm=Adam, Tuning Strategy=Full Tuning, Model=OPT-13B2026.01 | 70.8 | — | — | |
| FO-PrefixModel=Vicuna-7b-v1.52026.04 | 70.4 | — | — | |
| Hybrid-PromptModel=Vicuna-7b-v1.52026.04 | 70.1 | — | — | |
| MeZOModel Backbone=Llama-3-8B, Model Precision=16 bits, Memory Profiling=20.5GB2025.05 | 70 | — | — | |
| HiZOOOptimization Algorithm=HiZOO, Tuning Strategy=Full Tuning, Model=OPT-13B2026.01 | 69.3 | — | — | |
| BERT baseParams [M]=345, Operating Point=12026.01 | 69.1 | — | 0 | |
| BERT baseParams [M]=345, Operating Point=22026.01 | 69.1 | — | 0 | |
| ADEPTParams [M]=42, Operating Point=22026.01 | 68.8 | — | -53 | |
| F-PABEEParams [M]=258, Operating Point=22026.01 | 68.1 | — | -47 | |
| MeZO (LoRA)Optimization Algorithm=MeZO, Tuning Strategy=LoRA, Guiding Vectors=false, Model=OPT-13B2026.01 | 67.9 | — | — | |
| PABEEParams [M]=258, Operating Point=22026.01 | 67.7 | — | -46 | |
| Dominant-layer ZO FTModel Backbone=Llama2-7B, Fine-tuning Strategy=Full Fine-tuning, Optimization Algorithm=Dominant-layer ZO, Number of training examples=10002026.06 | 67.51 | — | — | |
| HiZOO (LoRA)Optimization Algorithm=HiZOO, Tuning Strategy=LoRA, Model=OPT-13B2026.01 | 67.5 | — | — | |
| Dominant-layer ZO LoRAModel Backbone=Llama2-7B, Fine-tuning Strategy=LoRA, Optimization Algorithm=Dominant-layer ZO, Number of training examples=10002026.06 | 67.44 | — | — | |
| BranchyNetParams [M]=258, Operating Point=22026.01 | 67.4 | — | -47 | |
| Shallow-DeepParams [M]=258, Operating Point=22026.01 | 67.2 | — | -48 | |
| MeZO-GV(Prefix)Backbone=OPT-1.3B, Optimization Method=MeZO, Adaptation Technique=Prefix Tuning, Guiding Vectors (GV)=true2026.01 | 66.8 | — | — | |
| QZOModel Backbone=Llama-3-8B, Model Precision=4 bits, Memory Profiling=6.3GB2025.05 | 66.8 | — | — | |
| MeZO (FT)Optimization Algorithm=MeZO, Tuning Strategy=Full Tuning, Guiding Vectors=false, Model=OPT-13B2026.01 | 66.1 | — | — | |
| DeeBERTParams [M]=258, Operating Point=22026.01 | 65.9 | — | -33 | |
| MeZO LoRAModel Backbone=Llama2-7B, Fine-tuning Strategy=LoRA, Optimization Algorithm=MeZO, Number of training examples=10002026.06 | 65.89 | — | — | |
| MeZO(Prefix)Backbone=OPT-1.3B, Optimization Method=MeZO, Adaptation Technique=Prefix Tuning2026.01 | 65.7 | — | — | |
| MeZO FTModel Backbone=Llama2-7B, Fine-tuning Strategy=Full Fine-tuning, Optimization Algorithm=MeZO, Number of training examples=10002026.06 | 65.34 | — | — | |
| MeZOModel Backbone=OPT-6.7B, Model Precision=16 bits, Memory Profiling=14.8GB2025.05 | 64.6 | — | — | |
| GPT-3# Params=175B, prompting_mode=zero-shot2021.07 | 63.5 | — | — | |
| Fine-tuningModel Backbone=Llama-2-7B, Model Precision=16 bits, Memory Profiling=26.0GB2025.05 | 63.2 | — | — | |
| MeZO-GV(LoRA)Backbone=OPT-1.3B, Optimization Method=MeZO, Adaptation Technique=LoRA, Guiding Vectors (GV)=true2026.01 | 62.8 | — | — | |
| Hybrid-LoRAModel=Llama-2-7b2026.04 | 62.5 | — | — | |
| Hybrid-PromptModel=OPT-1.3b2026.04 | 62.5 | — | — | |
| ICLEvaluation Protocol=In-context Learning, Model=OPT-13B2026.01 | 62.1 | — | — | |
| FO-LoRAModel=Llama-2-7b2026.04 | 62.1 | — | — | |
| ZO-AdaMU (Prefix)Optimization Algorithm=ZO-AdaMU, Tuning Strategy=Prefix Tuning, Model=OPT-13B2026.01 | 61.8 | — | — | |
| Zero-shot w/o finetuneModel Backbone=Llama2-7B, Fine-tuning Strategy=None, Optimization Algorithm=None, Number of training examples=10002026.06 | 61.73 | — | — | |
| MeZO(LoRA)Backbone=OPT-1.3B, Optimization Method=MeZO, Adaptation Technique=LoRA2026.01 | 61.7 | — | — | |
| QZOModel Backbone=OPT-6.7B, Model Precision=4 bits, Memory Profiling=4.8GB2025.05 | 61.7 | — | — | |
| Zero-ShotModel Backbone=Llama-2-7B, Model Precision=16 bits2025.05 | 61.7 | — | — | |
| CAQ-ZOBackbone=Qwen-2.5-1.5B, Precision=NF42026.05 | 61.5 | — | — | |
| Hybrid-LoRAModel=OPT-1.3b2026.04 | 61 | — | — | |
| ADEPTParams [M]=42, Operating Point=12026.01 | 60.8 | — | -75 | |
| MeZO-GV(FT)Backbone=OPT-1.3B, Optimization Method=MeZO, Adaptation Technique=Fine-tuning, Guiding Vectors (GV)=true2026.01 | 60.6 | — | — | |
| FO-PrefixModel=Llama-2-7b2026.04 | 60.6 | — | — | |
| Hybrid-PrefixModel=Llama-2-7b2026.04 | 60.6 | — | — | |
| FO-PromptModel=Llama-2-7b2026.04 | 59.9 | — | — | |
| Hybrid-PromptModel=Llama-2-7b2026.04 | 59.9 | — | — | |
| DistilBERTParams [M]=258, Operating Point=22026.01 | 59.7 | — | -40 | |
| Zero-shotEvaluation Protocol=Zero-shot, Model=OPT-13B2026.01 | 59.6 | — | — | |
| QZOModel Backbone=Llama-2-7B, Model Precision=4 bits, Memory Profiling=5.0GB2025.05 | 59.2 | — | — | |
| MeZOModel Backbone=Llama-2-7B, Model Precision=16 bits, Memory Profiling=14.8GB2025.05 | 58.1 | — | — | |
| MeZOBackbone=Llama-2-7B, Precision=BF162026.05 | 58.1 | — | — | |
| MeZO(FT)Backbone=OPT-1.3B, Optimization Method=MeZO, Adaptation Technique=Fine-tuning2026.01 | 57.4 | — | — | |
| CAQ-ZOBackbone=Llama-2-7B, Precision=NF42026.05 | 57.4 | — | — | |
| QuZOBackbone=Llama-2-7B, Precision=NF42026.05 | 56.5 | — | — | |
| F-PABEEParams [M]=258, Operating Point=12026.01 | 56 | — | -76 | |
| PABEEParams [M]=258, Operating Point=12026.01 | 55.8 | — | -75 | |
| LeRaCTraining Regime=LeRaC, Model=LSTM2022.05 | 55.71 | — | — | |
| Zero-Shot-QBackbone=Llama-2-7B, Precision=NF42026.05 | 55.6 | — | — | |
| Zero-ShotModel Backbone=OPT-6.7B, Model Precision=16 bits2025.05 | 55.2 | — | — | |
| FO-LoRAModel=OPT-1.3b2026.04 | 54.8 | — | — | |
| BranchyNetParams [M]=258, Operating Point=12026.01 | 54.7 | — | -76 | |
| Shallow-DeepParams [M]=258, Operating Point=12026.01 | 54.7 | — | -76 | |
| QZOModel=Llama-2-13B, Precision=2 bits, Memory Profiling=5.78GB, Evaluation Protocol=Zeroth-order optimization2025.05 | 54.5 | — | — | |
| GradFix (theta_B + delta^A)|Ds^c|=502025.10 | 54.25 | — | — | |
| conventionalTraining Regime=conventional, Model=LSTM2022.05 | 54.12 | — | — | |
| CBSTraining Regime=CBS, Model=LSTM2022.05 | 54.03 | — | — | |
| Zero-Shot-QModel Backbone=OPT-6.7B, Model Precision=4 bits2025.05 | 53.8 | — | — | |
| Zero-Shot-QBackbone=Qwen-2.5-1.5B, Precision=NF42026.05 | 53.8 | — | — | |
| ICLBackbone=OPT-1.3B, Learning Strategy=In-context Learning2026.01 | 53.4 | — | — | |
| Zero-Shot-QModel Backbone=Llama-2-7B, Model Precision=4 bits2025.05 | 53.4 | — | — | |
| Zero-shotBackbone=OPT-1.3B, Learning Strategy=Zero-shot2026.01 | 53.1 | — | — | |
| Zero-Shot-QModel=Llama-2-13B, Precision=2 bits, Memory Profiling=N/A, Evaluation Protocol=Zero-shot2025.05 | 53.1 | — | — | |
| Hybrid-PrefixModel=OPT-1.3b2026.04 | 52.7 | — | — | |
| QuZOBackbone=Qwen-2.5-1.5B, Precision=NF42026.05 | 52.6 | — | — |