Natural Language Understanding on SuperGLUE 1,000 examples (test)
86.7BoolQFT
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| FTModel=Mistral-7b, Optimization=Full-parameter fine-tuning with Adam2024.02 | 86.7 | 87.1 | 71.2 | 86.1 | 95.6 | 91.2 | 86.3 | |
| S-MeZOModel=Mistral-7b, Variant=Sparse MeZO2024.02 | 85.3 | 84.5 | 64.3 | 84.9 | 94.2 | 86.1 | 83.2 | |
| LoRAModel=Mistral-7b2024.02 | 84.8 | 87.4 | 68.2 | 83.9 | 95.6 | 91 | 85.2 | |
| R-MeZOModel=Mistral-7b, Variant=MeZO with Random Mask2024.02 | 84 | 78.7 | 63.2 | 83.1 | 92.4 | 84.1 | 80.9 | |
| MeZO - LoRAModel=Mistral-7b2024.02 | 83.5 | 80.1 | 60.7 | 82.6 | 93.8 | 86.9 | 81.3 | |
| MeZOModel=Mistral-7b2024.02 | 81.6 | 80.9 | 63.2 | 82.7 | 93.8 | 86.7 | 81.5 | |
| ICLModel=Mistral-7b, Strategy=In-Context Learning2024.02 | 76.7 | 78 | 61.4 | 71.3 | 94.6 | 90 | 78.7 | |
| Zero-ShotModel=Mistral-7b2024.02 | 69.3 | 55.2 | 50 | 57.1 | 55.5 | 84 | 61.9 | |
| ICLModel=LLaMA-7b, Protocol=In-Context Learning2024.02 | 67.4 | 54.5 | 52.7 | 58.7 | 81.2 | 84.4 | 66.5 | |
| Zero-ShotModel=LLaMA-7b2024.02 | 65.1 | 49.5 | 50.6 | 55.8 | 79.7 | 59.7 | 60.1 |