Commonsense Reasoning Suite (BoolQ, PIQA, SIQA, Win, OBQA, HellaSwag, ARC-E, ARC-C)
77.5BoolQ AccuracyPEML
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| PEMLBackbone=LLaMA2-7B, Training Setup=Joint training, Prompting=Same as Liu et al. [2024], Yang et al. [2025]2026.05 | 77.5 | 88.7 | 83.7 | 90.2 | 83.8 | 77.4 | 88.5 | 74.1 | 83 | |
| SMoAModel=Llama-3-8B, Params(%)=0.7002, r=32, K=22026.05 | 76.52 | 89.24 | 82.35 | 88.72 | 89.54 | 95.88 | 93.25 | 83.65 | 87.39 | |
| HiRAModel=Llama-3-8B, Params(%)=0.7002, r=322026.05 | 75.4 | 89.7 | 81.15 | 87.7 | 88.32 | 95.36 | 93.27 | 82.9 | 86.73 | |
| DoRAModel=Llama-3-8B, Params(%)=0.7002, r=322026.05 | 74.6 | 89.3 | 79.9 | 85.6 | 85.8 | 95.5 | 90.5 | 80.4 | 85.2 | |
| SMoAModel=Llama-3-8B, Params(%)=0.3513, r=16, K=22026.05 | 74.56 | 87.85 | 81.13 | 86.77 | 87.86 | 95.13 | 92.43 | 82.57 | 86.04 | |
| MeLoRAModel=Llama-3-8B, Params(%)=0.7002, r=32, K=22026.05 | 74.53 | 87.35 | 80.54 | 86.37 | 86.38 | 93.76 | 90.37 | 78.43 | 84.72 | |
| MoRAModel=Llama-3-8B, Params(%)=0.6997, r=322026.05 | 74.28 | 87.43 | 80.71 | 86.74 | 85.6 | 43.53 | 91.16 | 79.61 | 78.63 | |
| HiRAModel=Llama-3-8B, Params(%)=0.3513, r=162026.05 | 73.85 | 89.12 | 81.06 | 86.74 | 87.4 | 94.85 | 93.06 | 82.59 | 86.08 | |
| ChatGPTModel=ChatGPT2026.05 | 73.1 | 85.4 | 68.5 | 66.1 | 74.8 | 78.5 | 89.8 | 79.9 | 77.01 | |
| MeLoRAModel=Llama-3-8B, Params(%)=0.3513, r=16, K=22026.05 | 72.85 | 87.33 | 79.32 | 84.38 | 85.63 | 92.78 | 88.65 | 78.38 | 83.67 | |
| SMoAModel=Llama-2-7B, Params(%)=0.8256, r=32, K=22026.05 | 72.65 | 83.74 | 79.85 | 84.38 | 82.82 | 89.56 | 87.22 | 74.57 | 81.85 | |
| MoRAModel=Llama-2-7B, Params(%)=0.8241, r=322026.05 | 72.17 | 80.79 | 79.53 | 80.19 | 81.2 | 29.09 | 85.31 | 71.42 | 72.46 | |
| DoRABackbone=LLaMA2-7B, Training Setup=Joint training, Prompting=Same as Liu et al. [2024], Yang et al. [2025]2026.05 | 72 | 83.1 | 79.9 | 83 | 81.2 | 89.1 | 84.5 | 71 | 80.5 | |
| DoRAModel=Llama-2-7B, Params(%)=0.8256, r=322026.05 | 71.8 | 83.7 | 76 | 82.6 | 82.4 | 89.1 | 83.7 | 68.2 | 79.69 | |
| MeLoRAModel=Llama-2-7B, Params(%)=0.8256, r=32, K=22026.05 | 71.43 | 82.83 | 77.65 | 81.35 | 81.35 | 87.52 | 85.38 | 71.89 | 79.93 | |
| HiRAModel=Llama-2-7B, Params(%)=0.8256, r=322026.05 | 71.22 | 83.35 | 79.53 | 83.98 | 84.6 | 88.12 | 86.74 | 73.81 | 81.42 | |
| MTL-LoRABackbone=LLaMA2-7B, Training Setup=Joint training, Prompting=Same as Liu et al. [2024], Yang et al. [2025]2026.05 | 71 | 84.4 | 80.8 | 84.9 | 82.6 | 93.1 | 87 | 73.4 | 82.1 | |
| LoRAModel=Llama-3-8B, Params(%)=0.7002, r=322026.05 | 70.8 | 85.2 | 79.9 | 84.3 | 79 | 91.7 | 84.2 | 71.2 | 80.79 | |
| SMoAModel=Llama-2-7B, Params(%)=0.4128, r=16, K=22026.05 | 70.13 | 81.27 | 78.31 | 83.68 | 81.03 | 87.42 | 86.41 | 72.53 | 80.1 | |
| MeLoRAModel=Llama-2-7B, Params(%)=0.4128, r=16, K=22026.05 | 70.12 | 80.65 | 75.35 | 78.23 | 80.28 | 86.05 | 82.55 | 67.67 | 77.61 | |
| HiRAModel=Llama-2-7B, Params(%)=0.4128, r=162026.05 | 69.82 | 80.2 | 78.2 | 83.43 | 81 | 86.99 | 85.9 | 71.33 | 79.61 | |
| LoRABackbone=LLaMA2-7B, Training Setup=Joint training, Prompting=Same as Liu et al. [2024], Yang et al. [2025]2026.05 | 69.8 | 79.9 | 79.5 | 82.6 | 81 | 83.6 | 79.8 | 64.7 | 77.6 | |
| LoRAModel=Llama-2-7B, Params(%)=0.8256, r=322026.05 | 69.8 | 79.9 | 79.5 | 82.6 | 81 | 83.6 | 79.8 | 64.7 | 77.61 | |
| SSMLoRAModel=Llama-3-8B, Params(%)=0.6813, r=322026.05 | 68.57 | 84.38 | 77.95 | 84.33 | 78.87 | 90.5 | 85.38 | 70.81 | 80.1 | |
| SSMLoRAModel=Llama-2-7B, Params(%)=0.8024, r=322026.05 | 68.35 | 76.57 | 78.58 | 82.38 | 80.68 | 82.78 | 78.85 | 64.57 | 76.6 | |
| MoELoRABackbone=LLaMA2-7B, Training Setup=Joint training, Prompting=Same as Liu et al. [2024], Yang et al. [2025]2026.05 | 68 | 83.5 | 70.4 | 82.5 | 83.2 | 90.6 | 86.8 | 61.5 | 78.3 | |
| MultiLoRABackbone=LLaMA2-7B, Training Setup=Joint training, Prompting=Same as Liu et al. [2024], Yang et al. [2025]2026.05 | 66.5 | 65.8 | 62.8 | 79.3 | 75.4 | 79.2 | 76.7 | 59.6 | 70.7 | |
| Gated DeltaNet-2Architecture Group=Attention or hybrid models, Evaluation Protocol=zero-shot2026.05 | 62.57 | 72.2 | 41.5 | 58.56 | 33 | 58.46 | 71.89 | 36.69 | 53.97 | |
| KDAArchitecture Group=Attention or hybrid models, Evaluation Protocol=zero-shot2026.05 | 62.03 | 71.06 | 40.53 | 57.77 | 30 | 56.89 | 71.59 | 35.07 | 52.68 | |
| KDAArchitecture Group=Recurrent models, Evaluation Protocol=zero-shot2026.05 | 60.67 | 72.09 | 40.99 | 55.72 | 30.4 | 55.75 | 70.83 | 35.92 | 52.28 | |
| Mamba-2Architecture Group=Recurrent models, Evaluation Protocol=zero-shot2026.05 | 60.19 | 72.58 | 40.63 | 55.33 | 31 | 55.51 | 70.68 | 35.26 | 51.82 | |
| Gated DeltaNetArchitecture Group=Attention or hybrid models, Evaluation Protocol=zero-shot2026.05 | 60 | 70.06 | 40.97 | 56.83 | 30.6 | 57.5 | 70.41 | 35.15 | 52.25 | |
| P-TuningModel=Llama-3-8B, Params(%)=0.62402026.05 | 59.97 | 11.64 | 8.19 | 37.65 | 9.6 | 1.77 | 8.63 | 7.42 | 18.11 | |
| Gated DeltaNet-2Architecture Group=Recurrent models, Evaluation Protocol=zero-shot2026.05 | 59.54 | 72.8 | 40.58 | 57.85 | 31.6 | 56.84 | 72.43 | 38.23 | 53.11 | |
| TransformerArchitecture Group=Attention or hybrid models, Evaluation Protocol=zero-shot2026.05 | 59.42 | 70.21 | 39.74 | 55.85 | 25 | 56.12 | 69.23 | 33.84 | 50.86 | |
| Mamba-2Architecture Group=Attention or hybrid models, Evaluation Protocol=zero-shot2026.05 | 59.31 | 71.47 | 40.35 | 56.17 | 29.8 | 57.52 | 70.5 | 34.73 | 51.99 | |
| Gated DeltaNetArchitecture Group=Recurrent models, Evaluation Protocol=zero-shot2026.05 | 58.78 | 72.31 | 40.53 | 56.75 | 30.2 | 56.5 | 68.81 | 35.15 | 52.07 | |
| P-TuningModel=Llama-2-7B, Params(%)=0.74282026.05 | 58.75 | 36.02 | 0.2 | 0 | 0.8 | 0.01 | 1.98 | 0.17 | 12.24 | |
| Mamba-3 (MIMO)Architecture Group=Attention or hybrid models, Evaluation Protocol=zero-shot2026.05 | 57.98 | 71.98 | 40.99 | 57.06 | 29.4 | 58.19 | 70.54 | 38.48 | 52.72 | |
| Mamba-3 (SISO)Architecture Group=Attention or hybrid models, Evaluation Protocol=zero-shot2026.05 | 57.86 | 71.01 | 41.2 | 57.3 | 32 | 58.75 | 70.54 | 36.35 | 52.69 | |
| Mamba-3 (MIMO)Architecture Group=Recurrent models, Evaluation Protocol=zero-shot2026.05 | 57.74 | 72.36 | 40.89 | 55.78 | 30 | 56.49 | 72.38 | 38.07 | 52.39 | |
| PromptTuningModel=Llama-3-8B, Params(%)=0.00102026.05 | 56.85 | 45.05 | 36.13 | 50.12 | 29.2 | 14.01 | 32.74 | 31.57 | 36.96 | |
| PromptTuningModel=Llama-2-7B, Params(%)=0.00122026.05 | 55.93 | 12.35 | 30.5 | 40.57 | 9.4 | 6.91 | 8.63 | 6.06 | 21.29 | |
| Mamba-3 (SISO)Architecture Group=Recurrent models, Evaluation Protocol=zero-shot2026.05 | 55.9 | 72.31 | 41.76 | 56.2 | 31 | 55.58 | 70.45 | 34.56 | 51.42 |