Mathematical Reasoning on GSM8K (Accuracy, Avg.)
66.6GSM8K AccuracyFourierFT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| FourierFTModel Backbone=Llama 3B, n=512, #Param=28.7K, VRAM (GB)=19.80, Runtime (seconds/1K steps)=881.52025.02 | 66.6 | 51 | |
| ULPTModel Backbone=Llama 3B, Rank=256, #Param=8.7K, VRAM (GB)=17.17, Runtime (seconds/1K steps)=589.62025.02 | 66.4 | 50.5 | |
| VeRAModel Backbone=Llama 3B, Rank=8, #Param=115.1K, VRAM (GB)=17.97, Runtime (seconds/1K steps)=621.62025.02 | 65.7 | 49.8 | |
| ULPTModel Backbone=Llama 3B, Rank=2, #Param=6.2K, VRAM (GB)=17.17, Runtime (seconds/1K steps)=589.22025.02 | 65.6 | 50.8 | |
| VeRAModel Backbone=Llama 3B, Rank=1, #Param=114.7K, VRAM (GB)=17.97, Runtime (seconds/1K steps)=620.92025.02 | 65.5 | 50.5 | |
| FourierFTModel Backbone=Llama 3B, n=1024, #Param=57.3K, VRAM (GB)=19.80, Runtime (seconds/1K steps)=882.62025.02 | 65.5 | 50.5 | |
| ULPTModel Backbone=Llama 3B, Rank=64, #Param=6.8K, VRAM (GB)=17.17, Runtime (seconds/1K steps)=589.62025.02 | 65.4 | 51 | |
| PTModel Backbone=Llama 3B, #Param=30.7K, VRAM (GB)=17.18, Runtime (seconds/1K steps)=589.52025.02 | 65.3 | 49.2 | |
| VeRAModel Backbone=Llama 3B, Rank=4, #Param=114.9K, VRAM (GB)=17.97, Runtime (seconds/1K steps)=621.52025.02 | 65 | 49.7 | |
| IA3Model Backbone=Llama 3B, #Param=286.7K, VRAM (GB)=19.15, Runtime (seconds/1K steps)=636.12025.02 | 63.7 | 50.2 | |
| LoRAModel Backbone=Llama 3B, Rank=4, #Param=1.15M, VRAM (GB)=18.45, Runtime (seconds/1K steps)=618.52025.02 | 63.4 | 48.9 | |
| FourierFTModel Backbone=Llama 3B, n=128, #Param=7.2K, VRAM (GB)=19.80, Runtime (seconds/1K steps)=880.92025.02 | 63.1 | 42.5 | |
| LoRAModel Backbone=Llama 3B, Rank=1, #Param=286.7K, VRAM (GB)=18.44, Runtime (seconds/1K steps)=615.82025.02 | 62.9 | 47.5 | |
| ICLModel Backbone=Llama 3B, In-context shots=42025.02 | 62.5 | 43.2 | |
| LoRAModel Backbone=Llama 3B, Rank=8, #Param=2.29M, VRAM (GB)=18.46, Runtime (seconds/1K steps)=621.12025.02 | 62.2 | 50 | |
| ULPTModel Backbone=Llama 1B, Rank=64, #Param=4.7K, VRAM (GB)=9.78, Runtime (seconds/1K steps)=204.22025.02 | 42.4 | 35.6 | |
| ULPTModel Backbone=Llama 1B, Rank=256, #Param=6.7K, VRAM (GB)=9.78, Runtime (seconds/1K steps)=204.62025.02 | 41.4 | 33.9 | |
| VeRAModel Backbone=Llama 1B, Rank=8, #Param=41.2K, VRAM (GB)=10.02, Runtime (seconds/1K steps)=217.02025.02 | 40.9 | 35.2 | |
| LoRAModel Backbone=Llama 1B, Rank=8, #Param=852.0K, VRAM (GB)=10.23, Runtime (seconds/1K steps)=216.52025.02 | 40.2 | 32.5 | |
| PTModel Backbone=Llama 1B, #Param=20.5K, VRAM (GB)=9.79, Runtime (seconds/1K steps)=203.62025.02 | 40.2 | 32.5 | |
| LoRAModel Backbone=Llama 1B, Rank=4, #Param=426.0K, VRAM (GB)=10.22, Runtime (seconds/1K steps)=215.72025.02 | 40.1 | 33.7 | |
| IA3Model Backbone=Llama 1B, #Param=147.5K, VRAM (GB)=10.83, Runtime (seconds/1K steps)=229.22025.02 | 39.7 | 33.1 | |
| ULPTModel Backbone=Llama 1B, Rank=2, #Param=4.1K, VRAM (GB)=9.78, Runtime (seconds/1K steps)=203.82025.02 | 39.7 | 32.9 | |
| VeRAModel Backbone=Llama 1B, Rank=4, #Param=41.1K, VRAM (GB)=10.02, Runtime (seconds/1K steps)=216.62025.02 | 39.6 | 33.7 | |
| VeRAModel Backbone=Llama 1B, Rank=1, #Param=41.0K, VRAM (GB)=10.02, Runtime (seconds/1K steps)=216.02025.02 | 39.3 | 31.9 | |
| LoRAModel Backbone=Llama 1B, Rank=1, #Param=106.5K, VRAM (GB)=10.22, Runtime (seconds/1K steps)=215.12025.02 | 38.5 | 32.6 | |
| FourierFTModel Backbone=Llama 1B, n=1024, #Param=32.8K, VRAM (GB)=10.54, Runtime (seconds/1K steps)=278.52025.02 | 36.6 | 31.3 | |
| FourierFTModel Backbone=Llama 1B, n=128, #Param=4.1K, VRAM (GB)=10.53, Runtime (seconds/1K steps)=278.12025.02 | 35.8 | 28.7 | |
| FourierFTModel Backbone=Llama 1B, n=512, #Param=16.4K, VRAM (GB)=10.53, Runtime (seconds/1K steps)=278.52025.02 | 34.9 | 31.1 | |
| ICLModel Backbone=Llama 1B, In-context shots=42025.02 | 34.3 | 27.7 |