Reasoning on FLenQA 250 tokens
80AccuracyLIME+1
Evaluation Results
| Method | Links | |
|---|---|---|
| LIME+1Model Size=2B2025.12 | 80 | |
| LIME+1Size=2B, Matching criteria=generates and matches only one word2025.12 | 80 | |
| LIME+1Size=1B, Matching criteria=generates and matches only one word2025.12 | 80 | |
| LIME+1Model Size=2B2025.12 | 80 | |
| LIME+1Size=500M, Matching criteria=generates and matches only one word2025.12 | 70 | |
| LIMEModel Size=2B2025.12 | 52 | |
| LIMESize=2B, Matching criteria=match the first eight generated words2025.12 | 52 | |
| LIMEModel Size=2B2025.12 | 52 | |
| BaseModel Size=2B2025.12 | 42 | |
| BaselineSize=2B, Matching criteria=match the first eight generated words2025.12 | 42 | |
| Base (DCLM-BASELINE)Model Size=2B2025.12 | 42 | |
| LIMESize=500M, Matching criteria=match the first eight generated words2025.12 | 40 | |
| BaselineSize=1B, Matching criteria=match the first eight generated words2025.12 | 36 | |
| LIMESize=1B, Matching criteria=match the first eight generated words2025.12 | 28 | |
| BaselineSize=500M, Matching criteria=match the first eight generated words2025.12 | 22 |