Reasoning on FLenQA 1000 tokens
78.5AccuracyLIME+1
Evaluation Results
| Method | Links | |
|---|---|---|
| LIME+1Size=1B, Matching criteria=generates and matches only one word2025.12 | 78.5 | |
| LIME+1Size=500M, Matching criteria=generates and matches only one word2025.12 | 72.5 | |
| LIME+1Model Size=2B2025.12 | 65.3 | |
| LIME+1Size=2B, Matching criteria=generates and matches only one word2025.12 | 65.3 | |
| LIME+1Model Size=2B2025.12 | 65.3 | |
| LIMEModel Size=2B2025.12 | 47 | |
| LIMESize=2B, Matching criteria=match the first eight generated words2025.12 | 47 | |
| LIMEModel Size=2B2025.12 | 47 | |
| BaseModel Size=2B2025.12 | 34.8 | |
| BaselineSize=2B, Matching criteria=match the first eight generated words2025.12 | 34.8 | |
| Base (DCLM-BASELINE)Model Size=2B2025.12 | 34.8 | |
| LIMESize=1B, Matching criteria=match the first eight generated words2025.12 | 33 | |
| BaselineSize=1B, Matching criteria=match the first eight generated words2025.12 | 32.4 | |
| LIMESize=500M, Matching criteria=match the first eight generated words2025.12 | 30.5 | |
| BaselineSize=500M, Matching criteria=match the first eight generated words2025.12 | 12.5 |