Reasoning on FLenQA 3000 tokens
39.3AccuracyLIME+1
Evaluation Results
| Method | Links | |
|---|---|---|
| LIME+1Model Size=2B2025.12 | 39.3 | |
| LIME+1Size=2B, Matching criteria=generates and matches only one word2025.12 | 39.3 | |
| LIME+1Model Size=2B2025.12 | 39.3 | |
| LIMEModel Size=2B2025.12 | 30 | |
| LIMESize=2B, Matching criteria=match the first eight generated words2025.12 | 30 | |
| LIMEModel Size=2B2025.12 | 30 | |
| BaseModel Size=2B2025.12 | 28.2 | |
| BaselineSize=2B, Matching criteria=match the first eight generated words2025.12 | 28.2 | |
| Base (DCLM-BASELINE)Model Size=2B2025.12 | 28.2 |