Reasoning on FLenQA 2000 tokens
52.5AccuracyLIME+1
Evaluation Results
| Method | Links | |
|---|---|---|
| LIME+1Model Size=2B2025.12 | 52.5 | |
| LIME+1Size=2B, Matching criteria=generates and matches only one word2025.12 | 52.5 | |
| LIME+1Model Size=2B2025.12 | 52.5 | |
| LIMEModel Size=2B2025.12 | 44.3 | |
| LIMESize=2B, Matching criteria=match the first eight generated words2025.12 | 44.3 | |
| LIMEModel Size=2B2025.12 | 44.3 | |
| BaseModel Size=2B2025.12 | 40.3 | |
| BaselineSize=2B, Matching criteria=match the first eight generated words2025.12 | 40.3 | |
| Base (DCLM-BASELINE)Model Size=2B2025.12 | 40.3 |