Large Language Model Evaluation on RULER, LongBench-paragraph, and HumanEval-WildChat (Primary sweep)
0Mean Quality DeltaMerlin
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| MerlinCells evaluated=40 (24 RULER + 12 LongBench-paragraph + 4 HumanEval-WildChat)2026.05 | 0 | -4 | 0 | 16,000 | 82 |