Long-context understanding on InfiniteBench (test)
36.7En QA F1Llama 3 70B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Llama 3 70B2024.07 | 36.7 | 78.2 | |
| Llama 3 405B2024.07 | 30.5 | 83.4 | |
| Llama 3 8B2024.07 | 27.1 | 65.1 | |
| GPT-4o2024.07 | 19.1 | 82.5 | |
| GPT-42024.07 | 15.7 | 72 | |
| Claude 3.5 Sonnet2024.07 | 11.3 | — |