ELI5
Benchmarks
Task NameDataset NameSOTA ResultTrendResults
ELI5
31.15ROUGE-L
57
ELI5 (test)
27.13ROUGE-L
54
eli5-category (test)
1.313PPL
28
ELI5 informal
0.77Percentage of AmE Generations
20
ELI5
25.9Claim Correctness Score
19
ELI5
91.79Accuracy
16
ELI5
100AUC
16
ELI5
55.8Precision
15
ELI5 (test)
14.1MAE
14
ELI5 Wiki-answerable
26.6ROUGE-L Score
14
ELI5 Domain Generalization (DG-MGT)
91.79Accuracy
12
ELI5 (test)
0.08MAE
12
ELI5 (test)
10.33MAE
12
ELI5 (val)
31.5F1
11
ELI5 (test)
81.9Citation Recall
10
ELI5 (test)
47.43Fluency (mauve)
10
ELI5 (dev)
26.6ROUGE-L Score
9
ELI5 16-bit bitstrings LLAMA3.1-8B
77.8Message Accuracy
8
ELI5 KILT (test)
25.4F1
8
ELI5
0.186Correctness
8
ELI5 KILT (test)
11Retrieval Precision
8
ELI5
20.75R-L
8
ELI5 prompts 32-bit payload Llama3.1-8B (test)
50.9Bit Accuracy
7
ELI5 prompts 32-bit payload Llama3.1-8B (test)
70.7Bit Accuracy (10% Substitution)
7
ELI5 prompts 32-bit payload Llama3.1-8B (test)
70.7Bit Accuracy (10% Deletion)
7