Prompts
Benchmarks
Task NameDataset NameSOTA ResultTrendResults
prompts 500 sampled
0.335HPSv2
36
20 Prompts across 4 Task Categories
6.55Mean Expected Tokens per Speculation Step
20
500 randomly sampled prompts
81Similarity
16
Prompts 100 (evaluation)
94.1Distinct-N (WM)
14
160 prompts NQD (test)
8.67Character Development
13
1000 prompts (test)
88.2Succ. @≥ 1
8
prompts 10 randomly sampled
2.232Inference Time (s)
6
Prompts (test)
0.48MAP
6
1000 prompts (held-out)
1.1CPRSD l(f)
5
1,024 prompts (held-out)
4.81VQ
5
14 prompts 1000 panoramas of dimensions 512x4608
0.58Intra-LPIPS
4
400 prompts (test)
29.053HPSv2
4
1,172 Prompts (test)
677Win Count (CS)
3
1,172 prompts (test)
695CS Wins
3
50 prompts
1.667LogFreq (d)
3
15 distinct single-task prompts
10.63LLM Time
3
1000 prompts (test)
1Usefulness Score
2
180 held-out prompts Length Far OOD 8–12k
72Length Adherence Ratio
1
180 held-out prompts Length Near OOD 4–8k
87Length Adherence Ratio
1
180 Prompts Length ID 1–4k (test)
99Length Adherence Ratio
1
180 held-out prompts Far OOD 8–12k
44.1Story Quality Score
1
180 prompts Near OOD 4–8k (held-out)
48.2Story Quality Score
1