Knowledge Evaluation on SuperGPQA (Original)
11.01AccuracySTOC
Evaluation Results
| Method | Links | |
|---|---|---|
| STOCModel Scale=1.7B, Prompts=5-shot, Freezing Layers=Best reported2026.05 | 11.01 | |
| STOCModel Scale=0.6B, Prompts=5-shot, Freezing Layers=Best reported2026.05 | 10.76 | |
| LAMOLModel Scale=1.7B, Prompts=5-shot, Freezing Layers=Best reported2026.05 | 10.63 | |
| LAMOLModel Scale=0.6B, Prompts=5-shot, Freezing Layers=Best reported2026.05 | 10.6 | |
| NaiveModel Scale=0.6B, Prompts=5-shot, Freezing Layers=Best reported2026.05 | 9.48 | |
| NaiveModel Scale=1.7B, Prompts=5-shot, Freezing Layers=Best reported2026.05 | 9.4 |