Knowledge Evaluation on SuperGPQA Continual
15.85AccuracySTOC
Evaluation Results
| Method | Links | |
|---|---|---|
| STOCModel Scale=1.7B, Prompts=5-shot, Freezing Layers=Best reported2026.05 | 15.85 | |
| STOCModel Scale=0.6B, Prompts=5-shot, Freezing Layers=Best reported2026.05 | 15.24 | |
| LAMOLModel Scale=0.6B, Prompts=5-shot, Freezing Layers=Best reported2026.05 | 13.87 | |
| LAMOLModel Scale=1.7B, Prompts=5-shot, Freezing Layers=Best reported2026.05 | 13.35 | |
| NaiveModel Scale=0.6B, Prompts=5-shot, Freezing Layers=Best reported2026.05 | 10.98 | |
| NaiveModel Scale=1.7B, Prompts=5-shot, Freezing Layers=Best reported2026.05 | 9.6 |