Effect-size forecasting on Query2Effect level-0 queries (test)
0.2449RMSEGold RCT → ModernBERTr
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Gold RCT → ModernBERTrEvaluation Protocol=Supervised (Ceiling), Input/Training Source=Gold RCT, Backbone=ModernBERT-large2026.05 | 0.2449 | 0.1574 | 0.2342 | 0.4844 | 0.389 | 0.734 | 0.6197 | 0.687 | |
| Synthetic-RCT(GPT-5.2) → ModernBERTrEvaluation Protocol=Supervised, Input/Training Source=Synthetic-RCT (GPT-5.2), Backbone=ModernBERT-large2026.05 | 0.2601 | 0.1722 | 0.1359 | 0.382 | 0.3026 | 0.7308 | 0.583 | 0.6802 | |
| ModernBERTqEvaluation Protocol=Supervised, Input/Training Source=Policy queries, Backbone=ModernBERT-large2026.05 | 0.2659 | 0.1743 | 0.097 | 0.348 | 0.2988 | 0.7182 | 0.5771 | 0.6678 | |
| Synthetic-RCT(GPT-OSS-20B) → ModernBERTrEvaluation Protocol=Supervised, Input/Training Source=Synthetic-RCT (GPT-OSS-20B), Backbone=ModernBERT-large2026.05 | 0.2685 | 0.179 | 0.0807 | 0.3349 | 0.3002 | 0.7263 | 0.5771 | 0.6539 | |
| mean-effectEvaluation Protocol=Baseline2026.05 | 0.2788 | 0.1946 | -0.0797 | — | — | 0.7302 | 0.4846 | 0.453 | |
| GPT-5.2Evaluation Protocol=Prompted2026.05 | 0.3165 | 0.2374 | -0.2801 | 0.27 | 0.3008 | 0.6523 | 0.5233 | 0.6222 | |
| retrieval-lookupEvaluation Protocol=Baseline2026.05 | 0.388 | 0.2526 | -0.9235 | 0.2423 | 0.164 | 0.678 | 0.51 | 0.533 | |
| Gemini 2.5 FlashEvaluation Protocol=Prompted2026.05 | 0.4503 | 0.3542 | -1.5908 | 0.1711 | 0.2014 | 0.5936 | 0.5136 | 0.4137 | |
| GPT-OSS 120BEvaluation Protocol=Prompted2026.05 | 0.49 | 0.416 | -2.0682 | 0.1977 | 0.2213 | 0.598 | 0.4888 | 0.3298 |