Effect-size forecasting on Query2Effect non-health cross-sector (test ood)
0.2211RMSEGold RCT → ModernBERTr
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Gold RCT → ModernBERTrprotocol=Supervised, representation=Gold RCT2026.05 | 0.2211 | 0.1302 | 0.0628 | 0.271 | 0.2618 | 73.95 | 60.3 | 63.61 | |
| Synthetic-RCT(GPT-5.2) → ModernBERTrprotocol=Supervised, pipeline=Synthetic-RCT, intermediate_llm=GPT-5.22026.05 | 0.2241 | 0.1375 | 0.0591 | 0.265 | 0.223 | 73.33 | 57.88 | 63.56 | |
| Synthetic-RCT(GPT-OSS-20B) → ModernBERTrprotocol=Supervised, pipeline=Synthetic-RCT, intermediate_llm=GPT-OSS-20B2026.05 | 0.231 | 0.1402 | 0.0019 | 0.1818 | 0.1589 | 74.93 | 57.12 | 63.23 | |
| mean-effecttype=baseline2026.05 | 0.2444 | 0.181 | -0.1205 | — | — | 72.32 | 47.5 | 43.5 | |
| ModernBERTqprotocol=Supervised2026.05 | 0.2578 | 0.1773 | -0.0603 | 0.1095 | 0.1138 | 71.23 | 55.7 | 58.23 | |
| GPT-5.2protocol=Prompted2026.05 | 0.2782 | 0.2089 | -0.4513 | 0.2503 | 0.252 | 66.31 | 48.99 | 57.8 | |
| Gemini 2.5 Flashprotocol=Prompted2026.05 | 0.3793 | 0.3027 | -1.6972 | 0.2026 | 0.1988 | 67.66 | 45.64 | 43.77 | |
| retrieval-lookuptype=baseline2026.05 | 0.4084 | 0.2588 | -2.1269 | 0.0279 | 0.0379 | 68.23 | 45.23 | 44.02 | |
| GPT-OSS 120Bprotocol=Prompted2026.05 | 0.4505 | 0.3936 | -2.8057 | 0.1659 | 0.1918 | 66.78 | 45.04 | 34 |