ResearchBenchmarksRubric Generation on OpenAI HealthBench HARDFollow4.2BLEUSFT1.3922.1212.853.579Mar 21, 2026Evaluation ResultsMethodMethodLinksBLEUROUGE-1ROUGE-2ROUGE-LLLM Judge ScoreSFTMode=SFTMode=SFT2026.034.229.99.824.143.6RubricRAGThinking=nothinkThinking=nothink2026.033.931.110.325.152.3RubricRAGThinking=thinkThinking=think2026.03328.38.522.551.4Few-ShotMode=Few-ShotMode=Few-Shot2026.032.927.78.32250.5GRPOMode=GRPOMode=GRPO2026.032.729.18.923.150.6Zero-ShotMode=Zero-ShotMode=Zero-Shot2026.031.521.65.717.747.4