Medical Question Answering on PubMedQA
86Pass@1MedXIAOHE
Evaluation Results
| Method | Links | |
|---|---|---|
| MedXIAOHEMode=Thinking mode, Decoding=Greedy2026.02 | 86 | |
| Gemini 3.0 ProDecoding=Greedy2026.02 | 80.8 | |
| GPT-5.2 ThinkingMode=Thinking mode, Decoding=Greedy2026.02 | 79.8 | |
| Gemini 2.5 ProDecoding=Greedy2026.02 | 75.6 | |
| SHIFTTrain Size=0.1%2026.05 | 74.2 | |
| SC-EntropyTrain Size=0.2%2026.05 | 72.4 | |
| SC-EntropyTrain Size=0.1%2026.05 | 71.6 | |
| CoT SimilarityTrain Size=0.1%2026.05 | 71.4 | |
| ClusterTrain Size=0.1%2026.05 | 71 | |
| CoT SimilarityTrain Size=0.2%2026.05 | 71 | |
| SHIFTTrain Size=0.2%2026.05 | 70.8 | |
| Random (avg)Train Size=0.2%2026.05 | 70.72 | |
| Random (avg)Train Size=0.1%2026.05 | 70.56 | |
| CoreSetTrain Size=0.2%2026.05 | 70.4 | |
| Q-PPLTrain Size=0.2%2026.05 | 70 | |
| Qwen3-1.7BTrain Size=100%2026.05 | 69.6 | |
| Q-PPLTrain Size=0.1%2026.05 | 69.6 | |
| A-PPLTrain Size=0.2%2026.05 | 69.2 | |
| A-PPLTrain Size=0.1%2026.05 | 69 | |
| ClusterTrain Size=0.2%2026.05 | 69 | |
| CoreSetTrain Size=0.1%2026.05 | 68.6 | |
| Qwen3-1.7BTrain Size=N/A2026.05 | 66.2 | |
| Qwen2.5-7B + GRPO w/ VERL.Base Model=Qwen2.5-7B, Training=GRPO w/ VERL.2025.09 | 47.8 | |
| Qwen2.5-7B + GRPOBase Model=Qwen2.5-7B, Training=GRPO2025.09 | 46.6 | |
| Qwen2.5-7B + PPOBase Model=Qwen2.5-7B, Training=PPO2025.09 | 45.8 | |
| Qwen2.5-7B + PPO w/ VERL.Base Model=Qwen2.5-7B, Training=PPO w/ VERL.2025.09 | 45.4 | |
| Qwen2.5-7BBase Model=Qwen2.5-7B2025.09 | 43.6 | |
| Mathstral-7B-v0.1 + GRPO w/ VERL.Base Model=Mathstral-7B-v0.1, Training=GRPO w/ VERL.2025.09 | 24.4 | |
| Mathstral-7B-v0.1 + PPOBase Model=Mathstral-7B-v0.1, Training=PPO2025.09 | 23.1 | |
| Mathstral-7B-v0.1 + GRPOBase Model=Mathstral-7B-v0.1, Training=GRPO2025.09 | 22.6 | |
| Mathstral-7B-v0.1Base Model=Mathstral-7B-v0.12025.09 | 22.4 | |
| Mathstral-7B-v0.1 + PPO w/ VERL.Base Model=Mathstral-7B-v0.1, Training=PPO w/ VERL.2025.09 | 22.3 |