Medical Reasoning on Medical-O1-Reasoning-SFT (test)
0.5127WinsLLM-AutoDP
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| LLM-AutoDPComparison Baseline=All-Process, Evaluated Model=Qwen2.5-7B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 0.5127 | 0.0963 | 0.391 | |
| LLM-AutoDPComparison Baseline=All-Process, Evaluated Model=Gemma-2-9B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 0.5066 | 0.1003 | 0.3931 | |
| LLM-AutoDPComparison Baseline=RS, Evaluated Model=Gemma-2-9B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 0.4842 | 0.1178 | 0.398 | |
| LLM-AutoDPComparison Baseline=RS, Evaluated Model=Llama3.1-8B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 0.4793 | 0.0811 | 0.4396 | |
| LLM-AutoDPComparison Baseline=All-Process, Evaluated Model=Llama3.1-8B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 0.4769 | 0.0873 | 0.4358 | |
| LLM-AutoDPComparison Baseline=RS, Evaluated Model=Qwen2.5-7B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 0.4688 | 0.1017 | 0.4295 | |
| LLM-AutoDPComparison Baseline=SELA, Evaluated Model=Gemma-2-9B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 0.4688 | 0.0889 | 0.4423 | |
| LLM-AutoDPComparison Baseline=SELA, Evaluated Model=Qwen2.5-7B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 0.4421 | 0.0889 | 0.469 | |
| LLM-AutoDPComparison Baseline=SELA, Evaluated Model=Llama3.1-8B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 0.433 | 0.0108 | 0.5562 | |
| LLM-AutoDPComparison Baseline=No-Process, Evaluated Model=Qwen2.5-7B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 0 | 1 | 0 | |
| LLM-AutoDPComparison Baseline=No-Process, Evaluated Model=Llama3.1-8B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 0 | 1 | 0 | |
| LLM-AutoDPComparison Baseline=No-Process, Evaluated Model=Gemma-2-9B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 0 | 1 | 0 |