Medical Question Answering on Huatuo-26M-Lite (test)
59.31Win RateLLM-AutoDP
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| LLM-AutoDPComparison Baseline=No-Process, Evaluated Model=Gemma-2-9B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 59.31 | 8.42 | 32.27 | |
| LLM-AutoDPComparison Baseline=No-Process, Evaluated Model=Llama3.1-8B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 56.37 | 9.25 | 34.37 | |
| LLM-AutoDPComparison Baseline=No-Process, Evaluated Model=Qwen2.5-7B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 54.01 | 12.87 | 33.12 | |
| LLM-AutoDPComparison Baseline=RS, Evaluated Model=Gemma-2-9B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 48.97 | 18.18 | 32.85 | |
| LLM-AutoDPComparison Baseline=All-Process, Evaluated Model=Gemma-2-9B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 45.43 | 17.63 | 36.94 | |
| LLM-AutoDPComparison Baseline=SELA, Evaluated Model=Gemma-2-9B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 45.43 | 17.63 | 36.94 | |
| LLM-AutoDPComparison Baseline=RS, Evaluated Model=Qwen2.5-7B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 45.25 | 20.25 | 34.5 | |
| LLM-AutoDPComparison Baseline=RS, Evaluated Model=Llama3.1-8B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 44.44 | 18.33 | 37.23 | |
| LLM-AutoDPComparison Baseline=All-Process, Evaluated Model=Llama3.1-8B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 42.69 | 17.17 | 40.14 | |
| LLM-AutoDPComparison Baseline=SELA, Evaluated Model=Llama3.1-8B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 42.69 | 17.17 | 40.14 | |
| LLM-AutoDPComparison Baseline=All-Process, Evaluated Model=Qwen2.5-7B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 39.13 | 19.63 | 41.24 | |
| LLM-AutoDPComparison Baseline=SELA, Evaluated Model=Qwen2.5-7B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 39.13 | 19.63 | 41.24 |