Medical Question Answering on Huatuo-26M-Lite-100 (test)
93WinsLLM-AutoDP
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| LLM-AutoDPComparison Baseline=No-Process, Evaluated Model=Gemma-2-9B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 93 | 2 | 5 | |
| LLM-AutoDPComparison Baseline=No-Process, Evaluated Model=Llama3.1-8B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 91.37 | 4.25 | 4.37 | |
| LLM-AutoDPComparison Baseline=No-Process, Evaluated Model=Qwen2.5-7B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 64 | 25.37 | 10.625 | |
| LLM-AutoDPComparison Baseline=RS, Evaluated Model=Gemma-2-9B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 58.2 | 20.27 | 21.53 | |
| LLM-AutoDPComparison Baseline=All-Process, Evaluated Model=Gemma-2-9B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 57.32 | 11.68 | 31 | |
| LLM-AutoDPComparison Baseline=SELA, Evaluated Model=Gemma-2-9B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 56.9 | 18 | 25.1 | |
| LLM-AutoDPComparison Baseline=RS, Evaluated Model=Llama3.1-8B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 56.5 | 19.37 | 24.13 | |
| LLM-AutoDPComparison Baseline=All-Process, Evaluated Model=Llama3.1-8B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 55.25 | 12.38 | 32.37 | |
| LLM-AutoDPComparison Baseline=SELA, Evaluated Model=Llama3.1-8B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 54.75 | 12.75 | 32.5 | |
| LLM-AutoDPComparison Baseline=All-Process, Evaluated Model=Qwen2.5-7B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 52.13 | 19 | 28.88 | |
| LLM-AutoDPComparison Baseline=SELA, Evaluated Model=Qwen2.5-7B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 50.5 | 26.62 | 22.88 | |
| LLM-AutoDPComparison Baseline=RS, Evaluated Model=Qwen2.5-7B, Judge Model=Baichuan-M1-14B-Instruct2026.01 | 50.38 | 29.12 | 20.5 |