Real-life Chat LLM Evaluation on AlpacaEval 203 random samples 2
19.13LC Win RateZero-shot + MG
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Zero-shot + MGEvaluation protocol=Zero-shot, Guideline configuration=MG2025.06 | 19.13 | 7.07 | |
| LongGuideEvaluation protocol=Zero-shot, Guideline configuration=LongGuide2025.06 | 19.13 | 7.07 | |
| Few-shot + MGEvaluation protocol=Few-shot, Guideline configuration=MG2025.06 | 12.65 | 4.88 | |
| LongGuideEvaluation protocol=Few-shot, Guideline configuration=LongGuide2025.06 | 12.65 | 4.88 | |
| Few-shot + MG-OCGEvaluation protocol=Few-shot, Guideline configuration=MG-OCG2025.06 | 12.63 | 4.88 | |
| Zero-shot (ZS)Evaluation protocol=Zero-shot2025.06 | 11.08 | 3.17 | |
| Zero-shot + MG-OCGEvaluation protocol=Zero-shot, Guideline configuration=MG-OCG2025.06 | 8.42 | 3.9 | |
| Few-shot (FS)Evaluation protocol=Few-shot2025.06 | 8.08 | 2.68 | |
| Few-shot + OCGEvaluation protocol=Few-shot, Guideline configuration=OCG2025.06 | 7.73 | 3.45 | |
| Zero-shot + OCGEvaluation protocol=Zero-shot, Guideline configuration=OCG2025.06 | 4.73 | 2.44 |