Persona Alignment on Philosophical Persona Evaluation Set
32.24Mean ScoreFew-Shot
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Few-ShotBase Model=Qwen-3-4B2026.05 | 32.24 | 31.81 | — | |
| AlphaPOBase Model=Qwen-3-4B, Adaptation Examples=3002026.05 | 30.92 | 30.48 | 0.001 | |
| AlphaPOBase Model=Qwen-3-4B, Adaptation Examples=2002026.05 | 29.3 | 28.88 | 0.001 | |
| ORPOBase Model=Qwen-3-4B, Adaptation Examples=3002026.05 | 28.98 | 28.47 | 0.001 | |
| AlphaPOBase Model=Qwen-3-4B, Adaptation Examples=1002026.05 | 28.54 | 28.1 | 0.001 | |
| ORPOBase Model=Qwen-3-4B, Adaptation Examples=1002026.05 | 28.46 | 28.07 | 0.001 | |
| ORPOBase Model=Qwen-3-4B, Adaptation Examples=2002026.05 | 28.37 | 27.91 | 0.001 | |
| Zero-ShotBase Model=Qwen-3-4B2026.05 | 27.79 | 27.25 | 0.001 | |
| Few-ShotBase Model=Llama-3.2-3B2026.05 | 26.11 | 25.41 | 0.001 | |
| ORPOBase Model=Llama-3.2-3B, Adaptation Examples=3002026.05 | 25.13 | 24.49 | 0.001 | |
| AlphaPOBase Model=Llama-3.2-3B, Adaptation Examples=3002026.05 | 21.56 | 20.68 | 0.001 | |
| ORPOBase Model=Llama-3.2-3B, Adaptation Examples=2002026.05 | 21.49 | 20.69 | 0.001 | |
| AlphaPOBase Model=Llama-3.2-3B, Adaptation Examples=2002026.05 | 19.82 | 18.97 | 0.001 | |
| ORPOBase Model=Llama-3.2-3B, Adaptation Examples=1002026.05 | 18.66 | 17.78 | 0.001 | |
| AlphaPOBase Model=Llama-3.2-3B, Adaptation Examples=1002026.05 | 18.55 | 17.76 | 0.001 | |
| Zero-ShotBase Model=Llama-3.2-3B2026.05 | 17.65 | 16.83 | 0.001 |