Role-playing evaluation on RoleBench
85.67LLM-as-a-Judge ScorePersona-Pruner
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Persona-PrunerBackbone=Qwen2.5-7B-Instruct, Ratio=25%, Recovery finetuning=true2026.06 | 85.67 | — | |
| Persona-PrunerBackbone=Qwen2.5-3B-Instruct, Ratio=25%, Recovery finetuning=true2026.06 | 84.58 | — | |
| Persona-PrunerBackbone=Qwen2.5-7B-Instruct, Ratio=25%, Recovery finetuning=false2026.06 | 84.33 | — | |
| Persona-PrunerBackbone=Qwen2.5-7B-Instruct, Ratio=50%, Recovery finetuning=true2026.06 | 83.24 | — | |
| Persona-PrunerBackbone=Qwen2.5-3B-Instruct, Ratio=25%, Recovery finetuning=false2026.06 | 80.89 | — | |
| Persona-PrunerBackbone=Qwen2.5-7B-Instruct, Ratio=50%, Recovery finetuning=false2026.06 | 77.89 | — | |
| Persona-PrunerBackbone=Qwen2.5-3B-Instruct, Ratio=50%, Recovery finetuning=true2026.06 | 77.56 | — | |
| Qwen2.5-7B-InstructBackbone=Qwen2.5-7B-Instruct, Ratio=0%, Recovery finetuning=false2026.06 | 68.79 | — | |
| Qwen2.5-7B-InstructBackbone=Qwen2.5-7B-Instruct, Ratio=0%, Recovery finetuning=true2026.06 | 68.79 | — | |
| Persona-PrunerBackbone=Qwen2.5-3B-Instruct, Ratio=50%, Recovery finetuning=false2026.06 | 68.07 | — | |
| LLM-PrunerBackbone=Qwen2.5-7B-Instruct, Ratio=25%, Recovery finetuning=false2026.06 | 67.12 | — | |
| Adapt-PrunerBackbone=Qwen2.5-7B-Instruct, Ratio=25%, Recovery finetuning=false2026.06 | 48.51 | — | |
| LLM-PrunerBackbone=Qwen2.5-7B-Instruct, Ratio=50%, Recovery finetuning=false2026.06 | 46.7 | — | |
| Qwen2.5-3B-InstructBackbone=Qwen2.5-3B-Instruct, Ratio=0%, Recovery finetuning=false2026.06 | 45.51 | — | |
| Qwen2.5-3B-InstructBackbone=Qwen2.5-3B-Instruct, Ratio=0%, Recovery finetuning=true2026.06 | 45.51 | — | |
| Adapt-PrunerBackbone=Qwen2.5-3B-Instruct, Ratio=25%, Recovery finetuning=false2026.06 | 35.68 | — | |
| LLM-PrunerBackbone=Qwen2.5-7B-Instruct, Ratio=25%, Recovery finetuning=true2026.06 | 31.51 | — | |
| Adapt-PrunerBackbone=Qwen2.5-7B-Instruct, Ratio=50%, Recovery finetuning=false2026.06 | 27.85 | — | |
| Adapt-PrunerBackbone=Qwen2.5-7B-Instruct, Ratio=25%, Recovery finetuning=true2026.06 | 27.42 | — | |
| LLM-PrunerBackbone=Qwen2.5-3B-Instruct, Ratio=25%, Recovery finetuning=false2026.06 | 27.19 | — | |
| Depth PruningBackbone=Qwen2.5-7B-Instruct, Ratio=25%, Recovery finetuning=false2026.06 | 26.79 | — | |
| LLM-PrunerBackbone=Qwen2.5-7B-Instruct, Ratio=50%, Recovery finetuning=true2026.06 | 25.52 | — | |
| Adapt-PrunerBackbone=Qwen2.5-3B-Instruct, Ratio=25%, Recovery finetuning=true2026.06 | 24.15 | — | |
| Depth PruningBackbone=Qwen2.5-7B-Instruct, Ratio=25%, Recovery finetuning=true2026.06 | 23.25 | — | |
| Adapt-PrunerBackbone=Qwen2.5-7B-Instruct, Ratio=50%, Recovery finetuning=true2026.06 | 22.84 | — | |
| LLM-PrunerBackbone=Qwen2.5-3B-Instruct, Ratio=25%, Recovery finetuning=true2026.06 | 22.19 | — | |
| Depth PruningBackbone=Qwen2.5-3B-Instruct, Ratio=25%, Recovery finetuning=true2026.06 | 22.05 | — | |
| Adapt-PrunerBackbone=Qwen2.5-3B-Instruct, Ratio=50%, Recovery finetuning=true2026.06 | 20.86 | — | |
| SliceGPTBackbone=Qwen2.5-7B-Instruct, Ratio=25%, Recovery finetuning=true2026.06 | 20.25 | — | |
| Depth PruningBackbone=Qwen2.5-3B-Instruct, Ratio=50%, Recovery finetuning=true2026.06 | 17.3 | — | |
| LLM-PrunerBackbone=Qwen2.5-3B-Instruct, Ratio=50%, Recovery finetuning=true2026.06 | 16.12 | — | |
| Adapt-PrunerBackbone=Qwen2.5-3B-Instruct, Ratio=50%, Recovery finetuning=false2026.06 | 15.87 | — | |
| Depth PruningBackbone=Qwen2.5-3B-Instruct, Ratio=25%, Recovery finetuning=false2026.06 | 14.58 | — | |
| LLM-PrunerBackbone=Qwen2.5-3B-Instruct, Ratio=50%, Recovery finetuning=false2026.06 | 12.67 | — | |
| SliceGPTBackbone=Qwen2.5-3B-Instruct, Ratio=25%, Recovery finetuning=true2026.06 | 11.97 | — | |
| SliceGPTBackbone=Qwen2.5-7B-Instruct, Ratio=25%, Recovery finetuning=false2026.06 | 11.66 | — | |
| SliceGPTBackbone=Qwen2.5-7B-Instruct, Ratio=50%, Recovery finetuning=true2026.06 | 11.16 | — | |
| Depth PruningBackbone=Qwen2.5-7B-Instruct, Ratio=50%, Recovery finetuning=true2026.06 | 8.73 | — | |
| SliceGPTBackbone=Qwen2.5-3B-Instruct, Ratio=50%, Recovery finetuning=true2026.06 | 8.16 | — | |
| Depth PruningBackbone=Qwen2.5-3B-Instruct, Ratio=50%, Recovery finetuning=false2026.06 | 3.84 | — | |
| Depth PruningBackbone=Qwen2.5-7B-Instruct, Ratio=50%, Recovery finetuning=false2026.06 | 3.49 | — | |
| SliceGPTBackbone=Qwen2.5-3B-Instruct, Ratio=25%, Recovery finetuning=false2026.06 | 0.79 | — | |
| SliceGPTBackbone=Qwen2.5-3B-Instruct, Ratio=50%, Recovery finetuning=false2026.06 | 0.39 | — | |
| SliceGPTBackbone=Qwen2.5-7B-Instruct, Ratio=50%, Recovery finetuning=false2026.06 | 0.18 | — | |
| DPO-Qwen3-8BStage=DPO2026.05 | — | 37.1 | |
| GPT-4.12026.05 | — | 34.3 | |
| Qwen3-8BStage=Base2026.05 | — | 0 | |
| SFT-Qwen3-8BStage=SFT2026.05 | — | 28.6 |