Adversarial Tutor Leakage on HumanEval (test)
17.3Leak RateManually Defined Prompts
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Manually Defined PromptsAttack Type=Manually Defined Prompts, Tutor Type=Tutor w/ Reasoning, Tutor Model=Qwen-7B2026.04 | 17.3 | 6.347 | |
| Multi-Agent StudentAttack Type=Multi-Agent Student, Tutor Type=Tutor w/ Reasoning, Tutor Model=Qwen-7B2026.04 | 18.9 | 10.806 | |
| Base Student Adv. AgentAttack Type=Base Student Adv. Agent, Tutor Type=Tutor w/ Reasoning, Tutor Model=Qwen-7B2026.04 | 24.4 | 12.175 | |
| Student w/ ReasoningAttack Type=Student w/ Reasoning, Tutor Type=Tutor w/ Reasoning, Tutor Model=Qwen-7B2026.04 | 24.4 | 10.6 | |
| Finetuned Adv. AgentAttack Type=Finetuned Adv. Agent, Tutor Type=Tutor w/ Reasoning, Tutor Model=Qwen-7B2026.04 | 40.9 | 6.866 | |
| Manually Defined PromptsAttack Type=Manually Defined Prompts, Tutor Type=Base In-context Tutor, Tutor Model=Qwen-7B2026.04 | 74.6 | 3.85 | |
| Base Student Adv. AgentAttack Type=Base Student Adv. Agent, Tutor Type=Base In-context Tutor, Tutor Model=Qwen-7B2026.04 | 76.2 | 4.952 | |
| Student w/ ReasoningAttack Type=Student w/ Reasoning, Tutor Type=Base In-context Tutor, Tutor Model=Qwen-7B2026.04 | 78 | 4.039 | |
| Multi-Agent StudentAttack Type=Multi-Agent Student, Tutor Type=Base In-context Tutor, Tutor Model=Qwen-7B2026.04 | 78 | 4.516 | |
| Finetuned Adv. AgentAttack Type=Finetuned Adv. Agent, Tutor Type=Base In-context Tutor, Tutor Model=Qwen-7B2026.04 | 87.8 | 3.062 |