Automated Red-Teaming of LLM-Based Search Agents on SAFESEARCH
99.7Harm Score (Benign)GPT-5-mini
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| GPT-5-miniAgent Configuration=Deep research scaffold, Reasoning model=true2025.09 | 99.7 | 99.8 | 11.1 | 16.7 | 13.9 | 3.3 | 35.6 | 16.1 | |
| DeepSeek-R1Agent Configuration=Deep research scaffold, Reasoning model=true2025.09 | 99.6 | 99.5 | 47.2 | 28.6 | 31.2 | 8.9 | 37.4 | 30.6 | |
| GPT-4.1-miniAgent Configuration=Deep research scaffold, Reasoning model=true2025.09 | 99.3 | 96.3 | 76.1 | 63.9 | 60.6 | 9.4 | 77.2 | 57.4 | |
| GPT-5-miniAgent Configuration=LLM w/ tool calling, Reasoning model=true2025.09 | 99.2 | 99.4 | 0 | 2.8 | 10.6 | 1.1 | 10.1 | 4.9 | |
| GPT-4.1Agent Configuration=LLM w/ tool calling, Reasoning model=true2025.09 | 99 | 94.1 | 78.3 | 76.1 | 70 | 69.4 | 92.8 | 77.3 | |
| GPT-4.1Agent Configuration=LLM w/ search workflow, Reasoning model=true2025.09 | 98.9 | 92.9 | 92.8 | 88.9 | 69.4 | 76.7 | 97.2 | 85 | |
| Qwen3-235B-A22B-2507Agent Configuration=LLM w/ tool calling, Reasoning model=false2025.09 | 98.8 | 99.4 | 47.8 | 18.9 | 21.1 | 63.3 | 29.4 | 36.1 | |
| Qwen3-32BAgent Configuration=Deep research scaffold, Reasoning model=true2025.09 | 98.6 | 98.8 | 33.5 | 55.6 | — | 16.3 | 70.9 | 44.7 | |
| GPT-5Agent Configuration=LLM w/ tool calling, Reasoning model=true2025.09 | 98.5 | 99.4 | 1.1 | 2.2 | 13.3 | 0 | 8.3 | 5 | |
| GPT-5-miniAgent Configuration=LLM w/ search workflow, Reasoning model=true2025.09 | 98.4 | 99.4 | 10 | 19.4 | 15 | 6.1 | 43.9 | 18.9 | |
| Qwen3-8BAgent Configuration=Deep research scaffold, Reasoning model=false2025.09 | 98.3 | 98.2 | 43.3 | 55.6 | 50.6 | 14.6 | 64.4 | 45.8 | |
| DeepSeek-R1Agent Configuration=LLM w/ search workflow, Reasoning model=true2025.09 | 98.2 | 97.8 | 78.7 | 54.4 | 50.3 | 75.6 | 75 | 66.8 | |
| o4-miniAgent Configuration=LLM w/ tool calling, Reasoning model=true2025.09 | 98 | 96.6 | 41.7 | 57.8 | 37.8 | 16.1 | 65.6 | 43.8 | |
| DeepSeek-R1Agent Configuration=LLM w/ tool calling, Reasoning model=true2025.09 | 97.8 | 97.4 | 62.6 | 62.2 | 53.7 | 69.3 | 76.1 | 64.8 | |
| GPT-5Agent Configuration=LLM w/ search workflow, Reasoning model=true2025.09 | 97.6 | 99.2 | 13.9 | 16.7 | 21.1 | 5 | 35.6 | 18.4 | |
| Claude-Sonnet-4.5Agent Configuration=LLM w/ tool calling, Reasoning model=true2025.09 | 97.4 | 95.6 | 2.8 | 2.4 | 4.5 | 2.8 | 10.6 | 4.6 | |
| Qwen3-235B-A22BAgent Configuration=LLM w/ search workflow, Reasoning model=false2025.09 | 97.2 | 93.5 | 78.3 | 75 | 68.3 | 83.3 | 88.3 | 78.7 | |
| Qwen3-32BAgent Configuration=LLM w/ search workflow, Reasoning model=true2025.09 | 97.1 | 92.9 | 78.7 | 80.6 | 73.9 | 79.4 | 92.7 | 81.1 | |
| Qwen3-235B-A22BAgent Configuration=LLM w/ tool calling, Reasoning model=false2025.09 | 96.3 | 95.1 | 55 | 63.9 | 61.7 | 76.5 | 85 | 68.4 | |
| Claude-Haiku-4.5Agent Configuration=LLM w/ tool calling, Reasoning model=false2025.09 | 96.1 | 97.7 | 5 | 6.7 | 11.1 | 2.2 | 30 | 11 | |
| GPT-4.1-miniAgent Configuration=LLM w/ search workflow, Reasoning model=true2025.09 | 95.7 | 79.3 | 92.8 | 86.1 | 87.9 | 88.9 | 96.7 | 90.5 | |
| Kimi-K2Agent Configuration=LLM w/ tool calling, Reasoning model=false2025.09 | 95.7 | 91.1 | 43.3 | 42.2 | 42.8 | 47.2 | 60.6 | 47.2 | |
| o4-miniAgent Configuration=LLM w/ search workflow, Reasoning model=true2025.09 | 95.4 | 90.3 | 64.4 | 78.9 | 56.1 | 25.6 | 76.1 | 60.2 | |
| GPT-oss-120bAgent Configuration=LLM w/ search workflow, Reasoning model=false2025.09 | 95.3 | 96.4 | 80.3 | 86.4 | 65.4 | 57.7 | 85 | 75.6 | |
| Qwen3-32BAgent Configuration=LLM w/ tool calling, Reasoning model=true2025.09 | 95.2 | 94.2 | 58.9 | 63.3 | 61.7 | 63.3 | 81.7 | 65.8 | |
| GPT-4.1-miniAgent Configuration=LLM w/ tool calling, Reasoning model=true2025.09 | 94.8 | 86.9 | 70 | 73.3 | 75.6 | 79.4 | 90.6 | 77.8 | |
| Qwen3-8BAgent Configuration=LLM w/ search workflow, Reasoning model=false2025.09 | 94.3 | 90.6 | 76.9 | 76.2 | 86.1 | 91.1 | 93.9 | 85.5 | |
| Qwen3-235B-A22B-2507Agent Configuration=LLM w/ search workflow, Reasoning model=false2025.09 | 94.3 | 98.1 | 60.4 | 20.8 | 43.4 | 63.3 | 23.8 | 43.4 | |
| Claude-Sonnet-4.5Agent Configuration=LLM w/ search workflow, Reasoning model=true2025.09 | 94.2 | 94.9 | 35 | 9.6 | 17.1 | 18.6 | 18.1 | 19.8 | |
| Kimi-K2Agent Configuration=LLM w/ search workflow, Reasoning model=false2025.09 | 94 | 87.1 | 77.1 | 83.5 | 73.9 | 84.1 | 91 | 81.9 | |
| Qwen3-8BAgent Configuration=LLM w/ tool calling, Reasoning model=false2025.09 | 92.7 | 92 | 58.9 | 64.4 | 68.3 | 81.1 | 81.1 | 70.8 | |
| Gemini-2.5-ProAgent Configuration=LLM w/ tool calling, Reasoning model=true2025.09 | 86 | 88 | 59.8 | 66.1 | 61.7 | 21.1 | 83.8 | 58.5 | |
| Gemma-3-IT-27BAgent Configuration=LLM w/ search workflow, Reasoning model=false2025.09 | 85.9 | 87.6 | 87 | 89 | 84.1 | 80.3 | 92.7 | 86.6 | |
| Claude-Haiku-4.5Agent Configuration=LLM w/ search workflow, Reasoning model=false2025.09 | 85.8 | 94.7 | 35 | 25.6 | 21.7 | 8.9 | 47.2 | 27.7 | |
| Gemini-2.5-FlashAgent Configuration=LLM w/ tool calling, Reasoning model=false2025.09 | 76.1 | 77.6 | 60 | 74.6 | 69.4 | 68.7 | 83.9 | 71.3 | |
| Gemma-3-IT-27BAgent Configuration=LLM w/ tool calling, Reasoning model=false2025.09 | 72.4 | 74.5 | 62.8 | 69.3 | 79.4 | 63.9 | 90.6 | 73.2 | |
| Gemini-2.5-FlashAgent Configuration=LLM w/ search workflow, Reasoning model=false2025.09 | 71.6 | 76.2 | 70 | 85.8 | 81.7 | 58.3 | 89.7 | 76.8 | |
| Gemini-2.5-ProAgent Configuration=LLM w/ search workflow, Reasoning model=true2025.09 | 67.6 | 82.8 | 79.4 | 81.5 | 85 | 36.7 | 92.8 | 75.1 | |
| GPT-oss-120bAgent Configuration=LLM w/ tool calling, Reasoning model=false2025.09 | 60.4 | 88 | 59 | 64 | 56.1 | 47.7 | 63.5 | 58.1 |