Persuasion classification on Winning Arguments (test)
82.69AccuracyBlack-Box (Gemma2)
Evaluation Results
| Method | Links | |
|---|---|---|
| Black-Box (Gemma2)approach=Hybrid Interaction, classifier=Random Forest, features=LLM-generated features with pairwise interaction scores2025.11 | 82.69 | |
| Black-Box (Gemma2)approach=Hybrid Independent, classifier=Random Forest, features=LLM-generated features with independent belief update scores2025.11 | 82.32 | |
| Black-Box (LLaMA3)approach=Hybrid Independent, classifier=Random Forest, features=LLM-generated features with independent belief update scores2025.11 | 73.17 | |
| Black-Box (Mixtral)approach=Hybrid Interaction, classifier=Random Forest, features=LLM-generated features with pairwise interaction scores2025.11 | 73.17 | |
| Black-Box (LLaMA3)approach=Hybrid Interaction, classifier=Random Forest, features=LLM-generated features with pairwise interaction scores2025.11 | 73.11 | |
| Black-Box (Mixtral)approach=Hybrid Independent, classifier=Random Forest, features=LLM-generated features with independent belief update scores2025.11 | 72.55 | |
| Black-Box (LLaMA3)approach=Data-Driven, mode=zero-shot classification2025.11 | 64.93 | |
| MS-PS-MLPModel=OpenAI-o32026.01 | 64.53 | |
| Black-Box (Gemma2)approach=Theory-Driven, features=thresholded belief update scores2025.11 | 63.8 | |
| MS-PS-MLPModel=Gemma-3-12B2026.01 | 63.69 | |
| MS-PS-MLPModel=Gemini-1.52026.01 | 63.07 | |
| MS-PS-AVGModel=Gemma-3-12B2026.01 | 62.83 | |
| MS-PS-MLPModel=Gemini-22026.01 | 62.7 | |
| + ContextModel=Gemini-1.52026.01 | 61.96 | |
| MS-PS-AVGModel=Gemini-22026.01 | 61.83 | |
| Black-Box (Mixtral)approach=Data-Driven, mode=zero-shot classification2025.11 | 61.71 | |
| + Context + ExplanationModel=Gemini-22026.01 | 61.46 | |
| MS-PS-MLPModel=Llama-3.1-8B2026.01 | 61.34 | |
| + Context + ExplanationModel=Gemini-1.52026.01 | 61.09 | |
| + ContextModel=Gemini-22026.01 | 60.84 | |
| + ExplanationModel=Gemma-3-12B2026.01 | 60.72 | |
| MS-PS-AVGModel=Llama-3.1-8B2026.01 | 60.72 | |
| MS-PS-AVGModel=Gemini-1.52026.01 | 60.72 | |
| MS-PS-AVGModel=OpenAI-o32026.01 | 60.59 | |
| + Context + ExplanationModel=OpenAI-o32026.01 | 60.35 | |
| + ContextModel=Llama-3.1-8B2026.01 | 59.48 | |
| + ContextModel=Gemma-3-12B2026.01 | 59.48 | |
| + ExplanationModel=Gemini-22026.01 | 59.11 | |
| + ContextModel=OpenAI-o32026.01 | 58.98 | |
| + ExplanationModel=OpenAI-o32026.01 | 58.24 | |
| + Context + ExplanationModel=Gemma-3-12B2026.01 | 58.24 | |
| Black-Box (Mixtral)approach=Theory-Driven, features=thresholded belief update scores2025.11 | 58.2 | |
| Black-Box (Gemma2)approach=Data-Driven, mode=zero-shot classification2025.11 | 56.9 | |
| Independent ScoringModel=Llama-3.1-8B2026.01 | 56.66 | |
| Transparent (Logistic regression)approach=Data-Driven, features=term frequencies2025.11 | 56.5 | |
| + ExplanationModel=Gemini-1.52026.01 | 56.38 | |
| + ExplanationModel=Llama-3.1-8B2026.01 | 56.26 | |
| Independent ScoringModel=Gemini-22026.01 | 56.13 | |
| Independent ScoringModel=Gemini-1.52026.01 | 56.01 | |
| Independent ScoringModel=OpenAI-o32026.01 | 55.51 | |
| + Context + ExplanationModel=Llama-3.1-8B2026.01 | 54.52 | |
| Black-Box (LLaMA3)approach=Theory-Driven, features=thresholded belief update scores2025.11 | 54.3 | |
| Independent ScoringModel=Gemma-3-12B2026.01 | 53.78 |