Claim Verification on Truthful Claim
81AccuracyArgLLM
Evaluation Results
| Method | Links | |
|---|---|---|
| ArgLLMBackbone=Mixtral, Base Score Type=Estimated, Argument Depth (D)=1, Gradual Semantics (sigma)=DF-QuAD2024.05 | 81 | |
| ArgLLMBackbone=GPT-4o mini, Base Score Type=Estimated, Argument Depth (D)=1, Gradual Semantics (sigma)=DF-QuAD2024.05 | 81 | |
| ArgLLMBackbone=GPT-4o mini, Base Score Type=Estimated, Argument Depth (D)=2, Gradual Semantics (sigma)=DF-QuAD2024.05 | 81 | |
| Est. ConfidenceBackbone=GPT-4o mini, Prompting Method=Est. Confidence2024.05 | 79 | |
| Chain-of-ThoughtBackbone=GPT-4o mini, Prompting Method=Chain-of-Thought2024.05 | 79 | |
| ArgLLMBackbone=Mixtral, Base Score Type=Estimated, Argument Depth (D)=2, Gradual Semantics (sigma)=DF-QuAD2024.05 | 78 | |
| Direct QuestionBackbone=Gemma 2 9B, Prompting Method=Direct Question2024.05 | 78 | |
| Chain-of-ThoughtBackbone=Gemma 2 9B, Prompting Method=Chain-of-Thought2024.05 | 78 | |
| Direct QuestionBackbone=GPT-4o mini, Prompting Method=Direct Question2024.05 | 78 | |
| ArgLLMBackbone=GPT-4o mini, Base Score Type=0.5, Argument Depth (D)=2, Gradual Semantics (sigma)=DF-QuAD2024.05 | 78 | |
| Direct QuestionBackbone=Mixtral, Prompting Method=Direct Question2024.05 | 77 | |
| Est. ConfidenceBackbone=Mixtral, Prompting Method=Est. Confidence2024.05 | 77 | |
| ArgLLMBackbone=Mistral, Base Score Type=Estimated, Argument Depth (D)=1, Gradual Semantics (sigma)=DF-QuAD2024.05 | 76 | |
| Chain-of-ThoughtBackbone=Mixtral, Prompting Method=Chain-of-Thought2024.05 | 76 | |
| Chain-of-ThoughtBackbone=Mistral, Prompting Method=Chain-of-Thought2024.05 | 75 | |
| ArgLLMBackbone=Mistral, Base Score Type=Estimated, Argument Depth (D)=2, Gradual Semantics (sigma)=DF-QuAD2024.05 | 75 | |
| Est. ConfidenceBackbone=Gemma 2 9B, Prompting Method=Est. Confidence2024.05 | 74 | |
| Chain-of-ThoughtBackbone=GPT-3.5-turbo, Prompting Method=Chain-of-Thought2024.05 | 74 | |
| ArgLLMBackbone=GPT-4o mini, Base Score Type=0.5, Argument Depth (D)=1, Gradual Semantics (sigma)=DF-QuAD2024.05 | 74 | |
| Direct QuestionBackbone=Mistral, Prompting Method=Direct Question2024.05 | 73 | |
| Est. ConfidenceBackbone=Mistral, Prompting Method=Est. Confidence2024.05 | 73 | |
| ArgLLMBackbone=Gemma 2 9B, Base Score Type=Estimated, Argument Depth (D)=1, Gradual Semantics (sigma)=DF-QuAD2024.05 | 73 | |
| ArgLLMBackbone=Gemma 2 9B, Base Score Type=Estimated, Argument Depth (D)=2, Gradual Semantics (sigma)=DF-QuAD2024.05 | 73 | |
| Est. ConfidenceBackbone=GPT-3.5-turbo, Prompting Method=Est. Confidence2024.05 | 73 | |
| ArgLLMBackbone=GPT-3.5-turbo, Base Score Type=Estimated, Argument Depth (D)=1, Gradual Semantics (sigma)=DF-QuAD2024.05 | 73 | |
| ArgLLMBackbone=Mixtral, Base Score Type=0.5, Argument Depth (D)=1, Gradual Semantics (sigma)=DF-QuAD2024.05 | 72 | |
| ArgLLMBackbone=Mixtral, Base Score Type=0.5, Argument Depth (D)=2, Gradual Semantics (sigma)=DF-QuAD2024.05 | 72 | |
| ArgLLMBackbone=GPT-3.5-turbo, Base Score Type=Estimated, Argument Depth (D)=2, Gradual Semantics (sigma)=DF-QuAD2024.05 | 72 | |
| Chain-of-ThoughtBackbone=Llama 3 8B, Prompting Method=Chain-of-Thought2024.05 | 70 | |
| Direct QuestionBackbone=GPT-3.5-turbo, Prompting Method=Direct Question2024.05 | 70 | |
| Est. ConfidenceBackbone=Llama 3 8B, Prompting Method=Est. Confidence2024.05 | 69 | |
| ArgLLMBackbone=Llama 3 8B, Base Score Type=Estimated, Argument Depth (D)=2, Gradual Semantics (sigma)=DF-QuAD2024.05 | 69 | |
| Chain-of-ThoughtBackbone=Gemma 7B, Prompting Method=Chain-of-Thought2024.05 | 68 | |
| ArgLLMBackbone=Gemma 2 9B, Base Score Type=0.5, Argument Depth (D)=1, Gradual Semantics (sigma)=DF-QuAD2024.05 | 68 | |
| ArgLLMBackbone=Gemma 2 9B, Base Score Type=0.5, Argument Depth (D)=2, Gradual Semantics (sigma)=DF-QuAD2024.05 | 68 | |
| ArgLLMBackbone=Llama 3 8B, Base Score Type=Estimated, Argument Depth (D)=1, Gradual Semantics (sigma)=DF-QuAD2024.05 | 68 | |
| ArgLLMBackbone=Mistral, Base Score Type=0.5, Argument Depth (D)=2, Gradual Semantics (sigma)=DF-QuAD2024.05 | 67 | |
| Direct QuestionBackbone=Llama 3 8B, Prompting Method=Direct Question2024.05 | 66 | |
| ArgLLMBackbone=Mistral, Base Score Type=0.5, Argument Depth (D)=1, Gradual Semantics (sigma)=DF-QuAD2024.05 | 65 | |
| Direct QuestionBackbone=Gemma 7B, Prompting Method=Direct Question2024.05 | 65 | |
| ArgLLMBackbone=Gemma 7B, Base Score Type=0.5, Argument Depth (D)=1, Gradual Semantics (sigma)=DF-QuAD2024.05 | 64 | |
| ArgLLMBackbone=GPT-3.5-turbo, Base Score Type=0.5, Argument Depth (D)=2, Gradual Semantics (sigma)=DF-QuAD2024.05 | 64 | |
| ArgLLMBackbone=Gemma 7B, Base Score Type=Estimated, Argument Depth (D)=1, Gradual Semantics (sigma)=DF-QuAD2024.05 | 63 | |
| ArgLLMBackbone=Llama 3 8B, Base Score Type=0.5, Argument Depth (D)=1, Gradual Semantics (sigma)=DF-QuAD2024.05 | 63 | |
| Est. ConfidenceBackbone=Gemma 7B, Prompting Method=Est. Confidence2024.05 | 62 | |
| ArgLLMBackbone=Gemma 7B, Base Score Type=0.5, Argument Depth (D)=2, Gradual Semantics (sigma)=DF-QuAD2024.05 | 62 | |
| ArgLLMBackbone=Gemma 7B, Base Score Type=Estimated, Argument Depth (D)=2, Gradual Semantics (sigma)=DF-QuAD2024.05 | 62 | |
| ArgLLMBackbone=Llama 3 8B, Base Score Type=0.5, Argument Depth (D)=2, Gradual Semantics (sigma)=DF-QuAD2024.05 | 61 | |
| ArgLLMBackbone=GPT-3.5-turbo, Base Score Type=0.5, Argument Depth (D)=1, Gradual Semantics (sigma)=DF-QuAD2024.05 | 60 |