Paper Assessment Reasoning Extraction on STRICTA 1.0 (test)
87.6F1 ScoreGPT4o
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| GPT4ointeraction_condition=human oversight, graph_usage=true2024.09 | 87.6 | 79.5 | -0.094 | 19.4 | |
| Majority Baselineinteraction_condition=human oversight, graph_usage=N/A2024.09 | 85.4 | — | — | — | |
| GPT4ointeraction_condition=human oversight, graph_usage=false2024.09 | 82.8 | 77.6 | -0.154 | 18.8 | |
| Mixtralinteraction_condition=human oversight, graph_usage=true2024.09 | 82.2 | 79.4 | -0.077 | 16.1 | |
| humaninteraction_condition=programmatic, graph_usage=true2024.09 | 80.1 | 79.9 | -0.151 | 15.1 | |
| humaninteraction_condition=human oversight, graph_usage=true2024.09 | 80.1 | 79.9 | -0.158 | 15 | |
| GPT3.5tinteraction_condition=human oversight, graph_usage=true2024.09 | 78.9 | 80.5 | -0.125 | 21.4 | |
| GPT4ointeraction_condition=programmatic, graph_usage=true2024.09 | 72 | 78 | -0.186 | 13.9 | |
| Llama3interaction_condition=human oversight, graph_usage=true2024.09 | 65.7 | 78.6 | -0.141 | 14.5 | |
| Mixtralinteraction_condition=programmatic, graph_usage=true2024.09 | 55.9 | 76.1 | -0.149 | 12 | |
| GPT3.5tinteraction_condition=programmatic, graph_usage=true2024.09 | 53.1 | 75.9 | -0.178 | 16.3 | |
| Llama3interaction_condition=programmatic, graph_usage=true2024.09 | 17 | 75.2 | -0.274 | 9.8 |