Conversational Numerical Reasoning Question Answering on ConvFinQA v1 (test)
89.44Execution AccuracyHuman Expert Performance
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Human Expert Performance2022.12 | 89.44 | 86.34 | |
| APOLLOProtocol=Fine-tuning, Ensemble=true2022.12 | 78.76 | 77.19 | |
| APOLLOProtocol=Fine-tuning2022.12 | 76 | 74.56 | |
| FinQANetProtocol=Fine-tuning, Backbone=ROBERTa-large2022.12 | 68.9 | 68.24 | |
| T-5Protocol=Fine-tuning, Backbone=ROBERTa-large2022.12 | 58.66 | 57.05 | |
| GPT-2Protocol=Fine-tuning, Backbone=ROBERTa-large2022.12 | 58.19 | 57 | |
| General Crowd Performance2022.12 | 46.9 | 45.52 |