Conversational Numerical Reasoning Question Answering on ConvFinQA v1 (dev)
0.7846Execution AccuracyAPOLLO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| APOLLOProtocol=Fine-tuning, Ensemble=true2022.12 | 0.7846 | 0.7591 | |
| GPT-4Protocol=Prompting, Mode=zero-shot2022.12 | 0.7648 | — | |
| APOLLOProtocol=Fine-tuning2022.12 | 0.7647 | 0.7414 | |
| FinQANetProtocol=Fine-tuning, Backbone=ROBERTa-large2022.12 | 0.6832 | 0.6787 | |
| GPT-3.5-turboProtocol=Prompting, Mode=zero-shot2022.12 | 0.5986 | — | |
| GPT-2Protocol=Fine-tuning, Backbone=ROBERTa-large2022.12 | 0.5912 | 0.5752 | |
| T-5Protocol=Fine-tuning, Backbone=ROBERTa-large2022.12 | 0.5838 | 0.5671 | |
| BloombergGPTProtocol=Prompting2022.12 | 0.4341 | — |