Conversation Forecasting on Forecasting tasks 1.0 (mix)
20.7Brier Score (BS)GPT-4
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GPT-4scaling=post-hoc, prior_type=data prior2024.02 | 20.7 | 8.5 | — | |
| GPT-4scaling=post-hoc, prior_type=direct2024.02 | 21.1 | — | — | |
| GPT-4scaling=none, prior_type=data prior2024.02 | 21.8 | 3.8 | — | |
| Llama-2-chat 70B DFscaling=post-hoc, prior_type=direct2024.02 | 22 | — | — | |
| Llama-2-chat 70B DFscaling=none, prior_type=data prior2024.02 | 22 | 2.7 | — | |
| Llama-2-chat 70B DFscaling=post-hoc, prior_type=data prior2024.02 | 22.1 | 2.3 | — | |
| GPT-4scaling=none, prior_type=bad prior2024.02 | 22.7 | 9 | — | |
| Llama-2-chat 70B IFscaling=post-hoc, prior_type=direct2024.02 | 22.9 | — | — | |
| Llama-2-chat 70B IFscaling=post-hoc, prior_type=data prior2024.02 | 23.1 | -2 | — | |
| Llama-2-chat 70B DFscaling=none, prior_type=bad prior2024.02 | 23.6 | 5 | — | |
| GPT-4scaling=none, prior_type=direct2024.02 | 23.8 | — | — | |
| Llama-2-chat 70B DFscaling=none, prior_type=direct2024.02 | 25 | — | — | |
| Llama-2-chat 70B IFscaling=none, prior_type=direct2024.02 | 66.1 | — | — | |
| Llama-2-chat 70B IFscaling=none, prior_type=data prior2024.02 | 66.1 | -182 | — | |
| Llama-2-chat 70B IFscaling=none, prior_type=bad prior2024.02 | 66.1 | -160 | — | |
| GPT-4evaluation_subset=neg2024.02 | — | — | -11 | |
| GPT-4evaluation_subset=all2024.02 | — | — | -1.1 | |
| Llama-2-chat 70B DFevaluation_subset=neg2024.02 | — | — | -5.6 | |
| Llama-2-chat 70B DFevaluation_subset=all2024.02 | — | — | -1.3 | |
| Llama-2-chat 70B IFevaluation_subset=neg2024.02 | — | — | 3.6 | |
| Llama-2-chat 70B IFevaluation_subset=all2024.02 | — | — | -49.4 |