Machine Translation on WMT En-Hr 22 (COMET, BLEURT)
86.9COMETDUAL-REFLECT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DUAL-REFLECTbackbone=ChatGPT, prompting=zero-shot2024.06 | 86.9 | 76.4 | |
| ChatGPT + MAPSprompting=zero-shot2024.06 | 86.5 | 76 | |
| ChatGPTprompting=5-shot2024.06 | 86.4 | — | |
| ChatGPT + Refine_cosprompting=zero-shot2024.06 | 86.4 | 75.9 | |
| ChatGPT + Rerankprompting=zero-shot2024.06 | 86.3 | 75.4 | |
| ChatGPT + Self-Reflectprompting=zero-shot2024.06 | 86.3 | 75.8 | |
| ChatGPT + Refineprompting=zero-shot2024.06 | 86.1 | 75.6 | |
| ChatGPTprompting=zero-shot2024.06 | 85.9 | 75 | |
| DUAL-REFLECTbackbone=Vicuna-7B, prompting=zero-shot2024.06 | 72.9 | 60.4 | |
| Vicuna-7B + MAPSprompting=zero-shot2024.06 | 71.1 | 58.8 | |
| Vicuna-7Bprompting=5-shot2024.06 | 70.2 | 58.1 | |
| DUAL-REFLECTbackbone=Alpaca-7B, prompting=zero-shot2024.06 | 69.5 | 55.4 | |
| Vicuna-7Bprompting=zero-shot2024.06 | 69.3 | 57.7 | |
| Alpaca-7B + MAPSprompting=zero-shot2024.06 | 68.1 | 54.2 | |
| Alpaca-7Bprompting=5-shot2024.06 | 67.9 | 53.6 | |
| Alpaca-7Bprompting=zero-shot2024.06 | 65.9 | 53.2 |