Arithmetic Reasoning on MAWPS (5-fold cross val)
94.3AccuracyMSAT-DEDUCTREASONER
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| MSAT-DEDUCTREASONERInput Configuration=digit tokenization, Pre-training method=MSAT2023.06 | 94.3 | — | 2.3 | |
| Large language models w/ Chain-of-Thought promptingBackbone=PaLM 540B, Prompting strategy=Chain-of-Thought2023.06 | 93.3 | — | — | |
| PaLM 540B (CoT 8-shot)Model=PaLM 540B, Finetuning Strategy=None (Zero-shot/Few-shot), Prompting Strategy=8-shot Chain-of-thought2022.12 | 93 | 93.66 | — | |
| DEDUCTREASONERInput Configuration=symbolic masks2023.06 | 92 | — | — | |
| MSAT-ROBERTAGENInput Configuration=digit tokenization, Pre-training method=MSAT2023.06 | 91.6 | — | 3.2 | |
| DEDUCTREASONERInput Configuration=digit tokenization2023.06 | 91.6 | — | -0.4 | |
| ROBERTAGENInput Configuration=symbolic masks2023.06 | 88.4 | — | — | |
| ROBERTAGENInput Configuration=digit tokenization2023.06 | 84.1 | — | -4.3 | |
| T5 XXL (CoT Finetuned)Model=T5 XXL, Finetuning Strategy=CoT knowledge distillation, Prompting Strategy=Chain-of-thought2022.12 | 70.41 | 88.22 | — | |
| T5 XXL (Baseline)Model=T5 XXL, Finetuning Strategy=Finetuned on original target labels, Prompting Strategy=Direct2022.12 | 54.15 | — | — |