Logical Reasoning on CLUTRR rob_train_disc_23_all (test)
41.6AccuracyLlama-3.1-8B-it (w/ SLR)
Evaluation Results
| Method | Links | |
|---|---|---|
| Llama-3.1-8B-it (w/ SLR)Model=Llama-3.1-8B-it, Variant=w/ SLR, Decoding Strategy=Greedy decoding2025.06 | 41.6 | |
| Llama-3.1-8B-it (Base)Model=Llama-3.1-8B-it, Variant=Base, Decoding Strategy=Greedy decoding2025.06 | 29.7 | |
| Llama-3.1-8B-it (w/ DeepSeek-R1)Model=Llama-3.1-8B-it, Variant=w/ DeepSeek-R1, Decoding Strategy=Greedy decoding2025.06 | 18 |