Reasoning about Structured Pedagogical Configuration on language tutoring scenario 2024 (Level 2)
0.937Pearson Correlation (Bayesian vs LLM)APV
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| APVPIIE formalism=true, Prompting=Average across direct and CoT, Perspective=Average across first-person and assistant2026.07 | 0.937 | 0.935 | 0.958 | |
| GPT-4oPrompting=Average across direct and CoT, Perspective=Average across first-person and assistant2026.07 | 0.911 | 0.929 | 0.943 | |
| Llama 3.3-70BPrompting=Average across direct and CoT, Perspective=Average across first-person and assistant2026.07 | 0.876 | 0.923 | 0.922 | |
| Claude 3.5 SonnetPrompting=Average across direct and CoT, Perspective=Average across first-person and assistant2026.07 | 0.845 | 0.889 | 0.941 | |
| Gemini 2.0 FlashPrompting=Average across direct and CoT, Perspective=Average across first-person and assistant2026.07 | 0.788 | 0.901 | 0.925 | |
| o3-miniPrompting=Average across direct and CoT, Perspective=Average across first-person and assistant2026.07 | 0.716 | 0.869 | 0.712 | |
| o1Prompting=Average across direct and CoT, Perspective=Average across first-person and assistant2026.07 | 0.705 | 0.894 | 0.861 | |
| Llama 3.1-8BPrompting=Average across direct and CoT, Perspective=Average across first-person and assistant2026.07 | 0.608 | 0.813 | 0.701 | |
| Llama 3.2-3BPrompting=Average across direct and CoT, Perspective=Average across first-person and assistant2026.07 | 0.349 | 0.586 | 0.55 | |
| DeepSeek-R1Prompting=Average across direct and CoT, Perspective=Average across first-person and assistant2026.07 | 0.326 | 0.492 | 0.643 | |
| Gemma 3-4BPrompting=Average across direct and CoT, Perspective=Average across first-person and assistant2026.07 | 0.288 | 0.34 | 0.266 |