Pearson correlation with human judgment on AI-generated Explorable Explanations Balanced
0.676Functional CorrelationFSM-based (EE-Eval)
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| FSM-based (EE-Eval)Framework=FSM-based2026.06 | 0.676 | 0.69 | 0.728 | 0.58 | 0.696 | |
| VLM (Baseline 1)Framework=VLM (Baseline 1), Implementation=GPT-4o-mini2026.06 | 0.537 | 0.628 | 0.53 | 0.557 | 0.584 | |
| Unit test (Baseline 2)Framework=Unit test (Baseline 2), Implementation=LLM-generated Playwright tests2026.06 | -0.58 | -0.63 | -0.6 | -0.492 | -0.598 |