Student Simulation on Java 5
84AccuracyGPT-4o
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GPT-4oBehavior Prediction=Prototype Mapping, Solution Simulation=IO2025.05 | 84 | 3.8 | 3.24 | |
| GPT-4oBehavior Prediction=Prototype Mapping, Solution Simulation=CoT2025.05 | 84 | 3.8 | 3.38 | |
| GPT-4oBehavior Prediction=Prototype Mapping, Solution Simulation=Refine2025.05 | 84 | 3.8 | 3.46 | |
| LLaMA-3.3-70B-InstructBehavior Prediction=Prototype Mapping, Solution Simulation=IO2025.05 | 68 | 3.3 | 2.78 | |
| LLaMA-3.3-70B-InstructBehavior Prediction=Prototype Mapping, Solution Simulation=CoT2025.05 | 68 | 3.3 | 2.82 | |
| LLaMA-3.3-70B-InstructBehavior Prediction=Prototype Mapping, Solution Simulation=Refine2025.05 | 68 | 3.3 | 2.84 | |
| GPT-4oBehavior Prediction=Level+Random, Solution Simulation=IO2025.05 | 66 | 3.56 | 2.44 | |
| GPT-4oBehavior Prediction=Level+Random, Solution Simulation=CoT2025.05 | 66 | 3.56 | 2.64 | |
| GPT-4oBehavior Prediction=Level+Random, Solution Simulation=Refine2025.05 | 66 | 3.56 | 2.4 | |
| Claude-3.5-SonnetBehavior Prediction=Prototype Mapping, Solution Simulation=IO2025.05 | 56 | 3.3 | 2.68 | |
| Claude-3.5-SonnetBehavior Prediction=Prototype Mapping, Solution Simulation=CoT2025.05 | 56 | 3.3 | 2.44 | |
| Claude-3.5-SonnetBehavior Prediction=Prototype Mapping, Solution Simulation=Refine2025.05 | 56 | 3.3 | 2.9 | |
| GPT-4oBehavior Prediction=Level+Similarity, Solution Simulation=IO2025.05 | 56 | 3.28 | 2.24 | |
| GPT-4oBehavior Prediction=Level+Similarity, Solution Simulation=CoT2025.05 | 56 | 3.28 | 2.52 | |
| GPT-4oBehavior Prediction=Level+Similarity, Solution Simulation=Refine2025.05 | 56 | 3.28 | 2.1 | |
| Claude-3.5-SonnetBehavior Prediction=Similarity, Solution Simulation=IO2025.05 | 54 | 3.04 | 2.66 | |
| Claude-3.5-SonnetBehavior Prediction=Similarity, Solution Simulation=CoT2025.05 | 54 | 3.04 | 2.8 | |
| Claude-3.5-SonnetBehavior Prediction=Similarity, Solution Simulation=Refine2025.05 | 54 | 3.04 | 2.76 | |
| GPT-3.5Behavior Prediction=Prototype Mapping, Solution Simulation=IO2025.05 | 52 | 2.92 | 2.68 | |
| GPT-3.5Behavior Prediction=Prototype Mapping, Solution Simulation=CoT2025.05 | 52 | 2.92 | 2.98 | |
| GPT-3.5Behavior Prediction=Prototype Mapping, Solution Simulation=Refine2025.05 | 52 | 2.92 | 3.28 | |
| GPT-4oBehavior Prediction=Similarity, Solution Simulation=IO2025.05 | 50 | 2.9 | 2.74 | |
| GPT-4oBehavior Prediction=Similarity, Solution Simulation=CoT2025.05 | 50 | 2.9 | 2.78 | |
| GPT-4oBehavior Prediction=Similarity, Solution Simulation=Refine2025.05 | 50 | 2.9 | 2.74 | |
| LLaMA-3.3-70B-InstructBehavior Prediction=Random, Solution Simulation=IO2025.05 | 46 | 2.5 | 2.34 | |
| LLaMA-3.3-70B-InstructBehavior Prediction=Random, Solution Simulation=CoT2025.05 | 46 | 2.5 | 2.4 | |
| LLaMA-3.3-70B-InstructBehavior Prediction=Random, Solution Simulation=Refine2025.05 | 46 | 2.5 | 1.98 | |
| Claude-3.5-SonnetBehavior Prediction=Random, Solution Simulation=IO2025.05 | 46 | 2.88 | 2.46 | |
| Claude-3.5-SonnetBehavior Prediction=Random, Solution Simulation=CoT2025.05 | 46 | 2.88 | 2.46 | |
| Claude-3.5-SonnetBehavior Prediction=Random, Solution Simulation=Refine2025.05 | 46 | 2.88 | 2.8 | |
| Claude-3.5-SonnetBehavior Prediction=Level+Random, Solution Simulation=IO2025.05 | 46 | 2.48 | 2.24 | |
| Claude-3.5-SonnetBehavior Prediction=Level+Random, Solution Simulation=CoT2025.05 | 46 | 2.48 | 2 | |
| Claude-3.5-SonnetBehavior Prediction=Level+Random, Solution Simulation=Refine2025.05 | 46 | 2.48 | 2.36 | |
| Claude-3.5-SonnetBehavior Prediction=Level+Similarity, Solution Simulation=IO2025.05 | 46 | 2.56 | 2.5 | |
| Claude-3.5-SonnetBehavior Prediction=Level+Similarity, Solution Simulation=CoT2025.05 | 46 | 2.56 | 2.42 | |
| Claude-3.5-SonnetBehavior Prediction=Level+Similarity, Solution Simulation=Refine2025.05 | 46 | 2.56 | 2.4 | |
| GPT-3.5Behavior Prediction=Random, Solution Simulation=IO2025.05 | 46 | 2.36 | 2.44 | |
| GPT-3.5Behavior Prediction=Random, Solution Simulation=CoT2025.05 | 46 | 2.36 | 2.78 | |
| GPT-3.5Behavior Prediction=Random, Solution Simulation=Refine2025.05 | 46 | 2.36 | 3.22 | |
| GPT-3.5Behavior Prediction=Level, Solution Simulation=IO2025.05 | 46 | 2.74 | 2.52 | |
| GPT-3.5Behavior Prediction=Level, Solution Simulation=CoT2025.05 | 46 | 2.74 | 2.84 | |
| GPT-3.5Behavior Prediction=Level, Solution Simulation=Refine2025.05 | 46 | 2.74 | 2.86 | |
| GPT-3.5Behavior Prediction=Level+Random, Solution Simulation=IO2025.05 | 46 | 2.5 | 2.54 | |
| GPT-3.5Behavior Prediction=Level+Random, Solution Simulation=CoT2025.05 | 46 | 2.5 | 2.24 | |
| GPT-3.5Behavior Prediction=Level+Random, Solution Simulation=Refine2025.05 | 46 | 2.5 | 3.18 | |
| GPT-3.5Behavior Prediction=Level+Similarity, Solution Simulation=IO2025.05 | 42 | 2.48 | 2.42 | |
| GPT-3.5Behavior Prediction=Level+Similarity, Solution Simulation=CoT2025.05 | 42 | 2.48 | 2.44 | |
| GPT-3.5Behavior Prediction=Level+Similarity, Solution Simulation=Refine2025.05 | 42 | 2.48 | 2.96 | |
| GPT-4oBehavior Prediction=Level, Solution Simulation=IO2025.05 | 42 | 2.34 | 2.28 | |
| GPT-4oBehavior Prediction=Level, Solution Simulation=CoT2025.05 | 42 | 2.34 | 2.8 | |
| GPT-4oBehavior Prediction=Level, Solution Simulation=Refine2025.05 | 42 | 2.34 | 1.9 | |
| LLaMA-3.3-70B-InstructBehavior Prediction=Similarity, Solution Simulation=IO2025.05 | 34 | 2.32 | 2.2 | |
| LLaMA-3.3-70B-InstructBehavior Prediction=Similarity, Solution Simulation=CoT2025.05 | 34 | 2.32 | 2.34 | |
| LLaMA-3.3-70B-InstructBehavior Prediction=Similarity, Solution Simulation=Refine2025.05 | 34 | 2.32 | 2.04 | |
| LLaMA-3.3-70B-InstructBehavior Prediction=Level, Solution Simulation=IO2025.05 | 34 | 2.18 | 2.74 | |
| LLaMA-3.3-70B-InstructBehavior Prediction=Level, Solution Simulation=CoT2025.05 | 34 | 2.18 | 2.14 | |
| LLaMA-3.3-70B-InstructBehavior Prediction=Level, Solution Simulation=Refine2025.05 | 34 | 2.18 | 1.66 | |
| LLaMA-3.3-70B-InstructBehavior Prediction=Level+Similarity, Solution Simulation=IO2025.05 | 34 | 2.34 | 2.28 | |
| LLaMA-3.3-70B-InstructBehavior Prediction=Level+Similarity, Solution Simulation=CoT2025.05 | 34 | 2.34 | 2.48 | |
| LLaMA-3.3-70B-InstructBehavior Prediction=Level+Similarity, Solution Simulation=Refine2025.05 | 34 | 2.34 | 2.16 | |
| GPT-3.5Behavior Prediction=Similarity, Solution Simulation=IO2025.05 | 34 | 2.44 | 2.56 | |
| GPT-3.5Behavior Prediction=Similarity, Solution Simulation=CoT2025.05 | 34 | 2.44 | 2.44 | |
| GPT-3.5Behavior Prediction=Similarity, Solution Simulation=Refine2025.05 | 34 | 2.44 | 3.02 | |
| LLaMA-3.3-70B-InstructBehavior Prediction=Level+Random, Solution Simulation=IO2025.05 | 32 | 1.88 | 1.86 | |
| LLaMA-3.3-70B-InstructBehavior Prediction=Level+Random, Solution Simulation=CoT2025.05 | 32 | 1.88 | 2.18 | |
| LLaMA-3.3-70B-InstructBehavior Prediction=Level+Random, Solution Simulation=Refine2025.05 | 32 | 1.88 | 2 | |
| Claude-3.5-SonnetBehavior Prediction=Level, Solution Simulation=IO2025.05 | 30 | 1.8 | 1.86 | |
| Claude-3.5-SonnetBehavior Prediction=Level, Solution Simulation=CoT2025.05 | 30 | 1.8 | 1.58 | |
| Claude-3.5-SonnetBehavior Prediction=Level, Solution Simulation=Refine2025.05 | 30 | 1.8 | 1.32 | |
| GPT-4oBehavior Prediction=Random, Solution Simulation=IO2025.05 | 30 | 2.48 | 2.28 | |
| GPT-4oBehavior Prediction=Random, Solution Simulation=CoT2025.05 | 30 | 2.48 | 2.5 | |
| GPT-4oBehavior Prediction=Random, Solution Simulation=Refine2025.05 | 30 | 2.48 | 2.6 |