Code Generation on MBPP (Answer correctness rate)
29.9Answer Correctness RateLatent Thinking Optimization w. LRM
Evaluation Results
| Method | Links | |
|---|---|---|
| Latent Thinking Optimization w. LRMModel=Huginn-3.5B2025.09 | 29.9 | |
| Weighted Majority Voting w. LRMModel=Huginn-3.5B2025.09 | 29.5 | |
| Majority VotingModel=Huginn-3.5B2025.09 | 28.8 | |
| Self-Correction w. Confidence ScoreModel=Huginn-3.5B2025.09 | 28.8 | |
| Latent Thinking Correction w. CoE-CModel=Huginn-3.5B2025.09 | 28 | |
| Base ModelModel=Huginn-3.5B2025.09 | 27.8 | |
| Latent Thinking Correction w. CoE-RModel=Huginn-3.5B2025.09 | 27.6 | |
| Self-Correction w. Verbal EvaluationModel=Huginn-3.5B2025.09 | 22.6 |