Open-ended Generation on HumanEval+
0.44FDRSGR
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| SGRModel=Qwen3-4B-Instruct-2507, Uncertainty Score=logits-based2025.10 | 0.44 | 5.6 | 5.38 | — | |
| Conformal LabelingModel=DeepSeek-R1-Distill-Qwen-32B, Uncertainty Score=logits-based2025.10 | 0.61 | 100 | 100 | — | |
| FDR searchModel=DeepSeek-R1-Distill-Qwen-32B, Uncertainty Score=logits-based2025.10 | 0.62 | 100 | 100 | — | |
| SGRModel=DeepSeek-R1-Distill-Qwen-32B, Uncertainty Score=logits-based2025.10 | 0.64 | 39.48 | 39.49 | — | |
| SGRModel=Qwen3-4B-Thinking-2507, Uncertainty Score=logits-based2025.10 | 1.61 | 14.73 | 14.79 | — | |
| Conformal LabelingModel=Qwen3-4B-Thinking-2507, Uncertainty Score=logits-based2025.10 | 2.44 | 100 | 100 | — | |
| FDR searchModel=Qwen3-4B-Thinking-2507, Uncertainty Score=logits-based2025.10 | 2.44 | 100 | 100 | — | |
| Conformal LabelingModel=Qwen3-4B-Instruct-2507, Uncertainty Score=logits-based2025.10 | 4.76 | 72.46 | 71.06 | — | |
| FDR searchModel=Qwen3-4B-Instruct-2507, Uncertainty Score=logits-based2025.10 | 5.09 | 81.86 | 80.08 | — | |
| AI onlyModel=Qwen3-4B-Instruct-25072025.10 | — | — | — | 7.32 | |
| AI onlyModel=Qwen3-4B-Thinking-25072025.10 | — | — | — | 2.44 | |
| AI onlyModel=DeepSeek-R1-Distill-Qwen-32B2025.10 | — | — | — | 0.61 |