DRC Script Synthesis on Rule2DRC 1,000 problems
87.4F1 ScoreOracle
Evaluation Results
| Method | Links | |
|---|---|---|
| OracleBest-of-N=20, Backbone=GPT-OSS-120B2026.05 | 87.4 | |
| SplitTesterBest-of-N=20, Test Layouts=16, Backbone=GPT-OSS-120B2026.05 | 85.6 | |
| CodeMonkeyBest-of-N=20, Test Layouts=16, Backbone=GPT-OSS-120B2026.05 | 85.5 | |
| OracleBest-of-N=15, Backbone=GPT-OSS-120B2026.05 | 85 | |
| SplitTesterBest-of-N=15, Test Layouts=16, Backbone=GPT-OSS-120B2026.05 | 83.8 | |
| CodeMonkeyBest-of-N=15, Test Layouts=16, Backbone=GPT-OSS-120B2026.05 | 83.5 | |
| S*Best-of-N=20, Test Layouts=16, Backbone=GPT-OSS-120B2026.05 | 83.4 | |
| Generated Tests (8 Tests)Best-of-N=20, Test Layouts=8, Backbone=GPT-OSS-120B2026.05 | 83.1 | |
| Generated Tests (16 Tests)Best-of-N=20, Test Layouts=16, Backbone=GPT-OSS-120B2026.05 | 83 | |
| Generated Tests (8 Tests)Best-of-N=15, Test Layouts=8, Backbone=GPT-OSS-120B2026.05 | 81.5 | |
| S*Best-of-N=15, Test Layouts=16, Backbone=GPT-OSS-120B2026.05 | 81.4 | |
| Generated Tests (16 Tests)Best-of-N=15, Test Layouts=16, Backbone=GPT-OSS-120B2026.05 | 81.3 | |
| OracleBest-of-N=10, Backbone=GPT-OSS-120B2026.05 | 80.9 | |
| SplitTesterBest-of-N=10, Test Layouts=16, Backbone=GPT-OSS-120B2026.05 | 80.6 | |
| CodeMonkeyBest-of-N=10, Test Layouts=16, Backbone=GPT-OSS-120B2026.05 | 80 | |
| Generated Tests (16 Tests)Best-of-N=10, Test Layouts=16, Backbone=GPT-OSS-120B2026.05 | 78.6 | |
| Generated Tests (8 Tests)Best-of-N=10, Test Layouts=8, Backbone=GPT-OSS-120B2026.05 | 78.3 | |
| S*Best-of-N=10, Test Layouts=16, Backbone=GPT-OSS-120B2026.05 | 78.2 | |
| OracleBest-of-N=20, Backbone=GPT-OSS-20B2026.05 | 70.8 | |
| OracleBest-of-N=15, Backbone=GPT-OSS-20B2026.05 | 67.7 | |
| SplitTesterBest-of-N=20, Test Layouts=16, Backbone=GPT-OSS-20B2026.05 | 67.7 | |
| CodeMonkeyBest-of-N=20, Test Layouts=16, Backbone=GPT-OSS-20B2026.05 | 66.6 | |
| Generated Tests (8 Tests)Best-of-N=20, Test Layouts=8, Backbone=GPT-OSS-20B2026.05 | 65.2 | |
| Generated Tests (16 Tests)Best-of-N=20, Test Layouts=16, Backbone=GPT-OSS-20B2026.05 | 64.6 | |
| SplitTesterBest-of-N=15, Test Layouts=16, Backbone=GPT-OSS-20B2026.05 | 64.5 | |
| CodeMonkeyBest-of-N=15, Test Layouts=16, Backbone=GPT-OSS-20B2026.05 | 63.9 | |
| S*Best-of-N=20, Test Layouts=16, Backbone=GPT-OSS-20B2026.05 | 63 | |
| Generated Tests (8 Tests)Best-of-N=15, Test Layouts=8, Backbone=GPT-OSS-20B2026.05 | 62.9 | |
| OracleBest-of-N=10, Backbone=GPT-OSS-20B2026.05 | 62.7 | |
| Generated Tests (16 Tests)Best-of-N=15, Test Layouts=16, Backbone=GPT-OSS-20B2026.05 | 62.6 | |
| S*Best-of-N=15, Test Layouts=16, Backbone=GPT-OSS-20B2026.05 | 61.4 | |
| SplitTesterBest-of-N=10, Test Layouts=16, Backbone=GPT-OSS-20B2026.05 | 61 | |
| Generated Tests (8 Tests)Best-of-N=10, Test Layouts=8, Backbone=GPT-OSS-20B2026.05 | 59.5 | |
| Generated Tests (16 Tests)Best-of-N=10, Test Layouts=16, Backbone=GPT-OSS-20B2026.05 | 59.3 | |
| CodeMonkeyBest-of-N=10, Test Layouts=16, Backbone=GPT-OSS-20B2026.05 | 59.2 | |
| S*Best-of-N=10, Test Layouts=16, Backbone=GPT-OSS-20B2026.05 | 59.1 | |
| LLM-as-a-JudgeBest-of-N=20, Backbone=GPT-OSS-120B2026.05 | 47.6 | |
| LLM-as-a-JudgeBest-of-N=10, Backbone=GPT-OSS-120B2026.05 | 47.5 | |
| LLM-as-a-JudgeBest-of-N=15, Backbone=GPT-OSS-120B2026.05 | 46.6 | |
| Sample-1Best-of-N=1, Backbone=GPT-OSS-120B2026.05 | 44 | |
| LLM-as-a-JudgeBest-of-N=20, Backbone=GPT-OSS-20B2026.05 | 36.4 | |
| OracleBest-of-N=20, Backbone=Qwen3-30B-A3B-Instruct-25072026.05 | 35.9 | |
| LLM-as-a-JudgeBest-of-N=15, Backbone=GPT-OSS-20B2026.05 | 34.8 | |
| OracleBest-of-N=15, Backbone=Qwen3-30B-A3B-Instruct-25072026.05 | 34.7 | |
| LLM-as-a-JudgeBest-of-N=10, Backbone=GPT-OSS-20B2026.05 | 34.6 | |
| OracleBest-of-N=10, Backbone=Qwen3-30B-A3B-Instruct-25072026.05 | 32.9 | |
| SplitTesterBest-of-N=20, Test Layouts=16, Backbone=Qwen3-30B-A3B-Instruct-25072026.05 | 32.1 | |
| SplitTesterBest-of-N=15, Test Layouts=16, Backbone=Qwen3-30B-A3B-Instruct-25072026.05 | 31.5 | |
| CodeMonkeyBest-of-N=20, Test Layouts=16, Backbone=Qwen3-30B-A3B-Instruct-25072026.05 | 30.7 | |
| CodeMonkeyBest-of-N=15, Test Layouts=16, Backbone=Qwen3-30B-A3B-Instruct-25072026.05 | 30.5 | |
| Generated Tests (8 Tests)Best-of-N=20, Test Layouts=8, Backbone=Qwen3-30B-A3B-Instruct-25072026.05 | 30.5 | |
| Generated Tests (16 Tests)Best-of-N=20, Test Layouts=16, Backbone=Qwen3-30B-A3B-Instruct-25072026.05 | 29.8 | |
| Generated Tests (8 Tests)Best-of-N=15, Test Layouts=8, Backbone=Qwen3-30B-A3B-Instruct-25072026.05 | 29.7 | |
| SplitTesterBest-of-N=10, Test Layouts=16, Backbone=Qwen3-30B-A3B-Instruct-25072026.05 | 29.3 | |
| Generated Tests (16 Tests)Best-of-N=15, Test Layouts=16, Backbone=Qwen3-30B-A3B-Instruct-25072026.05 | 29 | |
| CodeMonkeyBest-of-N=10, Test Layouts=16, Backbone=Qwen3-30B-A3B-Instruct-25072026.05 | 28.9 | |
| Generated Tests (8 Tests)Best-of-N=10, Test Layouts=8, Backbone=Qwen3-30B-A3B-Instruct-25072026.05 | 28.4 | |
| Generated Tests (16 Tests)Best-of-N=10, Test Layouts=16, Backbone=Qwen3-30B-A3B-Instruct-25072026.05 | 27.8 | |
| S*Best-of-N=20, Test Layouts=16, Backbone=Qwen3-30B-A3B-Instruct-25072026.05 | 27.3 | |
| S*Best-of-N=15, Test Layouts=16, Backbone=Qwen3-30B-A3B-Instruct-25072026.05 | 26.9 | |
| S*Best-of-N=10, Test Layouts=16, Backbone=Qwen3-30B-A3B-Instruct-25072026.05 | 26.3 | |
| Sample-1Best-of-N=1, Backbone=GPT-OSS-20B2026.05 | 24.3 | |
| LLM-as-a-JudgeBest-of-N=20, Backbone=Qwen3-30B-A3B-Instruct-25072026.05 | 22.1 | |
| LLM-as-a-JudgeBest-of-N=15, Backbone=Qwen3-30B-A3B-Instruct-25072026.05 | 21.6 | |
| LLM-as-a-JudgeBest-of-N=10, Backbone=Qwen3-30B-A3B-Instruct-25072026.05 | 21 | |
| Sample-1Best-of-N=1, Backbone=Qwen3-30B-A3B-Instruct-25072026.05 | 20.4 |