Differential Expression on PerturbQA RPE1
53.06AccuracyBoN w/ Full Set
Evaluation Results
| Method | Links | |
|---|---|---|
| BoN w/ Full SetDecoding Strategy=Best-of-N, Reward Model=Full Set2026.03 | 53.06 | |
| BoN w/ Llama-PRM800KDecoding Strategy=Best-of-N, Reward Model=Llama-PRM800K2026.03 | 45.2 | |
| BoN w/ VersaPRMDecoding Strategy=Best-of-N, Reward Model=VersaPRM2026.03 | 43.73 | |
| BoN w/ Qwen-PRM800KDecoding Strategy=Best-of-N, Reward Model=Qwen-PRM800K2026.03 | 42.03 | |
| Majority VotingDecoding Strategy=Majority Voting2026.03 | 38.88 | |
| Greedy DecodingDecoding Strategy=Greedy2026.03 | 37.52 |