RO reformulation on Random
96.9AccuracyAutoREM
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| AutoREMBase LLM=DeepSeek-V3.2 (Chat Mode)2026.05 | 96.9 | 2,006 | |
| ACEBase LLM=DeepSeek-V3.2 (Chat Mode)2026.05 | 71.9 | 3,184 | |
| Expert PromptBase LLM=DeepSeek-V3.2 (Chat Mode)2026.05 | 68.8 | 3,734 | |
| Max ThinkingBase LLM=DeepSeek-V3.2 (Thinking Mode)2026.05 | 59.4 | 13,971 | |
| ReasoningBankBase LLM=DeepSeek-V3.2 (Chat Mode)2026.05 | 53.1 | 2,510 | |
| Base LLMBase LLM=DeepSeek-V3.2 (Chat Mode)2026.05 | 46.9 | 4,748 |