Story evaluation on OpenMEVA
0.2092Pearson CorrelationEvolvR
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| EvolvR2025.08 | 0.2092 | 0.2087 | 0.1881 | |
| AutoJ-13BModel Category=NLG Evaluation Open-Source Models2025.08 | 0.1871 | 0.1952 | 0.1611 | |
| TIGERScore-13BModel Category=NLG Evaluation Open-Source Models2025.08 | 0.1661 | 0.1649 | 0.1541 | |
| Qwen2.5-7B-Instruct + Pointwise CoT + GRPOModel Category=Finetune Open-source Models, Chain-of-Thought (CoT) configuration=Pointwise, Reinforcement learning (GRPO)=true2025.08 | 0.1605 | 0.1654 | 0.1415 | |
| InstructScore-7BModel Category=NLG Evaluation Open-Source Models2025.08 | 0.1578 | 0.1257 | 0.1046 | |
| Qwen2.5-7B-Instruct + Pointwise CoTModel Category=Finetune Open-source Models, Chain-of-Thought (CoT) configuration=Pointwise2025.08 | 0.1474 | 0.1503 | 0.1291 | |
| Themis-8BModel Category=NLG Evaluation Open-Source Models2025.08 | 0.1469 | 0.1719 | 0.1457 | |
| Qwen2.5-7B-InstructModel Category=Finetune Open-source Models2025.08 | 0.1061 | 0.1175 | 0.0976 | |
| Qwen2.5-7B-Instruct + GRPOModel Category=Finetune Open-source Models, Reinforcement learning (GRPO)=true2025.08 | 0.0947 | 0.0994 | 0.0897 |