Novelty Estimation on AI-Researcher (test)
0.37Pearson RIdeationEval
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| IdeationEvalEvaluation protocol=Decomposed Novelty Assessment2026.01 | 0.37 | 0.384 | |
| GPT-5 MiniEvaluation protocol=LLM-as-a-Judge2026.01 | 0.369 | 0.385 | |
| SPECTER2Evaluation protocol=Embedding-based Retrieval, Input representation=Concat2026.01 | 0.333 | 0.359 | |
| SPECTER2Evaluation protocol=Embedding-based Retrieval, Input representation=Summary2026.01 | 0.316 | 0.357 | |
| E5base-v2Evaluation protocol=Embedding-based Retrieval, Input representation=Concat2026.01 | 0.287 | 0.309 | |
| E5base-v2Evaluation protocol=Embedding-based Retrieval, Input representation=Summary2026.01 | 0.236 | 0.245 | |
| DeepSeek-R1-Distill-Llama-8BEvaluation protocol=LLM-as-a-Judge2026.01 | 0.223 | 0.233 | |
| Qwen3-8BEvaluation protocol=LLM-as-a-Judge2026.01 | 0.218 | 0.222 | |
| SciNCLEvaluation protocol=Embedding-based Retrieval, Input representation=Summary2026.01 | 0.169 | 0.144 | |
| GPT-5 NanoEvaluation protocol=LLM-as-a-Judge2026.01 | 0.168 | 0.211 | |
| SciNCLEvaluation protocol=Embedding-based Retrieval, Input representation=Concat2026.01 | 0.157 | 0.122 | |
| SPECTER2Evaluation protocol=Embedding-based Retrieval, Input representation=Title2026.01 | 0.114 | 0.096 | |
| E5base-v2Evaluation protocol=Embedding-based Retrieval, Input representation=Title2026.01 | 0.089 | 0.032 | |
| Llama3.1-8BEvaluation protocol=LLM-as-a-Judge2026.01 | 0.053 | 0.089 | |
| SciNCLEvaluation protocol=Embedding-based Retrieval, Input representation=Title2026.01 | 0.018 | 0.003 |