Knowledge Graph Factuality Evaluation on FActScore
84FActScoreGraphMERT
Evaluation Results
| Method | Links | |
|---|---|---|
| GraphMERTModel size=80M, #triples=138,201, LLM Judge=Gemini 2.0 Flash, Evaluation Mode=Context + General truth2025.10 | 84 | |
| GraphMERTModel size=80M, #triples=138,201, LLM Judge=Gemini 2.0 Flash, Evaluation Mode=Context only2025.10 | 79.7 | |
| Grok 4 FastModel size=1.7T, #triples=100,000, LLM Judge=Qwen3-32B, Evaluation Mode=Context + General truth2025.10 | 79.7 | |
| Grok 4 FastModel size=1.7T, #triples=100,000, LLM Judge=Qwen3-32B, Evaluation Mode=Context only2025.10 | 79.4 | |
| Grok 4 FastModel size=1.7T, #triples=100,000, LLM Judge=Gemini 2.0 Flash, Evaluation Mode=Context only2025.10 | 73.3 | |
| GraphMERTModel size=80M, #triples=138,201, LLM Judge=Qwen3-32B, Evaluation Mode=Context + General truth2025.10 | 72.2 | |
| GraphMERTModel size=80M, #triples=138,201, LLM Judge=Qwen3-32B, Evaluation Mode=Context only2025.10 | 69.8 | |
| Grok 4 FastModel size=1.7T, #triples=100,000, LLM Judge=Gemini 2.0 Flash, Evaluation Mode=Context + General truth2025.10 | 67.9 | |
| Qwen3-14BModel size=14B, #triples=100,000, LLM Judge=Gemini 2.0 Flash, Evaluation Mode=Context only2025.10 | 62.4 | |
| Qwen3-14BModel size=14B, #triples=100,000, LLM Judge=Gemini 2.0 Flash, Evaluation Mode=Context + General truth2025.10 | 57.6 | |
| Qwen3-32B (baseline)Model size=32B, #triples=515,460, LLM Judge=Gemini 2.0 Flash, Evaluation Mode=Context only2025.10 | 49.1 | |
| Qwen3-32B (baseline)Model size=32B, #triples=515,460, LLM Judge=Qwen3-32B, Evaluation Mode=Context + General truth2025.10 | 48.1 | |
| Qwen3-14BModel size=14B, #triples=100,000, LLM Judge=Qwen3-32B, Evaluation Mode=Context only2025.10 | 46.4 | |
| Qwen3-14BModel size=14B, #triples=100,000, LLM Judge=Qwen3-32B, Evaluation Mode=Context + General truth2025.10 | 45.7 | |
| Qwen3-32B (baseline)Model size=32B, #triples=515,460, LLM Judge=Gemini 2.0 Flash, Evaluation Mode=Context + General truth2025.10 | 44.3 | |
| Qwen3-32B (baseline)Model size=32B, #triples=515,460, LLM Judge=Qwen3-32B, Evaluation Mode=Context only2025.10 | 40.2 |