Kubernetes failure diagnosis on KubeFault 1.0 (test)
91.2EffectivenessMetaKube
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| MetaKubeBackbone Model=Qwen3-8B, Inference Strategy=MetaKube, Evaluation Protocol=GPT-5 Automated Assessment2026.03 | 91.2 | 90.8 | 87.3 | 92.5 | 92.5 | 90.5 | |
| GPT-4.1Backbone Model=GPT-4.1, Inference Strategy=GraphRAG, Evaluation Protocol=GPT-5 Automated Assessment2026.03 | 89.3 | 92.6 | 91.4 | 94.1 | 94.1 | 91.9 | |
| GPT-4.1-miniBackbone Model=GPT-4.1-mini, Inference Strategy=GraphRAG, Evaluation Protocol=GPT-5 Automated Assessment2026.03 | 79.8 | 81.3 | 78.4 | 85.2 | 85.2 | 81.2 | |
| MetaKubeBackbone Model=Qwen3-8B, Inference Strategy=MetaKube, Evaluation Protocol=Human Expert Evaluation2026.03 | 75.6 | 74.2 | 69.8 | 81.2 | 81.2 | 75.2 | |
| GPT-4.1Backbone Model=GPT-4.1, Inference Strategy=GraphRAG, Evaluation Protocol=Human Expert Evaluation2026.03 | 73.8 | 77.8 | 71.2 | 79.4 | 79.4 | 75.6 | |
| GPT-4.1Backbone Model=GPT-4.1, Inference Strategy=Zero-shot, Evaluation Protocol=GPT-5 Automated Assessment2026.03 | 72.1 | 74.3 | 69.8 | 78.9 | 78.9 | 73.8 | |
| Qwen3-8BBackbone Model=Qwen3-8B, Inference Strategy=GraphRAG, Evaluation Protocol=GPT-5 Automated Assessment2026.03 | 66.7 | 69.1 | 64.8 | 73.3 | 73.3 | 68.5 | |
| GPT-4.1-miniBackbone Model=GPT-4.1-mini, Inference Strategy=Zero-shot, Evaluation Protocol=GPT-5 Automated Assessment2026.03 | 61.5 | 63.8 | 59.2 | 69.7 | 69.7 | 63.6 | |
| GPT-4.1-miniBackbone Model=GPT-4.1-mini, Inference Strategy=GraphRAG, Evaluation Protocol=Human Expert Evaluation2026.03 | 59.3 | 61.7 | 57.4 | 68.9 | 68.9 | 61.8 | |
| GPT-4.1Backbone Model=GPT-4.1, Inference Strategy=Zero-shot, Evaluation Protocol=Human Expert Evaluation2026.03 | 56.4 | 59.2 | 52.7 | 64.8 | 64.8 | 58.3 | |
| Qwen3-8BBackbone Model=Qwen3-8B, Inference Strategy=Zero-shot, Evaluation Protocol=GPT-5 Automated Assessment2026.03 | 48.7 | 51.2 | 46.1 | 57.4 | 57.4 | 50.9 | |
| GPT-4.1-miniBackbone Model=GPT-4.1-mini, Inference Strategy=Zero-shot, Evaluation Protocol=Human Expert Evaluation2026.03 | 45.1 | 47.6 | 41.8 | 55.4 | 55.4 | 47.5 | |
| Qwen3-8BBackbone Model=Qwen3-8B, Inference Strategy=GraphRAG, Evaluation Protocol=Human Expert Evaluation2026.03 | 44.2 | 47.8 | 42.1 | 54.6 | 54.6 | 47.2 | |
| Qwen3-8BBackbone Model=Qwen3-8B, Inference Strategy=Zero-shot, Evaluation Protocol=Human Expert Evaluation2026.03 | 31.5 | 35.8 | 28.9 | 42.3 | 42.3 | 34.6 |