Knowledge-to-action Question Answering on 50 knowledge-to-action questions
4.37Overall ScoreText-chunk RAG
Evaluation Results
| Method | Links | |
|---|---|---|
| Text-chunk RAGModel=phi-4, Evaluator=Llama-3.3-it2025.10 | 4.37 | |
| Vanilla LLMModel=phi-4, Evaluator=Llama-3.3-it2025.10 | 4.24 | |
| KEOModel=phi-4, Evaluator=Llama-3.3-it2025.10 | 4.18 | |
| Text-chunk RAGModel=phi-4, Evaluator=GPT-4o2025.10 | 4.17 | |
| Vanilla LLMModel=phi-4, Evaluator=GPT-4o2025.10 | 4.01 | |
| Text-chunk RAGModel=mistral-nemo-it, Evaluator=Llama-3.3-it2025.10 | 4 | |
| KEOModel=phi-4, Evaluator=GPT-4o2025.10 | 3.96 | |
| Text-chunk RAGModel=gemma-3-it, Evaluator=Llama-3.3-it2025.10 | 3.96 | |
| KEOModel=gemma-3-it, Evaluator=Llama-3.3-it2025.10 | 3.93 | |
| Vanilla LLMModel=gemma-3-it, Evaluator=Llama-3.3-it2025.10 | 3.9 | |
| KEOModel=gemma-3-it, Evaluator=GPT-4o2025.10 | 3.86 | |
| Text-chunk RAGModel=gemma-3-it, Evaluator=GPT-4o2025.10 | 3.84 | |
| KEOModel=mistral-nemo-it, Evaluator=Llama-3.3-it2025.10 | 3.8 | |
| Vanilla LLMModel=mistral-nemo-it, Evaluator=Llama-3.3-it2025.10 | 3.78 | |
| Vanilla LLMModel=gemma-3-it, Evaluator=GPT-4o2025.10 | 3.75 | |
| Text-chunk RAGModel=mistral-nemo-it, Evaluator=GPT-4o2025.10 | 3.72 | |
| Vanilla LLMModel=mistral-nemo-it, Evaluator=GPT-4o2025.10 | 3.7 | |
| KEOModel=mistral-nemo-it, Evaluator=GPT-4o2025.10 | 3.68 |