API Question Answering on APIBench TensorFlow Hub (test)
88.91AccuracyAccurateRAG
Evaluation Results
| Method | Links | |
|---|---|---|
| AccurateRAGLLM Backbone=CodeGemma1.1-7b-it, Fine-tuning=true2025.10 | 88.91 | |
| AccurateRAGLLM Backbone=Llama-2-7B, Fine-tuning=true2025.10 | 88.03 | |
| RAFTLLM Backbone=Llama-2-7B, Fine-tuning=true, Chain-of-Thought (CoT)=GPT-4 CoT2025.10 | 86.86 | |
| RAFTLLM Backbone=Llama-2-7B, Fine-tuning=true, Chain-of-Thought (CoT)=false2025.10 | 83.21 |