Memory Imitation on InVivoGPT English-dominant users 1.0 (35 users, ~6.9k queries)
0.43BLEU Recall (GT)Qwen2.5-32B-it
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Qwen2.5-32B-itProtocol=Fine-tuned (FT), Inference time=130-390 ms per query2026.02 | 0.43 | 0.39 | 0.7 | 0.36 | 0.31 | 0.63 | 0.56 | 0.45 | 0.59 | |
| Qwen2.5-32B-itProtocol=In-Context Learning (ICL), Inference time=130-390 ms per query2026.02 | 0.41 | 0.35 | 0.68 | 0.27 | 0.22 | 0.66 | 0.53 | 0.41 | 0.69 | |
| Gemma3-27B-itProtocol=In-Context Learning (ICL), Inference time=130-390 ms per query2026.02 | 0.39 | 0.32 | 0.62 | 0.23 | 0.18 | 0.62 | 0.44 | 0.32 | 0.66 | |
| GPT-OSS-20BProtocol=Fine-tuned (FT), Inference time=130-390 ms per query2026.02 | 0.35 | 0.32 | 0.65 | 0.35 | 0.31 | 0.58 | 0.52 | 0.44 | 0.52 | |
| Gemma3-27B-itProtocol=Fine-tuned (FT), Inference time=130-390 ms per query2026.02 | 0.31 | 0.28 | 0.59 | 0.34 | 0.29 | 0.53 | 0.5 | 0.42 | 0.39 | |
| GPT-OSS-20BProtocol=In-Context Learning (ICL), Inference time=130-390 ms per query2026.02 | 0.19 | 0.16 | 0.49 | 0.2 | 0.18 | 0.55 | 0.37 | 0.32 | 0.5 |