Real-time Omni-modal Interaction on Real-time Omni-modal Interaction Various
200Model-side Latency (ms)Wan-Streamer
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Wan-StreamerInteraction=text/audio/video in/out, Comparison boundary=One end-to-end model; text I/O, speech, and synchronized visual response share one causal stream.2026.06 | 200 | 550 | — | — | — | 25 | |
| GPT-4o / Realtime APIInteraction=speech-to-speech, audio/vision input, Comparison boundary=Reported numbers mix model response, API TTFB, endpointing, and network.2026.06 | 232 | — | — | — | — | — | |
| MiniCPM-o 4.5Interaction=audio-video in, speech/text out, Comparison boundary=First-token/RTF metric; no visual avatar generation.2026.06 | — | — | — | 0.58 | 0.2 | — | |
| Qwen3/3.5-OmniInteraction=audio-video-text in, speech/text out, Comparison boundary=First-packet metric; no synchronized visual avatar generation.2026.06 | — | — | 234 | — | — | — |