# SeedRealtime Audio-Visual Full-Duplex LLM Released: Toward Omni-Modal Natural Interaction

ByteDance released SeedRealtime, a native audio-visual full-duplex LLM, and has fully rolled it out for large-scale deployment.

- End-to-end human evaluation shows it cuts audio-visual conversational pacing issues by half versus cascaded models.
- It proactively speaks up when it spots visual changes, like a key target appearing, and can embed tool calls in responses.
- It uses visual context to resolve homophone ambiguity and interpret temporal references in what it sees.
- It distinguishes bystander chatter and background noise to avoid false triggers in noisy environments.

## Why it matters

Developers building real-time AI assistants can now deploy a single end-to-end model that watches, listens, and speaks, replacing cascaded ASR-VLM-TTS pipelines and their latency.

## Sources

- [ByteDance Seed: SeedRealtime Audio-Visual Full-Duplex LLM Released: Toward Omni-Modal Natural Interaction](https://seed.bytedance.com/blog/seedrealtime-audio-visual-full-duplex-llm-released-toward-omni-modal-natural-interaction)

---

Summarized by dstilled on 2026-08-05. https://dstilled.ai/story/241930e4-106a-4b55-b09c-b40ab715c179
