ByteDance Seed Introduces SeedRealtime: a Native Audio-Visual Full-Duplex LLM That Watches, Listens and Speaks in One Model
ByteDance’s Seed staff has launched SeedRealtime, a native audio-visual full-duplex LLM. The mannequin fuses audio, video and textual content in a single unified structure. It interacts in actual time over steady multimodal streams, fairly than one flip at a time. Seed positions it as a step towards omni-modal interplay, and claims three breakthroughs: joint audio-visual understanding, proactive interplay, and pure conversational timing. The architectural goal is the cascade: chained ASR, VLM and TTS modules that add latency and lose info between phases. SeedRealtime as an alternative runs notion, understanding, decision-making and expression in parallel inside one end-to-end mannequin. Turn-taking strikes contained in the mannequin as nicely, changing the exterior voice-activity detector most real-time stacks nonetheless rely on.
Is it deployable?
It is partly deployable.
SeedRealtime is live inside the Doubao app, ByteDance’s client assistant. For this particular mannequin, ByteDance has printed no technical report, no parameter depend, no open weights, and no Volcano Engine or BytePlus endpoint. As a third-party staff, you can not combine it as of now. What is deployable proper now could be the thought: a validated reference structure, and a moved goalpost for anybody delivery real-time voice-plus-camera merchandise.
Interactive explainer
What is definitely new in the demos
Seed printed seven situations. Four are load-bearing.
- Identity binding throughout modalities: At a noisy group dinner, the mannequin matches names to faces as individuals are launched, then retains every voice tied to its identification — attributing conflicting journey preferences to the best speaker earlier than proposing a plan.
- Proactive speech from a held instruction: At the Hebei Museum, a consumer asks to be reminded when a particular bronze display stand seems. The digital camera retains panning; the mannequin watches and speaks up unprompted when the piece enters body. The identical habits reveals up on a ResNet paper — the mannequin tracks quick web page flips, spots the “3.4 Implementation” part, pauses by itself, and reads out studying charge, momentum and weight decay.
- Correction from visible state, not from a query: Watching an espresso workflow, the mannequin interrupts when complete beans go into the portafilter, then reads crema coloration and quantity and suggests shortening extraction by 2 to three seconds.
- Interference suppression pand off-screen reminiscence: At Beijing Daxing Airport, unrelated chatter about a flight doesn’t set off a reply. When the consumer really asks, the mannequin solutions from departure-board info that has already scrolled off display, and goes on-line for the baggage-carousel location.
Key Takeaways
- SeedRealtime is a native audio-visual full-duplex LLM — audio, video and textual content in one end-to-end structure.
- Turn-taking strikes contained in the mannequin; no exterior VAD decides when to talk.
- ByteDance’s personal human eval reviews pacing points halved versus cascaded stacks — no benchmark, no latency numbers.
- It is reside in the Doubao app, however there is no such thing as a technical report, no weights and no introduced API.
Check out the ByteDance Seed launch post and Seed models page. Also, be at liberty to comply with us on Twitter and don’t neglect to affix our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to associate with us for selling your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar and so forth.? Connect with us
The put up ByteDance Seed Introduces SeedRealtime: a Native Audio-Visual Full-Duplex LLM That Watches, Listens and Speaks in One Model appeared first on MarkTechPost.
