Text chat feels too flat
Most AI companions can write back quickly, but they cannot look present, react with visible listening behavior, or make a conversation feel face to face.
Wan Streamer is a prelaunch real-time AI video companion for face-to-face conversations with an AI that can see, hear, understand, and respond with synchronized voice, expression, and motion.
People do not communicate in clean text turns. We interrupt, pause, look away, smile, listen, and react. Wan Streamer is built around that product need: a real-time AI avatar that feels present instead of a chatbot placed behind a video renderer.
Most AI companions can write back quickly, but they cannot look present, react with visible listening behavior, or make a conversation feel face to face.
A voice-only agent can answer, but it cannot show gaze, expression, timing, or motion. Wan Streamer is designed around synchronized audio and video presence.
Traditional cascaded pipelines wait for separate speech recognition, language, speech synthesis, animation, and rendering modules before the user sees a response.
Language practice, coaching, emotional check-ins, and creative role play work better when the AI can hear, see, speak, and respond with natural motion.
Public Wan-Streamer research describes an end-to-end interactive foundation model for language, audio, and video as both input and output. Instead of relying on separate speech, language, animation, and video-generation modules, the model direction learns perception, reasoning, response timing, and cross-modal synchronization in one streaming system.
Wan Streamer turns that research signal into a focused product landing page: waitlist first, video companion experience next, then pricing, login, credits, storage, and privacy controls after the core beta flow is ready.
Wan-Streamer introduces a native-streaming interaction model where perception and response are not forced into strict turns. Wan Streamer turns that direction into a product waitlist.
The product concept centers on AI avatars that answer with voice, facial movement, gaze, and visible listening behavior instead of detached text or audio alone.
The research direction avoids stitching together separate VAD, ASR, LLM, TTS, animation, and video modules, reducing the latency and drift that make avatars feel mechanical.
The early product should focus on high-frequency emotional and conversational needs before expanding into business-facing digital humans. That keeps the first version clear: talk to a visible AI, test the interaction, save the best moments, and share only when you choose.
Open a Wan Streamer live video conversation when you want to talk without waiting for a friend to be available.
Practice role play, interview drills, pronunciation, and everyday dialogue with a visible partner.
Prototype Wan Streamer-style customer-facing AI hosts, onboarding guides, video concierges, and livestream assistants.
Turn natural conversations into 15-60 second highlights for TikTok, Reels, Xiaohongshu, or Douyin.
Join the waitlist for early access updates, beta character slots, product notes, and Wan-Streamer API availability tracking. The model is not publicly available as a finished product yet, so the first step is building a real audience before wiring payments and credits.
Key questions about Wan Streamer, Wan-Streamer, early access, privacy, and the current API status.
Wan Streamer is a prelaunch real-time AI video companion and digital human concept for users who want face-to-face AI conversation. The page tracks the Wan-Streamer research direction while preparing a product experience around video chat, voice, expression, and sharing.
No. wanstreamer.org is an independent product page. It references public Wan-Streamer research from the Wan Team / Alibaba Group and will clearly separate research status from product availability.
A public production Wan-Streamer API has not been confirmed. The current site is positioned as a waitlist and research-to-product tracker, not as an active API console.
Most avatar products are cascaded systems: speech recognition, language model, text-to-speech, animation, and rendering are separate. Wan-Streamer points toward native streaming audio-visual interaction in one model, which can reduce latency and improve synchronization.
The landing page is English-first for SEO, but the product direction should support Chinese and English conversations because real-time video companionship, language practice, and creator sharing are cross-market use cases.
Camera and microphone features require clear consent, visible controls, and a privacy-first beta. Saved clips should be opt-in, and private conversations should not become public sharing content without user action.
Launch-stage promise: no fake API access claims, no fake model availability, and no hidden pricing until the beta workflow is real.