Fuser Apps are here 🚀
Fuser Apps are here. Free generations for the next month 💫
Let's GobyKuaishou Technology
Transfer complex choreography, micro-expressions, and vocal performance from any reference clip onto your static character portrait
From static portrait to dynamic performance in three simple steps — upload your character, supply reference movement, and render synchronized video.
Engineered with Kuaishou's Omni One architecture to decouple actor choreography from character identity with realistic physics, lighting, and lip-sync.
Explore how complex choreography, dynamic lighting shifts, and vocal performances transfer across diverse artistic styles and character portraits.
Designed for filmmakers, fashion directors, virtual creators, and social teams translating human physical performance to stylized avatars.
Standard mode generates 720p video optimized for rapid iteration and testing, while Pro mode renders full 1080p footage with fine-grained skin textures and fabric weaves. Pro mode requires more compute credits and processing time, making it ideal for final client deliverables, whereas Standard provides a cost-effective path for prototyping choreography.
Use the Pro tier for close-up acting and pristine vocal lip-sync where facial micro-expressions and eye reflections are paramount. Standard tier works well for full-body dance drafts and broad gestures, but Pro delivers the spatial resolution needed to avoid facial drift and preserve fine mouth contours.
Video orientation aligns the character's body directly with the reference clip and supports clips up to 30 seconds for dynamic actions like dancing. Image orientation locks the character's initial framing to the source photo and supports up to 10 seconds, making it better suited for subtle acting and camera-following moves.
This model excels at single-person choreography, dance routines, facial micro-expressions, and speech transfer where the actor remains clearly visible throughout the shot. Its spatiotemporal mapping accurately computes 3D volume, dynamic lighting changes, and realistic clothing movement without flat image warping.
Avoid reference videos with quick camera cuts, severe handheld shake, or multiple subjects in frame. The model requires a continuous 3 to 30 second shot with a single subject whose limbs remain visible, and struggles with non-humanoid or heavily occluded character designs.
Yes, keeping the original sound enabled automatically synchronizes the character's lip movements to speech or singing in the reference clip. If your source video contains distracting background noise or wind, disable the audio toggle and add clean soundtrack audio in post-production.
Transfer complex choreography, micro-expressions, and vocal performance from any reference clip onto your static character portrait