Vidu S2 actually lets you edit video in real-time
Most AI video models are still stuck in the "render and pray" cycle—you prompt, wait for the cloud to crunch, and hope the result isn't a fever dream. Vidu S2 is trying something different by splitting into two modes: S2-Avatar for real-time digital humans and S2-Editing for live visual tweaks. Both hit 720p, which is a step up from the blurry messes we usually get with "real-time" tech.
The S2-Avatar side is surprisingly decent at following physical cues. I'm talking about arm swings, twisting, and leg movements that don't immediately collapse into spaghetti. You can toss in a reference photo of a person or a pet and suddenly you're talking to them via mic and camera. The weird part? You can inject objects or change their clothes mid-stream.
I decided to test this by resurrecting Steve Jobs. I fed the model a full-body reference photo and a detailed personality profile—basically telling it to be obsessed with design and user experience. I asked the digital Steve if the iPhone Duo met his expectations. Instead of sounding like a generic chatbot, he actually stayed in character, telling me the phone was smoother than expected but the price was a deterrent. I even made him dance. Seeing a guy in a black turtleneck bust a move is the kind of absurdity that actually proves the model can handle complex torso and leg movements without losing the character's identity.
Then there is Vidu S2-Editing, which is essentially "live Photoshop" for video. You upload a reference image and can trigger four specific changes:
- Virtual Try-on: Swapping clothes instantly. I tried a blue POLO, a beige coat, and a leather jacket; the fabric physics actually looked realistic as the character moved.
- Background Swaps: Moving from a beach to a cafe in a blink.
- Style Transfer: Changing the aesthetic of the live feed.
- Character Replacement: Swapping yourself for someone else. I used a clear half-body photo to turn myself into a different persona, and the perspective alignment was tight enough that it didn't look like a cheap sticker slapped on the screen.
Technically, this isn't just a fancy filter. They're using a Backbone-Refiner architecture. The Backbone handles the skeleton and physics with low latency, while the Refiner cleans up the 720p textures and lighting in milliseconds. To stop the video from "drifting" or hallucinating into a blob after a few seconds, they're using something called SRF (Self-Replay Forcing) to give the model a memory of what it just generated.
The whole thing is coordinated by a VLM Agent that acts as the brain, making sure that if you suddenly add a coffee cup to the scene, it fits the lighting and gravity of the current frame.
We're moving away from the "finished clip" era toward a "continuous world" where the video doesn't stop just because the prompt ended. If this keeps up, by late 2026, we might stop waiting for render queues entirely and just direct AI scenes live. With ByteDance reportedly working on their own real-time spatial video model under Zhang Yiming's supervision, the race to turn video into an interactive experience is officially on.


Frustrated with the lag in Runway Gen-3. I'm skeptical this is actually "real-time" or just a fancy cached preview for 720p.