The Chinese lab pairs A1 with a W1 world model and says a single shared state replaces the field's old split between short stable clips and chained live avatar pipelines. All numbers are Vivix's own.
Vivix has opened internal testing for A1, a real-time interactive video model, alongside a companion world model called W1. The company says the pair run on a single "unified streaming architecture" that treats video frames, audio, gesture, and conversation history as one continuously updated state instead of separate clips or a chained pipeline (Vivix A1 report).
Prior real-time avatars sidestepped the problem by chaining ASR, a language model, text-to-speech, and a motion or video generator, with each handoff shedding information. Short clip-based generators can see the whole segment at once and so avoid drift, but they cannot run live. A1 folds every signal into a shared state that updates token by token, with separate "local" and "global" priors that keep short actions and long-term character identity on track (QbitAI recap).
Vivix says A1 has roughly 30 billion active parameters, a 300-millisecond streaming-scheduler response, an average 0.6-second delay from input to first visible frame, and throughput above 10,000 video tokens per second on one consumer-grade GPU running native NVFP4 inference. The company also calls it the "first" model to unify multimodal reference, real-time interaction, and streaming generation in one architecture (Vivix W1 report).
Those numbers and the "first" framing come from Vivix's own technical blog, as relayed by QbitAI, and have not been independently reproduced. Whether A1 holds up on a third party's GPU and against current live-avatar products from Kuaishou, ByteDance, or U.S. labs is the next test.