The next contest in AI video is no longer prettier frames. It is how much a single generation can hold: how many seconds, how many references, and how finely a creator can edit inside it without breaking the take. That envelope is now the axis on which Veo, Sora, Kling, Pika, and Runway are being measured against each other. ByteDance's Seedance 2.5 is the latest bid to widen it.
Per the ByteDance Seed blog, the model generates up to 30 seconds of audio-video in a single pass, accepts up to 30 images, 10 video clips, and 10 audio clips as references, and adds timestamp-level editing. All three moves push the same unit: the usable clip — a frame the ByteDance Seed blog does not explicitly make, treating instead each as a distinct product capability. The 30-second figure, the reference counts, and the editing precision are vendor-supplied and have not been independently verified.
Most readers will hear this as "AI video is finally filmmaking." Hold the read that the field has re-priced its deliverable instead. A single-pass 30-second clip is not a finished film; it is the new minimum economic unit. Whoever holds the longest coherent single generation, swallows the most references, and lets the creator slice by timestamp owns the next layer of competition. Rivals will ship their own envelopes within months. The bet is that frame-by-frame polish was the wrong frontier: a single generation's depth was.
Reported by Sky for Type0, from One-take Creation, Flexible Referencing: Introducing Seedance 2.5. Read the original: seed.bytedance.com