MiniMax's H3 undercuts the main closed API alternative on per second price at 2K and is positioned to become the first open weights model to top the Artificial Analysis video editing leaderboard (a third party AI benchmarking site), with publicly
A new video generation model now handles editing, typography, sound, and pacing inside a single text prompt. MiniMax's H3, released this week, ships with native stereo audio at up to 2K resolution, charges less than one-third per second at 2K versus the main closed-API competitor, and is the first open-weights model to top the Artificial Analysis video editing leaderboard, according to the company's launch blog and a QbitAI report.
The model is a single pretraining pass that absorbs the production floor. H3 is described by MiniMax as a general-purpose omni-modal model: it jointly understands text, image, video, and audio, and generates finished video with native stereo audio in a single call, rather than chaining a text-to-image model, an image-to-video model, a voice synthesizer, and a music generator. The release ships four architectural pieces that make the unification work. Contextual Omni Representation lets a single representation carry the prompt across modalities. H3-VAE is the autoencoder for video. H3-Omni Transformer decodes the joint stream. In-Context Regeneration iterates on a clip without re-running the entire generation. Together they absorb tasks that used to live in separate tools: text-to-image, image-to-video, first-and-last-frame, subject and motion reference, style reference, voice, sound effects, and music, all under the same model, per MiniMax's technical documentation.
At the release price, 2K output costs less than one-third per second of what MiniMax calls the mainstream closed-API alternative, and 768p output costs less than half of mainstream 720p, per the launch blog. The blog does not publish an absolute per-second dollar figure, so the comparison is relative and sourced from MiniMax's own framing; an independent API pricing page was not available at the time of writing. A finished, edited clip at default 2K with voice, sound, and pacing, from one prompt, is the capability on the same page.
Open weights would also break the lock-in pattern of the previous wave. MiniMax has announced an open-weight release "in the coming days, subject to applicable laws and regulations," with hardware compatibility designed in from the start. If the weights ship on schedule, the model that tops the Artificial Analysis video editing leaderboard would be the first top-of-leaderboard video model that is auditable, modifiable, and not stranded behind a single vendor's API. The previous video-model wave was dominated by closed APIs at the top of the leaderboards, with open-weights entries trailing well behind on editing tasks.
A captioning pipeline distills roughly 100,000 inference tokens per sample down to about 4,000 tokens on average, per the launch blog. Lower token counts per generation mean lower compute per second of finished video, which is how the per-second price can fall without the model getting cheaper to train.
The use cases MiniMax calls out are the categories where one-prompt finished video is most directly substitutable for a small production team: film opening titles, product website hero clips, animated posters, and advertising and e-commerce creatives. A solo creator, a small marketer, or an educator can plausibly produce a finished, edited clip without hiring an editor or learning Premiere. The trade-offs are still real. Outputs cap at 15 seconds, and a 15-second clip is not a short film. A Hitchcock camera-movement prompt changes the framing, a character-sing reference prompt carries the voice across shots, and a vocals-matching-audio reference prompt locks the lip sync to a music track. Those are moves a human editor used to perform, now stated in language.
The Artificial Analysis ranking claim is reported by QbitAI, which characterizes H3 as the only open-weights entry at the top; an independent leaderboard pull was not completed in time for this piece, so the claim should be treated as a third-party leaderboard assertion rather than a verified benchmark. The open-weight release is announced but not yet shipped. The "coming days" wording is forward-looking, and a slip in the release date would push the portability argument from "capability is portable" to "capability is promised to be portable." The production-floor collapse stands either way. The open-weights lock-in dimension does not.
If the weights ship on the announced timeline, the first top-of-leaderboard open-weights video model becomes a real baseline for the next wave of video fine-tunes, distillations, and on-prem deployments. If they slip, the leaderboard claim still holds, and the portability story retreats to a roadmap.