A San Sebastián startup raised $570M to scale a compression method it says can shrink large AI models enough to run on phones, vehicles, and satellites.
Multiverse Computing, a San Sebastián startup, raised $570 million on Monday to scale a compression technique the company says can shrink large AI models by 80 to 95 percent, small enough to run inference on phones, cars, and satellites instead of a data center. The Series C, reported by The Quantum Insider carrying the company's release, values Multiverse at $1.7 billion pre-money, a fivefold step-up from its prior round.
The mechanism is tensor networks, a branch of math borrowed from physics. Multiverse's own materials describe the technique as "quantum-inspired," a label the company has used to market the work since at least its 2024 product announcement. The underlying method is documented in an arXiv paper from 2024 by co-founder Román Orús and collaborators. There is no quantum hardware in the product. The compression runs on standard GPUs and CPUs, and the company ships it as a software layer called CompactifAI.
The $570 million round, roughly €500 million, is co-led by Forgepoint Capital International, BNPP Solar Impulse Venture Fund, and Bullhound Capital, with participation from HP Inc., Orange Ventures, Scania Invest, and Santander. The lead mix is itself the story: a European industrial LP stack (auto, telco, bank) sitting alongside a U.S. cybersecurity-focused firm, with the company saying the capital is earmarked for R&D, sovereign AI infrastructure, and international expansion.
CompactifAI ships in two surfaces, a deployment product for embedded use and an API for cloud-style integration, and the company describes an intelligent router that decides per task whether to run the model locally on the device or offload to a server. The customer list Multiverse cites, including Bosch, Telefónica, Allianz, and Bank of Canada, is the company's own. None of the named enterprises are quoted in the release, and the round does not disclose contract size or scope, so the logos should be read as stated deployments rather than independently confirmed adoption.
The largest AI infrastructure checks of the last two years have gone into data centers, GPU clusters, and power purchase agreements. A $570 million round into a software compression layer is a bet that some share of AI capability can be decoupled from the data-center buildout, that inference can move to the edge where power and latency budgets are tighter, and that the resulting savings compound at the scale of phones, vehicles, and industrial cameras. The bet is now large enough to take seriously.
The 80 to 95 percent compression figure is Multiverse's own. The Quantum Insider carried the company release without an independent benchmark, and the arXiv paper describes the tensor-network method at a methodological level rather than as a head-to-head leaderboard result. The compression math is real; the headline number has not been independently reproduced in this packet, and the standard caveat on compression claims is that the percentage moves with model, hardware target, and acceptable accuracy loss.
Multiverse's September 2025 deployment with SOHMA AI, a behavioral health platform for youth athletes running fully on consumer devices via the company's compressed models, is the most concrete public example of the on-device claim working in production. The next test is whether the new capital extends that pattern to the named enterprise customers, where Bosch, Telefónica, Allianz, and Bank of Canada each have device fleets large enough to make the inference-cost math matter. The round closed Monday; the proof is in the deployments that follow.