A French solo founder is betting a software layer called KIE can pull more speed from the GPUs enterprises already own, but the 3,000 tokens per second demo ran on a 2B model, not a production LLM.
A solo founder in Paris is pitching a counter-thesis to the build-new-silicon playbook that dominates AI infrastructure spending: the next jump in AI inference speed will come from software, not chips. Kog, his one-person French startup, sells that argument to enterprises that already own the GPUs and want more from them.
The product is a software layer called the Kog Inference Engine, or KIE. It targets the part of LLM execution called decoding, where a model writes its answer one token at a time. Decoding is where latency lives, and where faster answers cost more compute. KIE runs on standard data-center cards: AMD's MI300X and NVIDIA's H200, the kind of hardware enterprises have already bought and depreciated, the same hardware that the rest of the inference-speed market is trying to displace.
The May Hacker News preview put the bet in front of readers. The post landed on the front page and, by the company's count, generated 200 tangible business leads, a real pipeline signal for an early-stage AI infrastructure shop. CEO Gaël Delalleau told TechCrunch that software engineering is the expected first use case, with design partners routing KIE through prompt-to-game and prompt-to-app workflows where each second of latency hits revenue.
The benchmark that anchors the story is a 3,000 tokens-per-second single-request result on a small, purpose-built model called Laneformer 2B, a roughly 2-billion-parameter network that Kog has now open-sourced. That is the number the company is showing on stage, in pitch decks, and in the Hacker News thread. It is also the line between what Kog has demonstrated and what it has not.
Kog's larger promise is 30x faster LLM inference, but the demo ran on a 2-billion-parameter model designed for the test, not on a production-scale LLM. Kog itself says the leap from the 2B demo to full-size models is the unfinished part. The company is open about that. After the May preview, prospective customers told the team they were not prepared to fine-tune small models for the speedup, so Kog has now refocused entirely on accelerating larger models. The 2B demo is independently inspectable on GitHub. The large-model promise is not, and the founder's own confidence is on the record alongside the question of whether the trick survives production-scale models.
Inference speed is now a paid product dimension. Cerebras Systems, a purpose-built AI chip company, went public in May 2026 on exactly that thesis: faster answers are a category, not a feature. Anthropic monetizes the same axis through Claude Fast Mode, charging for lower latency as a product tier. Both are pricing the latency of a token as a real line item in enterprise AI budgets.
Kog's bet is that the GPUs enterprises already own have more decoding throughput left in them. If KIE delivers at production scale, the math changes for anyone deciding whether to buy the next generation of data-center accelerators. If it does not, Kog joins the queue of software shops arguing a chip-shaped problem has a code-shaped answer, and the market will price that quickly.
The design partners are what grounds the speed claim. Software engineering teams waiting hours for code-assistant output are a real workflow, not a synthetic test. Prompt-to-game and prompt-to-app generations, where the user is watching a screen and the model is generating interactively, put latency on the revenue line. Those are the customers the 200 leads pointed at, and the use cases the 30x promise has to clear before the round is closed.
Cerebras priced the inference-speed thesis at IPO. Anthropic is already selling latency as a product dimension. The open-source Laneformer 2B is the part of the pitch anyone can verify today. The part that will settle the bet runs on full-size LLMs, on production workloads, on the GPUs enterprises have already paid for, and on a result Kog has not yet shown.