Two August 6 announcements, AMD's Taalas deal and OpenAI's GPT 5.6 Luna default, point to one market: inference splitting between specialized silicon and tiered model SKUs.
On August 6, 2026, AMD said it would acquire Toronto-based Taalas, a startup that etches a model's weights directly into mask-ROM silicon, and OpenAI made a nano-class model called GPT-5.6 Luna the default for free ChatGPT users. The two announcements describe the same market: inference is splitting into two lanes, specialized silicon and tiered model SKUs (distinct priced product tiers), both aimed at real-time, high-volume workloads.
Taalas's approach needs a definition. Most AI accelerators read model weights from HBM, fast DRAM that sits next to the processor. Taalas instead hardwires the weights into the chip itself, the same way a circuit permanently stores program code in read-only memory. The result is a model-specific integrated circuit, an MSIC: useful for one model, useless for any other. The tradeoff is throughput. Taalas's first test chip, HC1 on TSMC's 6-nanometer process, served Meta's Llama 3.1 8B at 16,960 tokens per second, a figure the company framed as 48 times an Nvidia GPU and 8.5 times a Cerebras system at the time of its original announcement. Those numbers are Taalas's own claims, not independent benchmarks. The next chip, HC2, targets about 20 billion parameters per device in summer 2026, with the company describing a pipeline of roughly 50 chips serving a trillion-parameter model.
The bet is that the future of inference is a model burned into silicon, served at the lowest possible latency, with no DRAM bottleneck in the path.
That bet is not new. Nvidia acquired Groq's assets in December 2025 for roughly $20 billion, and now sells LPX rack systems that pair its GPUs with thousands of Groq-derived language processing units for premium inference. AMD's framing is the same. Taalas's silicon will be folded into the Helios rack-scale system alongside Instinct GPUs and EPYC CPUs, so customers can mix a hardwired model for the high-volume path and a general GPU for everything else. AMD's press release on the deal does not disclose price or timeline beyond customary closing conditions. CNBC reports Taalas has raised $219 million in venture funding since its 2023 founding, and quotes CEO Ljubisa Bajic saying a model the company has never seen can be "realized in hardware in only two months."
OpenAI's announcement, on the same day, is the demand-side mirror. GPT-5.6 Luna is now the default model for free ChatGPT users with unlimited text chats, and the API prices input at $0.20 per million tokens with cached input at $0.02 per million, backed by a 1,050,000-token context window. The shape is intentional. Luna is a nano-class model, small and cheap, aimed at high-volume traffic. Plus and Pro subscribers get the updated GPT-5.6 Sol with a slider that controls how long the model thinks before answering, a premium tier for queries where the user wants to pay for reasoning depth. The two models are not just price points. They are different products for different jobs.
OpenAI's own internal evaluation found that responses containing at least one factual error were about 62 percent less common with Luna and 68 percent less common with Sol than with the previous GPT-5.5 Instant, on financial, medical, and legal prompts. That figure is from OpenAI, not an outside benchmark, and the prompt set is not published.
Read together, the two announcements describe a market that is segmenting rather than consolidating. One lane: model-specific silicon, where the weights are part of the chip and the chip exists to serve one model at high throughput. Other lane: tiered model SKUs, where the same vendor sells a small cheap model for volume and a large expensive model for depth, and the user picks. Both lanes are aimed at the same workload, real-time inference at scale, where latency and cost per token matter more than raw parameter count.
The watch item is whether the two lanes stay separate. If model-specific chips become cheap enough to deploy per customer, the tiered SKU structure collapses into a single hardware-plus-model package, and the API price of a model like Luna matters less than the cost of the rack it runs on. If the tiered model structure holds, the specialized-silicon play stays a niche for the largest workloads and a handful of latency-sensitive applications. The buyers are different today. AMD's customers are infrastructure teams, OpenAI's are end users. The underlying bet is the same: inference has to be cheap, fast, and everywhere, and the companies that win it will not do so with a single product.