Their pitch: real time text to speech running on Tenstorrent's RISC V servers — an open source chip architecture competing with Nvidia — with usage based pricing and no upfront commitment.
Tenstorrent and Smallest.ai said on October 5, 2026 that Smallest.ai's real-time text-to-speech model will run on Tenstorrent's open-architecture RISC-V servers — another attempt to put enterprise voice AI on non-Nvidia hardware. The pair are pitching banks, telecoms, and hospitals that need customer conversations to stay on their own premises.
The companies describe the result as "production-grade" and "available today," priced by usage. Smallest.ai's Lightning V2 model runs natively on Tenstorrent's Galaxy Blackhole servers, which combine Tensix AI cores with large on-chip memory and a fast interconnect. Smallest.ai claims the setup cuts inference cost by about 4x against a comparable Nvidia L40S system. A co-authored arxiv preprint backs the claim, but the comparison is narrow: streaming text-to-speech is a memory-bound workload, and the paper has no independent benchmarker.
No enterprise customer is named in the announcement, so the on-prem sovereignty story is still aspirational. Smallest.ai raised $13M in July 2026 to push enterprise voice; Tenstorrent has raised more than $1B from backers including Samsung, LG, and Hyundai.