Tencent's publicly downloadable Hy3 shipped on July 6, 2026. The same day, Simon Willison ran it through his pelican riding a bicycle prompt, and a folk benchmark for publicly downloadable AI went mainstream.
When Simon Willison asked Tencent's new open-weight Hy3 model to draw an SVG of a pelican riding a bicycle, the test that mattered was not whether the bird pedaled. It was whether it looked like a pelican at all. That question, known in AI circles as the pelican test, has become a quiet folk benchmark for which open-weight (publicly downloadable) models are actually getting better.
Tencent's Hy team released Hy3 on July 6, 2026: a 295-billion-parameter Mixture-of-Experts model (21 billion parameters active per query), Apache 2.0 licensed, with full weights on Hugging Face. The same day, independent AI writer and developer Simon Willison ran it through his pelican prompt and published the result. Willison has run versions of the prompt before, and it has become a community shortcut for the messy, visual work that benchmark leaderboards don't capture.
VentureBeat framed Hy3 as beating GLM 5.2 at roughly half the parameter count, except on coding. BenchLM and Artificial Analysis tell a similar story, but the pelican answers the leaderboards' blind spot: which model would a working developer actually deploy?
MarkTechPost's coverage of the launch and a Hacker News thread on the release show developers using the pelican as a tiebreaker. Whether the test scales past Willison's blog is the open question.