The National Institute of Standards and Technology's AI Technology Evaluation (AITE) program grades AI on data it cannot train against, starting with image tasks in August 2026.
NIST's new AI testbed gives models data they have never seen and does not let that data leave the room. The agency launched its AI Technology Evaluation (AITE) program on Monday as a voluntary safety testing vehicle that runs on an isolated environment, with the first round of evaluations set to start in August 2026.
At launch, the testbed covers image analysis only. Three domains anchor the first wave: quantum science, genomics, and public safety. Data providers must supply original datasets that are not publicly available, plus a meaningful task. Model developers then submit their models to be tested against that data. The data is not intended to serve as training material; it is meant to be a blind, one-shot read on capability.
NIST's press materials frame AITE as a step toward a universal rubric for AI capability, a shared federal yardstick rather than vendor-controlled benchmarks. The program sits inside the Trump administration's voluntary submission strategy and follows a May 2026 Commerce Department deal with Google DeepMind, Microsoft, and xAI routed through the Center for AI Standards and Innovation. Independent federal coverage confirms the same launch.
What remains unclear: which organizations will actually submit data and models, how scores will be released, and whether image-only analysis will expand to text or agentic tasks. AITE's TEVV program page lists the goals; the August start will show whether the field shows up.