ITSMBench runs frontier AI on 89 real enterprise service desk tasks. The best model completes 50.56%, and the publisher sells competing help desk software.
The 50.56% overall completion rate on a new open-source benchmark from Atomicwork and New Measure breaks into two distinct scores: the models that find the right API are not the models that finish the job. That split is what an enterprise procurement team should grade against.
The benchmark, ITSMBench, recreates a corporate service desk from 42 mocked software systems, almost 1,800 database tables, and more than 2,000 REST endpoints. It runs frontier models from OpenAI, Anthropic, xAI, and GLM through 89 tasks typical of L2 and L3 support work: identity and access management, device troubleshooting, networking, security operations, infrastructure, and engineering escalations. ITSM is the discipline that runs the internal help desk: password resets, laptop failures, access requests, outage triage. Atomicwork has a direct stake in the result. The company sells enterprise IT service management software with an AI workforce, so it competes with the very agents ITSMBench is grading. The benchmark was built jointly with New Measure.
Grok 4.5 leads tool discovery with an 83.0% API-find rate, but ranks near the bottom on completing the task. Anthropic's Claude Opus 5 is the opposite: the strongest executor at 63.5% completion, but the weakest discoverer at 70.2%. OpenAI's GPT-5.6 Sol sits between them. GLM-5.2 is the cheapest of the group at $0.25 per trial, but trails both Opus 5 and GPT-5.6 Sol on execution.
The benchmark's public leaderboard uses a different metric. Its top Pass@1 entry, a single-attempt success rate, is Claude Opus 5 at 46.07%, at an average cost of $1.75 per trial and 21,400 output tokens. Grok 4.5 follows at 45.39% Pass@1, cheaper at $0.71 per trial. The two statistics measure different things: overall completion across multiple attempts versus success on a single try. Buyers reading both numbers should know which one they are being sold.
The full repository of environments, tasks, and evaluations is open source. Anyone can rerun the numbers. That is a meaningful difference from a glossy vendor demo, and it lets enterprises test their own candidates before signing a contract. It is not the same as independent academic research.
The closest comparable project, ITBench from IBM Research, launched its ITBench-AA agentic benchmark in May for SRE, FinOps, and CISO workflows and found all evaluated models below 50%. ITSMBench narrows the same lens onto the service desk specifically. The two are related but distinct.
Atomicwork's framing leans harder. Co-founder and CEO Vijay Rayapati said in the release that "frontier models are brilliant at writing code, but they are completely blind to the hidden security landmines inside enterprise workflows," and called ITSMBench proof that "blindly trusting an AI agent right now is an invitation for an enterprise security breach." That is a vendor talking point attached to a reproducible data point. The data is real; the rhetoric is his. Security operations is one of six task categories in ITSMBench, and the leaderboard does not yet break out a per-category security score.
The next test for ITSMBench is whether the open repo attracts outside runs on new models, or whether the leaderboard stays vendor-curated. The public GitHub repository accepts community task submissions. If a model release next quarter is graded on ITSMBench by someone other than its publisher, the 50.56% number will start to mean something the vendors cannot move on their own.