Bruce Schneier and Prabhakar Raghavan pitch a 'Genie coefficient' for AI agents: a proposed score for the gap between what you ask and what you actually meant.
Ask an AI agent for coffee and the system may buy you a coffee plantation, or schedule a cup for delivery in three weeks. Both outcomes satisfy a literal reading of the request. Neither is what most people mean when they ask for coffee.
Bruce Schneier and Prabhakar Raghavan open an IEEE Spectrum op-ed with that example to argue that AI leaderboards are missing an evaluation layer. The piece pitches a proposed metric for the gap between what a user asks an AI to do and the unspoken assumptions the user has about how the AI should do it. They call that gap the Genie coefficient, a name borrowed from the storybook problem of wording a wish precisely enough to get what you actually want.
Today's benchmarks grade capability and treat intent as a constant. A model that completes a coding task, books a flight, or files an expense report is scored on whether the action happened, not on whether the action matched the user's underlying need. As agents move from chat windows into tools that browse, transact, and execute code, the difference between a successful task and a successful interpretation becomes the difference between a useful assistant and a costly mistake.
The structural case rests on a 1987 anchor. Schneier and Raghavan reach back to Terry Winograd and Fernando Flores, whose 1987 book Understanding Computers and Cognition argued that the meaning of any utterance is inseparable from the situation, prior conversation, and shared cultural context that surrounds it. A request to "get me coffee" assumes a person, a place, a price, and a time. A literal executor that ignores those assumptions is not malfunctioning. It is correctly following the only specification it has been given.
That shifts the failure away from the agent and onto the specification. Schneier and Raghavan write that AI agents "have enormous latitude to get it wrong" on underspecified requests and that they will "think outside the box" because they do not share the user's conception of where the box is. The Genie coefficient is the proposed score for how wide that latitude is in practice, on tasks a deployment would actually care about. A high coefficient would mean the agent completed the request in a way the user would not have wanted. A low coefficient would mean the agent either asked the right clarifying question or acted on the right default.
The proposal lands in a measurement gap capability leaderboards have not addressed. They track reasoning scores, code-completion rates, and tool-use success. None of them track whether a task scored as successful would have been called successful by the person who asked. An AI Weekly alert summarizing the proposal confirms the same authors and the same framing, and signals that the article is meant to start a measurement debate rather than ship a formula.
Schneier and Raghavan are candid about the limits of their own proposal. There is no published formula, no benchmark suite, and no leaderboard. The contribution is diagnostic: intent alignment is a separate axis from capability, and buyer, deployer, and regulator decisions made on capability scores alone are missing a layer the agents themselves cannot supply. Whether the field treats the Genie coefficient as the name for that layer, or builds a different one, the question the coefficient is meant to answer is the one today's leaderboards do not.
That question is the watch item. Capability scores will keep climbing through 2026 because the research path is well funded and the metrics are well defined. Intent-alignment scores do not exist as a public measurement category yet, and a metric that does not exist cannot be benchmarked against, audited, or compared across vendors. The next agent evaluation framework that names this layer will be the one that has to defend the term.