The OpenAI case study claims Circles' concierge lifted average revenue per user 22%, with 9% less churn and 65% autonomous resolution, in Singapore. None is independently verified.
An OpenAI case study about its customer Circles, a Singapore telco-software vendor, claims four headline results from the company's AI Concierge in Singapore: a 22% lift in ARPU (average revenue per user), a 9% drop in customer churn, a 65% autonomous resolution rate for support, and a 29% jump in developer efficiency. The numbers are all vendor-stated. The OpenAI case-study page that published them does not disclose methodology, measurement window, baseline, control group, or any third-party verification (OpenAI).
The mechanism that produced the figures matters as much as the figures themselves. Circles, founded in Singapore in 2014, sells digital-telco software to operators in 14 countries and also runs a consumer brand called Circles.Life in its home market. Its AI Concierge sits on top of OpenAI's API and routes customers through a multi-agent system the company calls CareX. Instead of the usual menu tree, the concierge pulls real-time billing and browsing context, then dispatches a request to a specialist agent for travel, billing, plan changes, or support. CareX, per Circles, handles 65% of those requests without handing them to a human (Circles).
The four metrics measure different things, and reading them as a single finding is what makes the case study easy to over-credit. ARPU and churn are business outcomes, revenue per customer and the rate at which they leave. The OpenAI page does not say whether the 22% figure is an A/B lift, a year-over-year comparison, or a campaign-period delta. Churn is similarly bracketed: 9% off what baseline, over what window, with what customer mix? The 65% autonomous resolution number depends entirely on what counts as resolved; the source does not give a threshold for handing a request to a human. The 29% developer-efficiency gain is attributed to OpenAI's Codex and is internal to Circles' engineering team. None of the four is a controlled experiment. None of the four has a peer-reviewed or third-party audit (OpenAI).
The same four numbers have already been redistributed as findings. PR Newswire carried the joint announcement, and other trade and financial outlets have syndicated the same claims, so trade readers now see them on the same footing as audited industry data. That redistribution is itself how vendor case studies become "evidence" in tech reporting: the case study lands, the wire repeats it, the trade publication rephrases it, and the methodology never appears. Circles says the deployment spans 14 countries and six continents; the OpenAI case study links the metrics to Singapore alone, so the geographic reach of the result is not the same as the geographic reach of the deployment (PR Newswire).
A reader who needs to act on this kind of announcement, to size a vendor, brief a procurement team, or write the next one up, can run it through a five-line filter before passing the numbers along. Ask for the methodology, including how ARPU and churn are computed and over what window. Ask for the baseline: the same metric on the same customer base before the concierge went live. Ask for a control, even a synthetic one, or a comparator from a peer deployment. Ask whether the 65% resolution rate has a published threshold for human handoff. Ask who audited the figures, or accept that nobody did. A vendor case study that answers all five is publishable. One that answers none is a press release dressed in percentages.
The OpenAI page and the Circles announcement are the first to fail the filter on the headline numbers. Until at least the ARPU and churn claims are pinned to a window and a baseline, the only thing this case study supports is that an OpenAI customer is running an AI concierge in production at non-trivial scale. The striking four-metric stack around it is claimed, not measured.