Pangram CEO Max Spero, who calls himself the 'slop janitor,' raised US$9mil this week, and independent work supports his AI detector's edge over rivals, but accuracy on real human writing is harder to verify.
More of the text people read online is now machine-made. A New York startup called Pangram said this week its detector can correctly flag AI-written text 9,999 times out of 10,000. The next question, for any reader, is who audits the auditor, and what happens when a real human gets mislabeled as a bot.
That question has stopped being academic. A Graphite analysis cited by The Star estimates roughly 50% of news articles and blogs on the internet are now AI-generated, a marketing-firm number rather than a peer-reviewed one, but a useful measure of the volume Pangram is selling against. "AI slop" is the industry's shorthand for the resulting flood: blog posts, product reviews, and SEO filler that an attentive reader can feel, even when no one can prove it.
Pangram is the most-funded name trying to clean it up. This week, Pangram closed a US$9 million (RM36.7 million) round led by Menlo Ventures, with ScOp Venture Capital, Haystack Ventures, and others participating. Max Spero sells a tool for telling machine text from human text, one paste at a time.
The product is simple by design. A user drops a paragraph into Pangram's interface and gets back one of three verdicts: AI, human, or AI-assisted. With the new round, the company released Pangram 4.0 and its first image-detection model. The headliner claim, attributed to the company, is that the model "can accurately identify 9,999 out of 10,000 times whether text was AI-generated," a self-reported number that has not been independently replicated on real journalism.
Independent academic work, though, gives the company's tools a real edge over its rivals. A 2025 Becker Friedman Institute audit at the University of Chicago put Pangram alongside OriginalityAI, GPTZero, and an open-source RoBERTa baseline, framing detector accuracy around the trade-off between false positives and false negatives. A 2026 peer-reviewed empirical study in Springer tested Pangram against GPTZero, Copyleaks, and Turnitin on 160 synthetic documents. Pangram was the only tool that did not significantly underestimate the share of AI text in advanced-generation cases.
Those are encouraging results, with one big caveat. Both studies ran on synthetic or non-newswriting corpora. Pangram's pitch to publishers is that it works on long-form journalism; the academic benchmarks do not directly test that case. The 99.99% headline number remains, in practice, a company number.
The harder question is the false-positive cost. A 2025 arXiv preprint by Russell, Karpinska, and Iyyer studied how well humans, including frequent ChatGPT writers, can detect LLM-generated text on their own. The answer was poorly. If trained readers miss most machine text, any detector that leans toward "AI" will catch more real prose than it should. That is a problem for the human freelancer whose submission gets flagged, the student whose essay is auto-rejected, and the publisher whose competitor's outlet is the one running the detector.
It is also a problem that vendors selling certainty have an obvious commercial reason to underplay. The more confidently a detector can name a culprit, the easier it is to sell. Pangram's "high-profile online publishers" customer list is not enumerated in the company's marketing; readers have to take the trust claim on the company's word. The auditors and the auditees are often the same business.
For now, the cleanest read on a 99.99% claim is the one the data already supports: Pangram is the best tool in a category that is not yet a science.