Hugging Face, a popular open source AI hub, hosts one click demo apps. An audit found 7 of the top 9 image editing demos strip clothing from real photos.
Hugging Face hosts open-source AI models. An audit of its most-popular image-editing demos found seven of the top nine are built to strip clothing from real photographs, and 73% of the prompts inside them were sexual. The platform disputes the framing; the mechanism the audit points to is the leaderboard that ranks those demos.
The audit comes from AI Forensics, a European nonprofit that studies platform abuse. Researchers manually tested Hugging Face "Spaces," zero-setup demo apps that anyone can run in a browser without writing code. They picked the nine top-ranked Spaces in the image-editing category and tried the obvious thing: feed in a clothed photo, see what comes back. Seven of the nine did the rest. "Nudify" workflows, with no adversarial tricks required, were the default behavior of the most-clicked tools.
The abuse overwhelmingly targets women. The prompts researchers watched flow through these Spaces are not generic; they name real people, often celebrities, often private women whose photos were scraped from social media. The images produced are nonconsensual intimate imagery, a category that is illegal to distribute in most US states and much of Europe, and that victims almost never fully scrub from the internet once it spreads.
Hugging Face does not train these models. The company hosts them and hosts the Spaces that wrap them in a one-click interface. That distinction is the platform's main defense, and it is partly accurate. The AI Forensics finding sits one layer up, at the discoverability surface rather than the model weights. The Spaces leaderboard is public, sortable, and updated in real time as users upvote. A new "nudify" app that lands on the leaderboard gets traffic before Hugging Face's trust-and-safety team can plausibly review it, and the traffic itself makes the app rank higher.
The platform's on-record response walks a careful line. After receiving a draft of the AI Forensics report before publication, Hugging Face says it ran the same tests and took action against Spaces that were intentionally misrepresenting their purposes. A spokesperson said the company is "always grateful for work that provides information on [nonconsensual intimate imagery] issues, but find misconceptions in this report make it harder to act upon." The company denies a "systemic lack of moderation" and points out that several of the artifacts the audit flagged were removed under existing procedures before the report went live.
Hugging Face's rebuttal is a real counter. AI Forensics' claim is a snapshot of what the leaderboard surfaces to a casual visitor on the day the researchers ran the study, not a claim that every flagged Space persists indefinitely. Hugging Face's claim is that it acts on individual Spaces after they are reported or noticed. Both hold at the same time, and the gap is the point. The mechanism that produces the next "nudify" Space is the same mechanism that lets Hugging Face's moderation team keep catching the last one.
The pattern is not unique to Hugging Face. WIRED and Engadget have both noted the parallel to X, where xAI's Grok image generator has become a destination for the same kind of content. The AI Forensics report assigns shared responsibility across model developers, hosting platforms, and downstream apps. Hugging Face is the case study here because the Spaces leaderboard makes the failure mode unusually legible, not because it is the only place it happens.
The watch item is whether Hugging Face changes the discoverability layer or only the moderation layer. Removing specific Spaces is necessary; ranking them by upvotes while doing so is what produces the surface the audit measured. AI Forensics recommends treating Spaces the way app stores treat apps: pre-review, not reactive takedown. Hugging Face has not committed to that, and the leaderboard is still live.