Reinforcement learning (RL) environments are simulated workplaces where agents learn by trial and error, replacing scraped text and chatbot ratings as the training ground for the next AI generation.
Google is in talks to pay more than $1.5 billion for Mechanize, a startup whose entire product is fake work. The company builds virtual work environments: simulated business software, multi-step tasks, and the tools agents need to practice until they can finish a real job.
The deal, reported last week by Business Insider, is not closed. If it lands, the Mechanize team would join Google and license its technology for model evaluation and development. The dollar figure and the buyer are the clearest sign yet that the next generation of AI will not be trained on the open internet. It will be trained on work.
Mechanize is selling into a category that labs now call "RL environments," short for reinforcement learning environments. The shorthand obscures a simple idea: instead of having a model read text and predict the next word, you put it inside a simulation, give it a goal, and let it learn by trying things and getting scored on the outcome. Think of it as the difference between reading every book on plumbing and actually fixing a sink under supervision.
The category is the training ground for agents: AI systems meant to do multi-step work, like filing an expense report, debugging a script, or triaging a customer queue, rather than answer a single question. Scale AI, a major supplier in this space, writes that the industry is moving from static datasets and human preference feedback to simulated environments where agents learn by trial and error.
The first LLM wave was trained on two kinds of data. The first was enormous text scraped from books and the open web. The second was feedback from contractors who rated the model's chatbot answers for helpfulness and harmlessness. That recipe produced assistants that could explain a tax form or draft an email but could not, on their own, reconcile a month of receipts, log into a vendor portal, and post a payment.
The bottleneck is not raw knowledge. It is the gap between "I know what this step looks like" and "I can chain twenty of those steps together without losing the thread." Reinforcement learning inside a simulated workplace is the bet that an agent will close that gap only by practicing multi-step work end to end, with a score attached to completion rather than to a single answer's tone.
Mechanize founder Tamay Besiroglu announced the company in April as a bet that "automated employees" are the next platform shift.
Mechanize is one path. Meta is taking a parallel one. According to the same Business Insider report, Meta has been collecting employees' keystrokes, mouse movements, clicks, and other screen activity, framing the goal as teaching AI how people actually use computers: the keyboard shortcuts, the navigation habits, the long-running workflows that no static dataset captures.
The two programs are not equivalent. Mechanize is building artificial offices. Meta is recording the real ones and using them to train models that will, presumably, operate inside offices like them. Together they sketch the same shift: the data that used to come from the public web is now coming from work, simulated or observed.
Two things are still unknown.
First, the public evidence that simulated work produces hour- or day-long task completion at scale is thin. TechCrunch reported last September that the bet on environments was real, but the field has not yet published benchmarks that show agents trained this way reliably out-performing the text-plus-feedback models on long, messy jobs. Mechanize's announcement drew sharp public backlash in Fortune over the "automated employees" framing.
Second, the consent question around recording real workers is unresolved. Meta's program, as described in the Business Insider report, sits at the intersection of employee data and training data without a settled public framework for either.
The next test will be whether the labs release benchmarks that test agents on multi-day work, not multi-step chat. Until then, the $1.5B is a bet, and the fake offices are the rehearsal room.