Synthetic data — text and code written by other AI — is the proposed fuel. A Northeastern researcher is testing whether it can teach real reasoning or just its shape.
AI may be running out of human-generated text to learn from. The proposed fix is "synthetic data," text and code written by other AI models, and the small teams of elite human experts who can produce that material are vanishingly rare. A Northeastern University graduate student is betting his research career on whether that material can teach reasoning, or only its shape.
Harsh Raj, a master's student in Northeastern's Khoury College of Computer Sciences, puts the problem this way. "It's very hard to collect data which trains the model to be better than humans, because there are very few humans who can create that data," Raj told Northeastern Global News. His method: have small teams of domain specialists generate text and code, then test where cloud-hosted frontier LLMs, or large language models, fail on it.
The work sits inside professor David Bau's lab. Raj is finishing a research-engineer co-op at Scale AI, after an earlier stint at Bespoke Labs.
The unresolved question, which Raj himself raises, is whether synthetic data teaches models to reason or only to mimic the surface of expert work.