Artificial intelligence researchers are confronting a looming challenge: what happens when large language models finally consume the full breadth of human knowledge? At Northeastern University, graduate student Harsh Raj is among those charting the next frontier, exploring how synthetic data might one day help AI systems think with the sophistication of Nobel‑level scientists.
When Human Knowledge Runs Out
As AI systems grow more powerful, experts predict a future in which algorithms outpace the limits of general human knowledge. Raj, a graduate student in Northeastern’s Khoury College of Computer Sciences, is already working inside that future. His research examines how to supply LLMs with information that only the world’s most elite thinkers possess.
“It’s very hard to collect data which trains the model to be better than humans, because there are very few humans who can create that data,” Raj told Northeastern Global News.
The dilemma is stark: if only a handful of people can produce the insights needed to push AI beyond human capability, how can researchers gather enough of that data to train the next generation of models?

Synthetic Data as a New Source of Intelligence
Raj’s work spans two high‑profile co‑ops, first at Bespoke Labs and now at Scale AI. At Bespoke Labs, he helped identify failure points in cloud‑hosted AI models and created synthetic data—artificially generated information designed to mimic the quality of human‑produced writing or code.
Creating synthetic data that matches human nuance is a formidable challenge. Computing power is easy to acquire, Raj said, but finding or generating data capable of supporting super‑human knowledge capabilities remains largely uncharted territory. His work aims to close that gap by designing datasets that help LLMs reason more effectively and avoid common failure modes.
At Scale AI, Raj now publishes research on failure taxonomies, cataloging the many ways an LLM can fall short of user expectations. His role involves probing the weaknesses of cutting‑edge models, reading scientific literature, attending conferences, and collaborating with industry experts to propose targeted improvements.
Automating Discovery
As AI labs push toward the limits of human knowledge, Raj believes the next step is training models on the collective expertise of entire scientific teams—not just individual experts. The goal is to build systems capable of performing not only the work of a skilled coder, but the output of a full startup staff.
You want to create data which is automating science, automating discovery, he said, noting that only a tiny fraction of people have experience operating at that level.
A Researcher at the Edge of AI’s Future
Raj plans to continue exploring the boundaries of AI knowledge after graduation, armed with the experience gained through his Northeastern co‑ops. Bau believes this work has given him confidence to contribute meaningfully to fundamental research.
Source: Northeastern Global News


