Interview: AI’s hunger games - A lucrative data market is exploding to feed insatiable LLMs
Nomad Data's CEO Brad Schneider spoke with Sharon Goldman of Venture Beat's The AI Beat on the exploding market for LLM training and inference data.
The Growing Appetite for Data
Last week, I wrote about Mark Zuckerberg’s comments about Meta’s AI strategy, which includes one special advantage: a massive, ever-growing internal dataset training its Llama models.
Zuckerberg boasted that on Facebook and Instagram there are hundreds of billions of publicly shared images and tens of billions of public videos.
But it turns out that the training data required for Meta, OpenAI, or Anthropic AI models is just the beginning of understanding how data functions as the diet that sustains today’s large language models.
When it comes to AI’s growing appetite for data, it is the ongoing inference required by every large company using LLM APIs that is turning AI models into the insatiable equivalent of the classic Hasbro Hungry Hungry Hippos game.
Highly-specific datasets are often needed for AI inference
“[Inference is] the bigger market, I don’t think people realize that,” said Brad Schneider, founder and CEO of Nomad Data. The New York City company, founded in 2020, has built its own LLMs to help match over 2,500 data vendors to data buyers, which includes an exploding number of companies needing often obscure, highly-specific datasets for their own LLM inference use cases.
Rather than serving as a data broker, Nomad offers data discovery—so companies can, in natural language, search for specific types of data.
Finding the right AI data ‘food’
Certainly, training data is important, but Schneider pointed out that even if you have perfect data to train the model, it is trained once—but inference can happen thousands of times every minute.
The problem, however, has always been to find just the right data “food.” For the typical large enterprise company, starting with internal data will be key, Schneider said.
Media companies and data selling
We’ll all read about large media companies negotiating to license their data to OpenAI and other LLM companies. But Schneider says that Nomad Data is also signing up media companies and other corporations as data vendors.
The hunger games of LLM data
The bottom line is that the LLM hunger supply chain is basically a never-ending circle. Schneider explained that Nomad Data uses LLMs to find new data vendors. Once those vendors are onboard, the company uses LLMs to help people find the data they are looking for.
AI training data, he reiterated, is an immeasurably small piece of this market. The most exciting part is LLM inference, as well as customized training.