HomeArticle

Who is creating, selling and grading test questions for large models? A discussion on the wild unregulated growth of the AI data industry

硅谷1012026-10-08 10:09
What exactly is this industry selling?

With the improvement of AI model capabilities and increasingly complex training demands, a group of new companies providing training data for model developers are growing rapidly: AfterQuery announced a $300 million valuation when it closed its Series A funding round in April this year, and by September, according to media reports, its new funding round has pushed its valuation to $3.2 billion. Another company providing training data and agent environments, Bespoke Labs, also announced in July this year that its combined seed and Series A financing has reached $40 million.

Despite the rising valuations, the AI data industry remains extremely opaque to the public. What exactly are these companies selling?

In the past, when we mentioned data labeling, we easily associated it with tagging images and scoring model responses. But as AI evolves from answering questions to executing tasks, the deliverables of data companies are also changing: From scoring criteria written by experts, to reinforcement learning environments that integrate tasks, tools and verification mechanisms. They need to find industry-savvy people, obtain real industry materials, and then translate professional knowledge and work experience into training signals that models can learn from.

At the same time, new AI evaluation benchmarks are emerging in endlessly. When the scores on the leaderboards rise, which specific capabilities are actually improved? When we produce targeted data for evaluation, when can it improve the real user experience, and when will it turn into mere score-chasing? For researchers who train models every day, what exactly is the data they need most?

In this episode of the Silicon Valley 101 podcast, we invite Yunzhong He, who leads post-training and evaluation research at Scale AI, and Yiyou Sun, postdoctoral researcher at UC Berkeley and researcher of the Agents’ Last Exam project, to break down this low-transparency, fast-changing business from the perspectives of both industry and academia.

We talked about the different backgrounds of data companies including Scale, Mercor and Surge, and also discussed the opportunities for small teams in synthetic environments and vertical domains. From purchasing a game source code, to verifying tasks submitted by experts, to designing incentive mechanisms that encourage people to share their experience, the difficulties of high-quality data are often hidden beyond the leaderboards.

Below are the highlights of this conversation:

01

Who is running the AI data business?

Yiwen: There are all kinds of data companies on the market now, how should we categorize them?

Yunzhong He: There are several major players, some rising stars, and even small workshops with only a few people. The most interesting point in my opinion is that they may have completely different origins. For example, some started out in recruitment, like Mercor and Handshake; Scale AI is a traditional data vendor that grew out of a crowdsourcing network; some new players focus on synthetic data; the latest entrants are even companies from certain industries that have pivoted to the data business. It is worth mentioning embodied AI: in the current industry landscape, many model companies have not made profits yet, but they have invested huge amounts of energy in data construction, so they naturally started the data business as a side line. Many people from all walks of life have also discovered the value of their own data, and then pivoted to set up data companies. I think these are roughly the several categories.

Several major players, such as Mercor, Surge AI and Scale AI, are basically trying to cover everything. Besides, some rising stars perform well in specific fields, like BigCode or Snorkel AI, which are more research and engineering oriented. I guess they mainly produce environments for code generation or tool calling, and have caught the industry trend this year.

In addition, I have also seen some small companies that solve problems in specific vertical fields, and some of them are not even doing labeled data, but act as data brokers. For example, a person from Hollywood who holds a huge amount of copyrighted data and knows how to sell content from different fields. The latest emerging trend is vertical domain specialization, where teams composed of industry experts build extremely high-quality solutions for one specific field.

I think there are still plenty of opportunities. Large data vendors have their own advantages, as they have more resources when entering a new field. I believe the entire market pie is getting bigger and bigger.

Yiwen: You mentioned that companies like Mercor are so-called recruitment companies, what kind of business model is recruitment?

Yunzhong He: Mercor started as a recruitment company. It has its own automated AI interview product, whose core value proposition is that it can find a sufficient number of experts in a relatively short time, and it has built such an expert network. I think Surge AI does more of this work in-house. There are tradeoffs that basically every data vendor has to face: if you train your own full-time staff, their capabilities are definitely stronger, but the flexibility is lower, and the production capacity may be worse; if you rely more on a crowdsourcing network, the flexibility is higher, but it is harder to guarantee the quality of each individual, which involves the problem of quality inspection. The advantage of companies like Scale AI is that we have accumulated experience in quality inspection, not only on data, but also on personnel management.

02

How is the AI data industry shifting?

Yiwen: In fact, we are focusing on this topic because roughly from 2024 to now, the demand for data has undergone a major change. Before 2024, when we talked about data, most people thought of manual labeling, which is relatively basic data, and we needed a huge amount of data for pre-training to train models. But since 2024, as Mercor mentioned before, the rise of Deep Search has led the entire industry to shift to so-called Agentic data.

We need rubric as you just mentioned, including RL environments, and now more real workflow data. Can you sort out these concepts for us first? Let's talk about what changes the industry has gone through from crowdsourced data, that is, manually labeled data, to agentic data?

Yunzhong He: Right, first of all, pre-training is still very important. But indeed, from the interest of people around me, everyone has experienced a shift: people feel that pre-training is a bit boring now. Back then, the o-series models came out, everyone was working on reasoning, and found that RL worked, so everyone's interest immediately shifted to RL.

One of the most interesting problems in RL is that I need a verifiable signal. Verifiable signals are very simple in fields like code or mathematics, especially the earliest reasoning models that were built on mathematical problems: we can easily know if a multiple-choice question is answered correctly or incorrectly, or if a calculated number is right. But then people started to think: now RL works for mathematics, what about other fields? For example, in healthcare, finance, or more basic instruction following, can we do RL there? The problem here is that there are objective standard answers, unlike in the previous RLHF (Reinforcement Learning from Human Feedback) stage, where the reward model basically has no standard answers, and only learns people's preferences.

There are standard answers here, but the standard answers are relatively unstructured. It's like answering a history exam question: there is a standard answer, but everyone writes it differently, so it's very hard to do pattern matching to check if options A, B, C, D are all correct. That's when the rubric approach comes in, which is equivalent to having a teaching assistant grade your paper according to the scoring rules written by the professor.

Because the scoring rules are pre-written by the professor, the TA doesn't need to be particularly smart. This is a core concept in large model data, called Weak-to-strong supervision, can we use weak models to supervise strong models. The premise for this to work is that an expert has written down all the knowledge in his mind, so that the weak model can grade the papers according to that knowledge. That's what rubric data is.

This approach has been popular for a long time, and it is still very important now. But the easily accessible rubrics in various expert domains have been quickly collected almost completely. The rise of companies like Mercor or Surge AI is actually closely related to this expert rubric segment. But moving forward, people need to get more hands-on. For example, current code models, or the ALE project that Yiyou is working on, even though it is not focused on code, it requires hands-on operations to solve all kinds of problems.

At this point, it is no longer just about verifying the output of a chat. We need to work in an agent environment, which may require providing some tools, and more complex verification methods. It still includes rubrics, but we may need to check whether variables in this environment are executed, whether the database is written to or deleted. Even if we are generating documents that contain images, we still need rubrics, but we need an agent to perform the check, which makes the scenarios much more diverse.

A major direction that people are pursuing now is environments, which are no longer just pre-defined scoring rules, but include a complete set of definitions of "what I need to do" and "in what environment I do it", including the required tools, the sandbox to execute tasks, and a set of verification methods that are closely related to the task. The verification methods are very diverse, such as Programmatic check, rubric, and preference selection where we let people choose which output they prefer.

There is no broad consensus in this field yet. It feels a bit like the rubric era back then, where every company wants different things. For example, Terminal-Bench does not use rubrics at all, it is purely programmatic. There are different schools of thought here, which is very interesting.

Yiwen: Yiyou, would you like to use the specific example of ALE to talk about what elements are needed? (Note: ALE, Agents' Last Exam, is jointly initiated and constructed by UC Berkeley RDI and more than 300 industry experts, aiming to evaluate the ability of AI agents to complete real professional tasks. It allows experts to submit projects they have done before, and equips tasks with execution environments and verification mechanisms to test whether agents can complete long-duration, economically valuable work)

Yiyou Sun: Actually my view is a little opposite to Yunzhong's. Yunzhong thinks that rubric has almost reached its limit now, but my view is that the rubric segment, especially the subjective judgment questions from experts, is far from being fully explored. I can start with the original intention of building ALE. The original intention of ALE is the same as what Yunzhong said: how can we extend the rapid development trend of vibe coding to "vibe everything".

Why did vibe coding develop so rapidly back then? Because there are a huge number of benchmarks in the coding field, and coding is relatively easy to verify, it is a deterministic rubric. People are eager to get high scores on these leaderboards, so the coding field developed extremely rapidly. Using the same logic, if we can build various evaluation systems in other fields, especially in engineering fields that support objective scoring, even if people chase scores on these leaderboards, as long as the scores are meaningful and the questions are valuable, the improvement of model capabilities will bring tangible benefits to people's daily work.

That's exactly our original intention of building ALE. Of course, there are countless different professions in the world, so in the first version, we only made a first attempt with relatively wide coverage. We will continue to work hard to expand the coverage, and hopefully one day it can become the real "Last Exam". When you can solve this benchmark, you can basically solve the real valuable problems in certain fields of human society.

According to the official ALE website, its evaluation covers 55 non-manual professions Image source: ALE Blog

Yunzhong He: I agree with Yiyou's point of view, training in various fields is far from reaching its end, many things have not been represented through the rubric approach, so there is still huge space here. What I just said about rubric reaching the end is more about the industry trend: the static single-turn or multi-turn chat that outputs a text and then uses rubric to check, which is purely knowledge-based. It's not even the end, because different labs have different levels of capabilities.

The big trend this year is focusing on coding, and then coding is gradually shifting to knowledge work, where people need to get hands-on. At this time, the demand for text rubric is a little lower, but it is still very large, and there are still many unsolved problems.

03

Where is the boundary between doing evaluation and selling data?

Yiwen: I want to trace back the development of evaluation in recent years and its relationship with selling data. People think that benchmarking and selling data are two different things that cannot be confused. Scale AI previously launched Humanity's Last Exam, the HLE benchmark. Can you tell us about this benchmark and how you think it is connected to data?

Humanity's Last Exam score curve Image source: lastexam.ai

Yunzhong He: That's a good question. In principle, an evaluation reflects the gap between the model's current state and the ideal state we want. It is more of a scientific research benchmark, showing where we want the model to reach and where it is currently lacking. But the connection between this and selling data is also very interesting. Because very often, the data that improves the model's capabilities is closely related to the evaluation method of the leaderboard. The hard part here is to find a balance. If I produce a large amount of data that is almost identical to the benchmark, then the model is basically doing Bench-maxxing.

But large model training has various dimensions. If I want to improve a certain type of capability, I need to focus on a specific dimension, which is definitely correlated with the benchmark. This is the balance that everyone needs to find. Because people's expectations for AGI are too high, they think the model can do everything that humans can do, so they can construct a benchmark for various fields, and then find that the model still has a gap with the ideal state.

Usually, the people who construct the benchmark have a deep understanding of that specific field, which naturally gives them a certain degree of authority to sell data, and people will trust the data they hold. That's how this business comes into being. Now ALE is also a very popular benchmark, and I believe many people have asked Yiyou if he sells data, or VCs may want to invest in him, and gradually this turns into a business.

I think it has both pros and cons. This ecosystem is still developing wildly, and we really need to carefully weigh how much of it is for scientific research and how much is for Bench-maxxing. But there are indeed a lot of profit incentives, so it is a very prosperous market now.

Yiyou Sun: Right, I want to give everyone a warning here. As Yunzhong mentioned, people who build benchmarks now receive all kinds of temptations to sell their benchmark data. I even heard privately that some data companies have crossed this red line, they sell the data they use for evaluation. This is a line that no one should ever cross, you must not let your business pollute the authority of your benchmark, otherwise you will not only destroy your own benchmark, but also ruin other people's models.

I personally think that benchmark and data selling have different roles and different positions in this ecosystem. I believe benchmarks are for the future, and data products are for the present. Benchmarks are oriented to the future, we hope model companies can solve the problems that have not been solved yet. When we built ALE in January and February this year, we were especially afraid that the benchmark would be too difficult and lose its meaning, because the average score of most models at that time was only 10% to 20%, and most questions got zero points.

I was very worried that this evaluation would lose its meaning. But there is no need to worry about that, because benchmarks are designed for future models. Data is more about meeting the demands generated by the benchmark, so now people chasing scores on this leaderboard is