Anthropic Labs Revealed: 20-person team, two-week period to decide retention or departure, 80% of ideas are allowed to fail
A team of around 20 people has built Claude Code, MCP and Claude Design.
But in the eyes of Ben Mann, co-founder of Anthropic and head of Labs, most of the team's attempts should have been allowed to fail in the first place.
According to a report by Business Insider, Anthropic's internal Labs team reviews its product prototypes on a roughly two-week cycle: should they move forward, or adjust their direction? Ideas that lack prospects will be terminated, or merged into other projects. Products that truly work will "graduate" from Labs and be handed over to independent teams for continued development.
The success rate of ideas Mann cited is only 20% to 30%. Some explorations that failed to take shape independently will also be absorbed into other prototypes or products. This figure, paired with the success of Claude Code, forms the most distinct feature of Labs: accepting that a large number of attempts will not make it to the end, and reserving resources for a small number of directions worthy of continued investment.
As Anthropic pushes forward with commercialization, the importance of this mechanism is on the rise. According to Business Insider's reporting, the company is preparing for an IPO. How to turn cutting-edge model capabilities into products that users are willing to use continuously will be a question that Labs needs to continue to answer.
How Labs Operates Like a Startup Incubator
Mann recalled in an interview that a few years ago, he even had to convince the company to start releasing products. At that time, a question that needed to be discussed was: could the company rely on charitable funding to continuously advance advanced AI research.
In 2024, he created Labs to set aside space for free exploration of new product ideas. Since then, this team has successively incubated Claude Code, the Model Context Protocol (MCP) that connects AI agents with external data and tools, and Claude Design launched this year. Mike Krieger, co-founder of Instagram, and other people also participated in it.
Ben Mann
Mann describes Labs as a true "startup incubator" in every sense.
This organizational form has a clear reference. He mentioned Bell Labs, and also joined Area 120, Google's internal incubation project, in 2018. Google X also provides an idea: use a dedicated team to explore directions with high uncertainty, and then let mature projects gradually become independent.
What makes Labs special is more reflected in its distance from model research. The capabilities of cutting-edge models are constantly changing. A product that is not easy enough to use today may cross the practical threshold after the next round of model upgrades. If the product team can know in advance which capabilities are being improved, they will have the opportunity to judge earlier what is worth doing and what can be tried.
Claude Code was born under exactly these conditions.
The Birth of Claude Code
According to Mann, as early as before the relevant development was launched at the end of 2024, researchers had revealed to his team that the new model showed potential in agent programming. This gave an important signal: building products around programming can try to let the model take on more practical operations.
Mann assigned this direction to Boris Cherny, who had just joined the company. Cherny initially proposed a code analysis tool, but Mann thought the idea was not large enough. He said in the interview that new employees in particular need to significantly raise their goals, because there is still much unknown about what agents can actually do.
Boris Cherny
This sentence corresponds to a practical problem in AI product development: if you always design products based on the capabilities you are already familiar with, you may underestimate the tasks that the model can take on next.
The prototype Cherny subsequently built was well-received internally. In February 2025, Anthropic launched Claude Code as a preview version. It uses the underlying model to write, edit and run code, and subsequent model upgrades have continuously promoted product development.
Claude Code has become an important driving force for Anthropic's growth. Cherny now leads this product.
A Product Mechanism That Allows Failure
Looking back at this process, Labs' advantage does not come from an isolated creative idea. The research team first discovers changes in capabilities, the product team builds prototypes based on this, internal usage verifies demands, and subsequent model advancements continue to improve the experience. The closer the connection between these steps, the more opportunities the product has to seize technological progress in a timely manner.
However, seeing capabilities in advance does not mean that every idea is worthy of long-term investment.
Labs conducts a "stick or pivot" evaluation of projects approximately every two weeks. This cycle is used to check the direction, and does not mean that all products have to be completed within two weeks.
If an idea performs poorly, the team can terminate it, or incorporate its valuable parts into other explorations. Participants then move on to new tasks. Mann mentioned that some attempts were eventually integrated into products such as Claude's Chrome extension.
Therefore, the 20% to 30% success rate cannot be simply interpreted as the rest of the work being completely wasted. Even if a prototype does not grow into an independent product, it may still leave behind reusable functions, technical experience or demand judgments.
Another rule is equally critical: when the project team usually has more than four members, it will "graduate" from Labs.
Both Claude Code and Claude Design have been transferred to independent teams within the company in this way. Labs thus maintains a small scale and continues to explore the next batch of directions. This also explains why it has always been a team with relatively frequent personnel turnover. After the project matures, some members leave with it; Labs then adds new members to start new attempts.
Two-way Feedback Between Research and Products
Proven projects require more engineering investment and long-term maintenance, while small teams need to retain space to adjust directions. Allowing mature projects to become independent in a timely manner helps prevent Labs from being completely occupied by daily operations.
It also has corresponding management costs. Personnel are constantly changing, new members need to quickly understand research progress, and the team must continuously rebuild collaborative relationships. The effectiveness of the biweekly evaluation depends on the quality of judgment on prototype performance.
From Mann's description, the communication between Labs and the research team also proceeds in both directions. Product developers use new models to find opportunities, and gaps in practical tasks in turn push researchers to pay attention to new capabilities.
Mann summed up one of Labs' goals as expanding the "action space" of AI, that is, enabling powerful models to do more things in the real world. This statement grounds the goal of product work on specific tasks. Programming requires models to be able to manipulate code, design requires models to generate more appropriate visual results, and cross-language usage requires the system to understand the expressions of more people. Whether a product is viable will expose the capabilities that the model still needs to supplement.
Ecological Challenges After Commercialization
As these products enter more markets, Anthropic will also face more complex business relationships.
Forrester analyst Mike Gualtieri pointed out to Business Insider that the tools launched by Anthropic may directly compete with software that enterprise customers are already using. Claude Design is an example, as it enters the design software space occupied by companies such as Adobe and Figma.
There is a constant tension here: Anthropic not only provides model capabilities to developers and enterprises, but also builds its own products based on these capabilities. The more features the platform adds, the more opportunities it will have to overlap with other software products.
For users, more built-in capabilities may reduce tool switching. For ecosystem partners, they need to judge what differentiated value their own products can still provide. How Anthropic balances the two sides will affect the attractiveness of its platform.
If the company completes its listing in the future, Labs will also need to reserve space for highly uncertain exploration under clearer business objectives. It is impossible to judge its listing time based solely on this report, nor can it pre-determine how R&D investment will be adjusted after listing.
At least in Mann's vision, Labs' direction does not stop at office and development tools.
He hopes that the team can participate in breakthroughs in biological research, clean energy storage and other real-world applications in the future. He believes that accelerating basic scientific research will help AI bring practical benefits to more people.
These are still visions, and there is still a long way to go before verifiable results are achieved. But they continue the working method that Labs has already adopted: observe the capability boundaries, find problems worth solving, and then test the ideas through prototypes.
Claude Code has proved that one of these paths works. The next test for Labs is to continuously find the next task worth investing in, between the ever-changing model capabilities and user demands.
For this team of around 20 people, the key may never be to make every project succeed. Whether it can timely stop attempts with limited prospects and give promising products sufficient resources determines how far this incubation mechanism can go.
Original article link: https://www.businessinsider.com/anthropic-labs-team-ai-innovation-ipo-2026-9
This article is from the WeChat Official Account "Machine Heart", republished by 36Kr with authorization.