A startup that aims to disrupt the CPU industry.
Recently, a relatively new CPU startup has emerged, that is Nuvacore founded by senior engineers Gerard Williams III, John Bruno and Ram Srinivasan who have worked at Apple, Nuvia and Qualcomm before.
Judging from their resumes, all of them have rich experience in custom CPU design. For example, Williams previously co-founded Nuvia, which was later acquired by Qualcomm. Besides, Nuvacore has also hired David Williamson, former head of CPU engineering at Apple, as Senior Vice President of Hardware Engineering.
It needs to be emphasized that the background of the company's founding team is not a simple stack of resumes. These people have built the most disruptive CPU architectures in the past two decades. They are not managers who manage chip projects, but architects who have made decisions on key microarchitecture choices such as out-of-order execution depth, branch prediction strategy, memory subsystem design and power management. These decisions directly affect the performance of chips, and their efforts have achieved huge success.
It is worth mentioning that this company has received support from HSG. Considering the past successful experience of HSG, the future of this company is particularly worthy of attention.
Why CPU?
The choice of the CPU track is naturally related to the background of the team. On the other hand, it is also related to the change of industry demand.
As everyone has discussed in the past few months, large market participants have been moving in a certain direction to cope with Agentic AI. NVIDIA's Vera CPU, the AGI CPU previously released by Arm, and Intel's positioning of Xeon as an inference processor all show that the main role of CPU in Agentic AI is orchestration: managing tool calls, routing between models, and processing control flows that GPUs and accelerators are not good at originally. This generalization is very accurate - but the actual situation is far more than that.
Think about what an agent system does all day long. It won't pop up, process a batch of tasks, and then shut down. On the contrary, it runs continuously. It maintains context information in conversations spanning hours or even days, continuously extracts inference traces from memory, refreshes vector storage, and checks existing knowledge before deciding the next action. These memory accesses are everywhere. They are irregular, unpredictable, and completely different from the clear and repeatable patterns that CPU cache architects have been optimizing for the past 30 years.
The CPU does not run independently. It needs to call multiple tools, start sub-agents, wait for responses, and feed back the results to the next decision. At the same time, it also shares the same system resources with the adjacent GPU, the underlying storage architecture, and the network connecting all other systems. Sounds complicated? It is. Moreover, any tiny efficiency loss in this loop will not stay small all the time, it will accumulate continuously. At production scale, millions of agent interactions run simultaneously, and the requirements for latency and consistency far exceed the design capabilities of today's server CPUs.
The design of the current generation of high-performance CPUs is based on workload assumptions that predate the agent operation model. Server CPUs are throughput-optimized for predictable enterprise workloads, while client CPUs are optimized for burst performance and aggressive power management. Neither design can perfectly match a processor that needs to coordinate millions of agent interactions continuously, efficiently and with low latency.
Based on the above thinking, our optimistic interpretation of Nuvacore is roughly as follows: the team looked at the current market landscape and came to the conclusion that no one has yet fundamentally built a processor suitable for Agentic AI. Arm's AGI CPU is just an improvement on existing products. The design concept of NVIDIA Vera is to run in parallel with the GPU. Intel Xeon has made some extensions on the traditional architecture.
All in all, the whole market is iterating based on assumptions that already existed before people really understood the characteristics of agent workloads, which further strengthens their confidence in entering this field.
The Mystery of Instruction Set Architecture
After confirming the CPU track, they resolutely entered this track.
The company aims at "general-purpose CPU cores" to meet "the high-intensity and continuous needs of advanced AI systems and agent computing", which is a very interesting statement. They are not designing TPUs, nor matrix accelerators, let alone dedicated inference engines like Cerebras and Groq. What they are building is a CPU. The core argument of Nuvacore is that this CPU must bring fundamental changes to the field of agent computing. This is a very interesting point of view.
In an article published on September 29, the company introduced the unique features of WarpCore they developed. According to the introduction, the product adopts the so-called "core-first" development mode:
"Core First is a method that breaks through the traditional industry construction mode. It gives our NUVACORE team the freedom to first build the absolute best and highest-performance core, and then choose the instruction set that is most suitable for the system we - or our partners - want to build."
Nuvacore said that they will not choose the instruction set architecture at the beginning of the project, but can first build most of the basic CPU core IP, and then select the instruction set architecture according to the system that itself or its partners ultimately want to build.
This is completely different from the traditional CPU design method, in which the target instruction set architecture (ISA) usually determines the design direction from the very beginning. Nuvacore said that decoupling part of the microarchitecture from the instruction set allows its engineers to focus more freely on performance, energy efficiency and chip area.
The company has not yet disclosed which ISA WarpCore will eventually adopt. However, its recruitment materials mention that RISC-V, Arm64 and x86 architectures are all related to its engineering work, which indicates that the underlying design has fully considered the flexibility of ISA during development.
Of course, this does not mean that the instruction set can be simply added to the core at the final stage. Decoding, privilege levels, memory ordering, exception handling and software compatibility still pose major constraints on CPU design. With Nuvacore, it is obviously possible to develop enough performance-critical mechanisms before finalizing the instruction set architecture (ISA).
The company's published recruitment information reveals some clues about its basic work content. The company is recruiting engineers engaged in branch prediction, instruction scheduling, register renaming, out-of-order execution, load/store design, cache and memory subsystem, consistency and performance modeling. In addition, the company is also recruiting positions in RTL, verification and physical design, covering links such as tape-out and post-silicon verification.
The legal issues of finally selecting the instruction set architecture (ISA) have not yet been resolved. RISC-V is relatively simple because it does not require Arm-style architecture licensing. An Arm implementation based on Nuvacore's own microarchitecture would require the corresponding Arm architecture authorization at a certain link, and given that Arm's business is selling CPU design IP, obtaining authorization may be tricky, especially considering Arm's past legal disputes with Qualcomm/Nuvia. Finally, the threshold for x86 is higher, because its implementation rights are strictly controlled.
This leads to another possibility: WarpCore may be designed to be reusable CPU IP for partners with relevant architecture licenses to adapt. The company has not yet clarified whether its long-term business model will focus on full CPUs, licensable core IP, semi-custom chips, or some combination of the three.
At present, key details such as the process node, cache hierarchy, pipeline configuration, vector functions or tape-out plan of WarpCore have not been announced, but it deliberately keeps the instruction set architecture (ISA) decision open, which makes it one of the more unusual brand-new CPU projects currently under development.
The Difficult Problems Nuvacore Needs to Solve
With the above background, we can see that there are some real challenges here. First of all, the past achievements of the Nuvacore team are mainly concentrated in the client side, not the data center field. Their reputation was built at Apple, which designs chips for a closed ecosystem with known heat dissipation performance, known workloads, known software stacks, and only one single customer: Apple itself. Although Nuvia was determined to enter the data center field when it was founded in 2019, it was acquired before launching server products. The product actually launched by Qualcomm is Oryon - an excellent laptop chip, but after all, it is only a laptop chip.
The unknown side of data center work is: it is not glamorous engineering. At various conferences, no one will get applause for reliability, availability and serviceability (RAS) - at least not at those most popular conferences. But when you need to ensure that the CPU on the production rack runs continuously for five years, this is the work that really matters. Detect and isolate errors before they cause node downtime; cooperate perfectly with the hypervisor stack; have sufficient understanding of NUMA topology so that it will not become an unexpected bottleneck. This is the essence of data center work.
Intel and AMD have been doing the same thing on the x86 architecture for decades. This capability has been deeply integrated into their organizational structure, which cannot be seen on the product roadmap. It is a deep-rooted organizational memory derived from years of hard communication with enterprise customers when encountering failures at two o'clock in the morning. Arm's Neoverse project is also worth mentioning. The reason why these chips can succeed in the data center is not because of their excellent benchmark test results, but because Arm has spent years working in depth with companies such as AWS, Microsoft and Google to jointly solve a series of exhaustive data center-specific requirements that you will never see in press releases. This cooperation time is not optional, but a necessary condition for success.
There is no reason to think that the Nuvacore team cannot do this work; the required skills can be learned and recruited. But this is completely different from the projects they developed before, and the gap cannot be underestimated. It is worth noting that the company's published recruitment information includes the head of operating system, head of firmware, head of telemetry and observability, and head of software verification. The job list is correct. The question is whether such an execution timeline is realistic considering the patience of HSG and potential customers.
There is another challenge that has not been fully paid attention to here: is the data center really ready for the next generation of chips? I don't think so - but the reason may not be what you think. It is not a matter of technical maturity, but a matter of fatigue. Data center architects are not idle, hoping to have more suppliers to evaluate. CIOs are already tired of managing existing systems, and are working hard to integrate the new generation of AI infrastructure with the existing environment.
All companies that bring new chips into the store always underestimate the cost required for customers to finally decide to adopt the chips. Evaluating chips, certification, integration, training teams, and providing support when failures occur... all of these will consume the limited resources of enterprises.
To really get a newly designed CPU into the data center, there are usually one of the following three situations: the hyperscale data center adopts it as a custom chip; a mature OEM commits to building a platform around it; or the performance and efficiency difference is large enough to force cost pressure to prompt manufacturers to re-evaluate.
I think Nuvacore is pursuing the third path. But to take this path, it also needs to win the design order of hyperscale data centers (more likely) or establish a cooperative relationship with OEMs (maybe) before delivering mass-produced chips. To do this, a lot of trust capital needs to be invested, because this product will take at least two to three years to complete tape-out. I suspect the Nuvacore team already has part of this trust capital. As for whether the funds are sufficient and whether the schedule is consistent with the multi-year product roadmap planning cycle of hyperscale data centers, it is still unknown.
Finally, can we talk a little about the word "CPU"? Because I feel that it is now undertaking a lot of work beyond its original design.
At present, "CPU" is more of a positioning term than an accurate architectural description. What we are really discussing, whether it is Nuvacore, Arm's AGI CPU, NVIDIA Vera or Intel Xeon with AMX technology, is closer to "the host processor in heterogeneous computing systems in the AI era". This is a more accurate statement, although it will never appear in the product specification. So everyone is using the word "CPU" now, hoping that customers can understand its real meaning on their own.
Nuvacore will face this problem directly. What exactly they are building will directly affect how you view their competitive position. If they are building a high-performance general-purpose core that needs to have excellent IPC and energy efficiency under various workloads, it will be a real tough battle. On the other side, there are Intel, AMD and Arm, all of which have decades of customer relationship and certification history. This battle is not without a chance to win, but you'd better come up with a product that is far better than theirs.
On the other hand, if Nuvacore is rethinking the execution-level agent inference microarchitecture - for example, irregular memory access patterns, continuous state management, and low-latency orchestration requirements - then what it builds, I think, is difficult to be simply classified into any existing category. This may make business negotiations with non-hyperscale customers more difficult, because you are not just selling chips, but selling a brand new way of thinking to buyers who are already overburdened. But it is also a more interesting problem at the same time.
Skeptics believe that this is a world-class team chasing a data center market that they have never been involved in, and their products may take several years to complete tape-out. At the same time, hyperscale data center operators are developing chips independently, and enterprise IT departments are also facing resource constraints in a new round of certification cycle. The statement of "rewriting the rules of chips" is either a real description of a truly different architectural approach, or a common trick used by all chip startups to prove their own existence.
We can't tell which interpretation is correct until the company shows the chip. The support of HSG shows that this powerful financial company with outstanding performance in chip market evaluation and investment believes that the optimistic interpretation is more likely than the pessimistic one. This is certainly important, but after all, it is not a product yet. In the data center field, the gap between a striking architectural concept and a qualified and deployable CPU is exactly why most ambitious chip companies choose to remain silent.
This article is from the WeChat official account "Semiconductor Industry Observer" (ID: icbank), written by the Editorial Department, and authorized for release by 36Kr.