HomeArticle

OpenAI has lost the "God of CUDA Kernels"

机器之心2026-08-17 16:56
Scott Gray has left his position to independently explore neuroscience-inspired AI methods

Today, a tweet posted a couple of days ago went viral, with only 5 words in its content: Scott Gray has left OpenAI.

This renowned engineer, who is hailed as "the God of CUDA Kernels and the world's most powerful GPU programmer", has become another heavyweight talent that left OpenAI this year. Although he has not posted a tweet to confirm the news and his LinkedIn page has not been updated yet, his personal profiles on 𝕏 and Bluesky have both been changed from "GPU Geek at @OpenAI" to "Former GPU geek at @OpenAI", and he has added a line in front that "he is currently independent and exploring some neuroscience-inspired AI approaches"

An engineer who has stayed at OpenAI for ten years announced his departure with only one line of profile.

It is such a coincidence. Exactly ten years ago on August 16, OpenAI released a blog post titled "Team update", introducing five full-time members who joined the company in that month. The list included Dario Amodei, Filip Wolski, Jack Clark, Scott Gray and Zain Shah.

https://openai.com/index/team-update-august/

Today, ten years later, Dario is the CEO of Anthropic, Jack Clark is in charge of policy at Anthropic, and Scott Gray has marked his experience at OpenAI as "former".

Up to now, OpenAI has not issued any statement on Scott Gray's departure, nor has Gray himself released a detailed explanation. The common source of all current reports is the change of his social account profile, and a breaking post that spread on X afterwards. There is no public information about the specific departure date, his next destination, and whether he has joined a new company.

It is worth mentioning another detail on his account. Gray's last original 𝕏 tweet was posted on November 20, 2023, with only one sentence: "OpenAI is nothing without its people." That was the third day after Sam Altman was dismissed by the board of directors, more than 700 employees signed a joint letter demanding the board to resign, and this sentence was copied and pasted repeatedly inside OpenAI at that time, which was regarded as a show of collective allegiance. After that, although he replied to some tweets (the replies also stopped after February 2025), he never posted any original tweets again.

Scott Gray Profile

We have published an article specially introducing Scott Gray before, you can refer to "The God of CUDA Kernels, the world's most powerful GPU programmer? Who is this behind-the-scenes guru at OpenAI". We will make a brief review here.

His reputation came from his time at Nervana Systems. Back then, most developers relied on NVIDIA's CUDA C/C++ and official libraries such as cuBLAS and cuDNN, multi-layer software abstraction shielded hardware details, and also formed the performance ceiling. Gray judged that these abstraction layers must be bypassed, so he wrote maxas, an assembler for the Maxwell architecture, which allows developers to write SASS machine code directly, allocate registers manually, manage memory latency, and control the instruction pipeline.

To prove this path is feasible, he handwrote an SGEMM kernel with maxas. On the GM204 GPU, this kernel reached 98% of the theoretical peak hardware performance, and was 4.8% faster than NVIDIA's official closed-source cuBLAS, which was also handwritten by experts. The subsequent maxDNN extended the same methodology to convolution: it stably reached 93% to 95% of the computational efficiency on all convolutional layers of AlexNet, while the efficiency of cuDNN in the same period fluctuated greatly between 32% and 57%.

This is the methodological foundation of all his subsequent work: do not accept the performance upper limit given by the abstraction layer.

In that OpenAI team update post in August 2016, the official only gave two technical evaluations of him, which roughly meant that he had focused on optimizing the performance of deep networks on GPUs at Nervana before, and his assembly-level optimization on dense linear algebra and convolution is "still the fastest to this day". The same paragraph also mentioned that when he is not writing code, he usually reads the latest research in neuroscience and related fields.

It seems that his new direction " neuroscience-inspired AI approaches " ten years later is actually a return to that unchanged interest.

After joining OpenAI, his role changed from an "optimizer" to an "enabler". In 2017, he co-released the block-sparse GPU kernel with Alec Radford and Durk Kingma. Different from unstructured sparsity that removes individual weights, block sparsity cuts the weight matrix into fixed-size blocks and zeros the entire blocks, the dedicated kernel completely skips the zero-value blocks during calculation, and can be several orders of magnitude faster than cuBLAS that processes dense matrices or cuSPARSE that processes general sparse matrices. This set of kernels was open-sourced at that time, which directly gave birth to a series of subsequent sparse attention work.

Schematic diagram of OpenAI block sparse kernel, https://cdn.openai.com/blocksparse/blocksparsepaper.pdf

Following this line, the authors of "Generating Long Sequences with Sparse Transformers" in 2019 are Rewon Child, Scott Gray, Alec Radford and Ilya Sutskever — they reduced the time and memory overhead of attention from O(n²) to O(n√n), and the paper explicitly listed "fast attention training kernel" as one of the three contributions.

After that, his name appeared in the author list of GPT-3, the author list of "Scaling Laws for Neural Language Models", the author list of DALL·E, and the technical report of OpenAI Five.

During this period, he also got the title of "the God of CUDA Kernels". That was in September 2025, former OpenAI employee Rohan Pandey posted on 𝕏 that only about one person in the company was responsible for the CUDA kernel on the inference side, and his colleagues called the attention kernel he wrote "the Bob kernel", which executes trillions of times every day on hundreds of thousands of GPUs. After the post became popular, the comment section generally guessed that "Bob" was Scott Gray.

Related posts also mentioned that there may be less than a hundred people in the world who can write high-performance CUDA kernels for the training process (especially backpropagation). Because this skill requires proficiency in parallel computing theory, GPU hardware architecture and deep learning algorithms at the same time, while most practitioners stay at the application layer or inference optimization layer.

And now, Scott Gray has left.

In 2026, OpenAI has lost many employees

OpenAI has lost many talents this year. According to the inventory of 𝕏 blogger Chubby, 12 senior leaders have left OpenAI this year, covering operations, business, product, research, security, ethics and hardware.

A few days ago, Brad Lightcap, who worked at OpenAI for 8 years and served as CFO and COO, announced his departure on 𝕏, saying he would "start something new"; two days later, Denise Dresser, the chief revenue officer who took office only last December, announced her departure, and Dali Rajic, former president and COO of Wiz, took over the position.

On the business and product line, Fidji Simo, CEO of the application business, took a leave of absence for health reasons in April, and stepped down from his full-time position to become a part-time consultant in July; former CMO Kate Rouch stepped down in April for treatment needs; Srinivas Narayanan, CTO of the enterprise business, also announced his departure in April. On the research and product side, Kevin Weil, who led OpenAI for Science, and Bill Peebles, head of Sora, announced their departure on the same day in April, Weil is currently running an AI science startup; Barret Zoph returned to OpenAI from Thinking Machines Lab in January this year to lead enterprise-level AI sales, and left again in June, the company did not explain the reason.

The three people on the security and ethics line left intensively in July: Johannes Heidecke, head of the security systems, left during the reorganization of the security team, and his responsibilities were merged into the research department led by Mia Glaese; Joshua Achiam, chief futurist, left after nearly nine years, the mission alignment team he previously led was disbanded in February this year; Chloé Bakalar, head of ethics, left after less than a year of joining. Earlier in March, Caitlin Kalinowski, head of robotics and consumer hardware, resigned because she disagreed with the company's cooperation with the Pentagon, she said on LinkedIn that related issues should go through more full deliberation, and later added that her objection was mainly aimed at the governance process.

This list also misses one person: in January this year, Jerry Tworek, vice president of research in charge of inference models, left to start his own business after nearly seven years, the reason is that he wants to do some research directions that "are difficult to carry out inside OpenAI".

Conclusion

Scott Gray's departure may not immediately change OpenAI's model release rhythm. The engineering system accumulated over ten years will not stop functioning just because one person leaves. But for a company that increasingly relies on scale, efficiency and cost control, losing such an engineer who can push hardware performance to the limit and has deeply participated in the development of multiple generations of core models is still a change that cannot be taken lightly.

The timing of this departure is also worthy of attention: OpenAI is preparing to go public, and become a giant enterprise that pursues revenue and participates in the global infrastructure competition. Now, the people who first shaped it are also re-judging what they want to do next.

Scott Gray's choice is not to join another cutting-edge lab or announce to start a business, he just said in his personal