OpenAI Executives' Firsthand Account: How We Built a Jev Competitor in Just One Week
The results of Computer Use are clearly very easy to verify, so why has the progress of this technology been so slow? When will AI truly learn to operate computers?
In June this year, this query raised by well-known podcast host Dwarkesh Patel sparked considerable controversy across the industry.
Three months later, OpenAI showcased Dot, Computer Use and a series of new developer interfaces at its DevDay on September 29. Each Dot can access an independent cloud-based Linux computer, which can not only open browsers, but also run desktop applications; users can entrust it with tasks that originally required them to complete step by step on web pages or in software. After the product demo, this question has become even more worthy of further exploration: what exactly can Computer Use achieve at the current stage, and what changes have taken place in the capabilities that support agents to complete tasks?
After DevDay, The AI Engineer Podcast from Latent Space had an in-depth discussion with the CUA team and the head of the API platform at OpenAI. The podcast first interviewed Ari Weinstein, co-founder of Sky and current head of Computer Use Product and Engineering at OpenAI. Ari believes that the most critical progress in the past year is that agents are better at troubleshooting errors and retrying after encountering obstacles. They no longer can only click step by step by looking at screenshots: page structures, accessibility information and self-written code can all become ways for them to operate software. However, when agents are able to access websites and process payments, at which links humans should step in to check remains a question that the product must answer.
Nikunj Handa, Head of OpenAI API Products, shared various practices developed at the API layer for agents during the interview: from allowing the model to continue reasoning while tools are executing, to fast decision-making, performance optimization and context management. When talking about the Decisions API, Nikunj Handa admitted that it was inspired by Jev and was not originally on the R&D roadmap at the very beginning. About a week after the project started, the team had already completed a runnable prototype. Instead of retraining the model, they reused Luna's weights, constrained the output to structured results, and specially optimized the inference system to minimize the time to first decision return, so that multiple problems can be processed in parallel in batches; it is more like an extremely fast classification and decision layer that can be used for customer service ticket classification, evaluation and scoring, and Computer Use. It can also cooperate with GPT Live to make agents more responsive when quick judgment is needed.
Starting from the query of "Why is AI still not good at operating computers", the two interviews gradually zoom in the camera: how agents find that they have made a mistake, how they decide the next step, and how they keep working in long and trivial real tasks. The demo at DevDay presented the results, and this interview shared their practices in Computer Use from different perspectives of products and platforms.
TL;DR:
Host Vibhu: Dot provides each agent with an independent cloud computer. How should users explore its capabilities?
Ari Weinstein: Dot can use browsers and desktop applications on a cloud-based Linux computer. A practical starting point is to sort out your daily computer tasks that take up a lot of time, and try to delegate suitable ones to the agent.
Host Swyx: There is a view that Computer Use has barely made progress in the past two years. What is the actual situation?
Ari Weinstein: In the past, agents usually could start a task, but easily got stuck after encountering problems. Today, they are better at troubleshooting errors, adjusting methods and retrying. This is one of the most notable changes in the past year.
Host Vibhu: Does the improvement of capabilities mainly come from the model or the agent runtime framework?
Ari Weinstein: Both are playing a role. Agents can now combine screenshots, accessibility information, page structures and Playwright to operate software, and they can also write code to execute multiple steps at once; the model's own capabilities and speed are also improving.
Host Vibhu: What are the main bottlenecks for Computer Use in the next stage?
Ari Weinstein: Bottlenecks are distributed in multiple links such as models, inference, runtime frameworks and information presentation. As the execution speed of agents increases, the time consumed by operations like waiting for websites to load has become increasingly noticeable.
Host Swyx: What should developers pay attention to when using Computer Use capabilities through the Agents API?
Ari Weinstein: You should limit the websites and applications that agents can access according to tasks, and ask for user consent before operations that may have important consequences such as payment. Reliability and appropriate security checks are the foundation of building user trust.
Host Vibhu: How will Computer Use change the software development process?
Ari Weinstein: Agents can actually open and test the software they develop. In this way, writing code and verifying results can form a closed loop, reducing the situation where humans take over all testing after development is completed.
Host Vibhu: What notable changes are there in the APIs released this time?
Nikunj Handa: Asynchronous function calls allow the model to continue executing while the tool is running, and then get the results afterwards; mid-prompt steering allows developers to add messages during the model's operation. These capabilities help handle tasks with frequent tool calls and long execution time.
Host Swyx: How did OpenAI start developing the Decisions API?
Nikunj Handa: The launch of Jev attracted the attention of users and internal teams, and the Decisions API was not on the development roadmap before that. The team first verified whether the prototype was feasible, and then set out to optimize the speed of decision return.
Host Vibhu: What problems is the Decisions API suitable for solving?
Nikunj Handa: The clear use case at present is fast classification, such as processing customer service tickets. It may also be used in certain tasks that require quickly selecting the next action, but fast decision-making requires different capabilities from completing complex, long-cycle tasks.
Host Swyx: If both use Luna, what is the difference between Decisions API and ordinary structured output?
Nikunj Handa: The first version does not retrain the model, but is based on the existing Luna weights, imposes constraints on the output, processes multiple problems in parallel, and specially optimizes the time to return the first decision.
Host Swyx: How do long-running agents deal with the context limit?
Nikunj Handa: Developers can set a threshold to let the Responses API automatically compress the context; they can also call /compact to control the compression timing by themselves.
1
When AI has its own computer, what tasks can be delegated to it?
Host Vibhu: Today is OpenAI DevDay, and we are recording this special interview on site.
Host Swyx: We are the first podcast interview you accepted after the live stream ended.
Host Vibhu: Today we have Ari here, who leads the product and engineering team for Computer Use agents. Before diving deep into this technology, can you first review what has been released today?
Ari Weinstein: We just came out of the keynote, and there are several releases related to Computer Use worth noting today.
The first is Dot, a new personal assistant product with some highly anticipated Computer Use features built in.
There is also GPT-6.1 Sol, a very outstanding new model. I think it is especially suitable for Computer Use, because it has advantages in both cost and speed. We seem to have announced that its cost is one fifth of Astra's; specifically for Computer Use, the cost is only one seventh, which is really remarkable.
The Agents API now also has Computer Use capabilities, and developers can build products using the same implementation as in Codex and ChatGPT. The demo also showed existing features such as app shots, which can quickly bring the content you are processing on your computer into Codex and ChatGPT.
There is also native Computer Use on Mac: Roman let it automatically take screenshots of its own applications, and while it is operating the apps, he can still do other things on his computer. So, the keynote was really wonderful.
Host Swyx: Not to mention the Decisions API. Let me ask directly: are all of these powered by the same model behind the scenes, or different models distilled from the same dataset? Specifically, does Computer Use use the Decisions API, or are the two relatively independent?
Ari Weinstein: The Decisions API is very interesting, it has several new capabilities: it can infer in parallel, does not perform thinking and reasoning, and the model it uses is smaller than the one we use for Computer Use. This makes it very fast, but its capabilities are slightly weaker when handling complex tasks with long execution cycles. How to combine these methods is still an open question to be explored. I am really looking forward to seeing what everyone will build with the Decisions API.
Host Vibhu: One interesting point is that every Dot now comes with its own personal computer.
Ari Weinstein: Exactly.
Host Vibhu: So they seem to be able to retain the working environment more persistently. You have been using it for a while, how should everyone explore its capability boundaries? In what directions should they try? I myself often ask it to handle customer service issues now. For example, "This thing is broken, I don't want to log in or do identity verification", you go find the relevant information and solve the problem. What else should everyone try to take it a step further?
Ari Weinstein: Dot is a very interesting product, because each Dot can use its own Linux virtual computer in the cloud, which is different from our other products. In the past, we usually provided a cloud browser, or let agents access your own computer; now you have a whole Linux computer in the cloud that belongs to you. It can run full desktop applications and also use browsers. I think Computer Use is so powerful and so anticipated precisely because it allows agents to do anything a human can do. All software in the world was originally designed for humans, and now agents can also use these software, so you can delegate work to them. So anything you would do on a computer, you can let Dot do. Which specific tasks are the most useful really depends on who the end user is and what things are valuable to their life.
I would suggest first thinking about what things you usually spend time on, and then seeing if you can delegate these tasks to agents. Anything you would do on a computer, you can let Dot do. Which specific tasks are the most useful depends on who the end user is and what things are valuable to their life. I suggest first thinking about what things you usually spend time on, and then seeing if you can delegate these tasks to agents.
Host Swyx: Right, for example, booking flights, shopping, and to be honest, even playing games and similar things are possible, right?
Ari Weinstein: Absolutely. I recently ordered a meal delivery service to eat healthier. I really like this service because it allows me to customize each meal very meticulously, like "how many grams of chicken I want, how many grams of rice". But the operation is so complicated that it took me two hours to place one order. Later I found that I could let Computer Use help me place the order, and it finished in fifteen minutes. With GPT-6.1 Sol, it completes this task eight times faster than me, saving me two hours at the same time. I think this type of task particularly reflects its value.
Host Swyx: As a creator, I can immediately name my primary use case: automating operations on YouTube. Many features of YouTube do not open APIs, you can only put it into a virtual machine and let the agent run it by itself. For example, A/B testing, or posting community posts, these have no APIs because they hate developers.
Ari Weinstein: I have also heard people from the developer experience team say that. They often use it to handle things on YouTube, and it works really well.
2
Computer Use has entered a new stage: better at understanding interfaces, and better at troubleshooting and speeding up
Host Swyx: I want to ask a slightly sharper question. We have friends who run a very influential AI podcast, and they have a well-known view: Computer Use has not made any progress in the past two years. This statement is very interesting, and you should be one of the most qualified people in the world to talk about this topic. What progress has actually been made?
Ari Weinstein: I remember they said that a few months ago, and I hope they have changed their minds now, because Computer Use has been completely different from what it used to be.
Host Swyx: You have spent almost your entire career doing some form of computer automation, from working on Shortcuts at Apple, to Sky, and then joining OpenAI. Can you talk about the main line running through this experience? What drives you? What things were impossible in the past, and what milestones have there been?
Host Vibhu: Let me add one more question: what is the most significant change from the Codex Computer Use last week to today? Is it the model, Dots, or the agent runtime harness? Apart from the whole history, I also want you to make it clear what exactly today's release has changed.
Ari Weinstein: I have always been passionate about automation, about helping people automate tasks, because it can save time in life and let people focus on things that are more important to them, instead of operating the computer in detail. That's why we built those products. I used to work at Apple, then founded Sky, and finally joined OpenAI, and this experience makes me very excited.
Host Swyx: It feels like you've been trying to bypass Apple's restrictions, until Apple said: "Okay, we might as well hire you and let you do these things from the inside." Is that right?
Ari Weinstein: Working there was indeed a great experience. Looking back at Sky, it's interesting that we were also doing Computer Use at that time, but the model's capabilities were much weaker. In just the past year, the model's Computer Use capabilities have become extremely strong. The biggest change I see is that in the past they could start a task relatively reliably, but they would run into problems halfway through; now, they are very good at troubleshooting errors, retrying, and reviewing which methods work and which don't.
The field of Computer Use itself is also evolving, and we have adopted more technologies. Current Computer Use often writes code. If you manually expand the tool calls in Codex, you will see that it does not just execute one action at a time, but writes JavaScript code to hand over to the computer for execution, sometimes completing many actions at once, with improved speed and capabilities. We have adopted more multimodal interaction methods such as accessibility interfaces, and the model may use screenshots, or accessibility interfaces, or Playwright. It can choose many different mechanisms according to the task at hand. The acceleration of the model itself is also quite amazing.
As for what is different today, I think we have been continuously improving Computer Use, so looking at the changes in a single day may not be as meaningful as looking at the past one or two months. However, I think the Computer Use in Dot and the new model released today are very worth looking forward to.
Host Vibhu: In the keynote, Tejal mentioned that the speed of Computer Use has increased sevenfold, and it performs much better on several benchmark tests. How do you measure that? As you said, Computer Use is constantly evolving. Do these improvements come from the agent runtime framework, models, or post-training? What changes have the new models brought?
Ari Weinstein: We actually have many measurement methods, some of which test different combinations and configurations of the agent runtime framework. The situation here is a bit complicated, because officially launched products will have more security checks, and different configurations will be adopted according to the needs of the current task. So there are many measurement methods, but no matter which one we use, the improvements we see are quite consistent. These improvements sometimes come from the runtime framework, and sometimes from the model. One result that impressed me deeply is that compared with Astra, the cost-effectiveness improvement of GPT-6.1 in Computer Use even exceeds the drop in its own base cost. It's really great to see that.
Host Swyx: There is a diagram in the live stream that I really like, showing the Pareto frontier of you continuously improving this curve. You have also talked a lot about how to improve together with the runtime framework. Can you give a few examples that made you suddenly realize something? It can be either the model driving the framework to improve, or the framework driving the