Deployment of global computing power nodes: GoodVision AI builds a full-domain dynamic scheduling AI Compute Grid
Jensen Huang published a post about AI factories, which received over a million views in less than 12 hours.
The post did not discuss more powerful models, but focused on land, power, computer rooms and complete computing power systems. He judged that infrastructure, rather than algorithms, is becoming the main constraint on AI growth. Because these things that sound far away from ordinary users are determining an increasingly realistic question: why are model calls getting cheaper, but enterprises' AI bills may grow faster than their business?
The answer lies in a seemingly ordinary task:
A software company asked an Agent to assist in troubleshooting online faults. It needs to read alarms and operation logs, locate abnormal codes, retrieve internal documents, analyze upstream and downstream dependencies, generate repair plans, then call tools to run tests and check results. What the engineer sends may be just one instruction, but the Agent has to perform dozens of steps continuously in the background to complete it. Each step may call a model, and each call consumes computing power.
These calls are not the same type of call. Sorting out logs does not require the most powerful model, but judging the cause of a fault may require complex reasoning; some codes can be handed over to cloud models, while logs containing customer data must remain inside the enterprise; ordinary analysis can wait a few more seconds, but alarms from production systems cannot afford long network round trips. If all tasks are sent to the same model and the same data center, the enterprise will not necessarily get better results, but will definitely bear higher costs and greater latency.
This is the true meaning of AI entering the infrastructure stage. Competition no longer only occurs between models, but also in every allocation behind the models: what task should use what model, run on which node, and be completed at what cost and speed. When such choices happen millions of times every day, what determines the enterprise's AI cost is not only how much computing power it buys, but also whether the computing power is used in the right place.
This is also the field that GoodVision AI is laying out. It uses modular AI Factory to undertake local and regional inference demands. A single factory can deploy up to 1.5MW of inference computing power in a 200-square-meter space, and relevant project plans have covered Japan, South Korea and the United States. But what is more noteworthy than how much computing power is built is that GoodVision AI is trying to answer another question on top of these infrastructures:
When a task is split into more and more calls, who decides where each request should go?
The money spent on training is often one-time, but the money spent on inference is permanent
In the past few years, the industry's attention has been almost entirely focused on training - parameter scale, training clusters, and leaderboard results. This makes sense, as that is the stage that determines the upper limit of model capabilities.
But training and inference are two completely different ways of spending money.
Training is like building a power plant: huge investment, concentrated cycle, once it is built, it is done, and a model is mostly trained only once. Inference is like power transmission and consumption: every call, every step executed by every Agent, and the operation of every terminal device consumes computing power in real time. A model is trained once, but inference will happen billions or trillions of times.
GoodVision AI saw this difference very early. David Wang, Founder and CEO, founded the company in 2019, focusing on the infrastructure gap that emerged when AI capabilities moved from "demonstrable" to "scalable delivery". In his words, AI infrastructure is gradually shifting from focusing on "where the computing power is" to focusing on "how each Token is efficiently produced, routed and delivered". While the industry focuses on the upper limit of model capabilities, GoodVision AI focuses on who manages the operating costs below that upper limit.
Today, this judgment is becoming the industry consensus.
The figure Jensen Huang presented at GTC 2026 is that the computing demand for AI inference has increased by about a million times in the past two years. Correspondingly, global capital expenditure on AI infrastructure exceeded 400 billion US dollars in 2025, and the 2026 forecast will further exceed 600 billion US dollars. AI infrastructure has not cooled down along with the training competition, but has entered a longer construction cycle due to the explosion of inference demands.
Agents also make inference costs more complex.
What Agents change is not only the consumption of Tokens, but also the way of consumption. In the past, the user asked a question and the model answered one; now, an Agent needs to read materials, call tools, check results, and continue to act according to feedback in order to complete a task. A task is thus split into dozens or even hundreds of model calls.
Tokens are changing from "conversation consumption" to "system consumption". When calls enter the main business processes of enterprises, costs are no longer occasional API fees, but become operating expenses that continue to grow with the business.
The true meaning of this for enterprises is: the cost structure of AI has changed from one-time procurement to a bill that is always updating. As long as the bill is continuous, efficiency is no longer an engineering problem that can be optimized later - it is an operational problem from day one.
As a result, a seemingly contradictory phenomenon begins to appear: the same enterprise may face both surplus and shortage of computing power at the same time. High-end computing power is queued up for competition, while a large number of simple tasks that could originally be completed by lightweight models locally in dozens of milliseconds occupy the most expensive resources and travel around the earth before the answers are sent back.
What really needs to be solved is how to allocate computing power: which requests should use flagship models, which requests can be handed over to lightweight models, and which requests should remain local.
Every call is a multiple-choice question
This is exactly the layer that GoodVision AI wants to solve. David Wang has many years of deep experience in the cloud computing industry, with resumes spanning IBM, AWS, Alibaba Cloud and Tencent Cloud. He once served as IBM Partner, Senior Leadership Team Member of AWS, participated in the early team building of Alibaba Cloud, and later served as Head of Tencent Cloud North America.
These experiences make him pay more attention to the practical application of AI in enterprises. The technologies used by enterprises are constantly changing, from cloud servers and GPUs to model APIs and Agents, but the demand has always been the same: obtain truly useful AI capabilities at lower costs and in a more stable way.
GoodVision AI's first foothold was cloud services, which is also its first layer of entry connecting enterprises with global computing resources. Facing AWS, Microsoft Azure, Google Cloud, as well as enterprises' own private clouds and local computer rooms, it helps customers complete cloud migration, architecture design, multi-cloud governance, operation and maintenance, and continuous optimization. The combination depends on business scale, cost, data security and global deployment requirements, without being bound by a single cloud vendor.
The value of this layer of business is first reflected in "helping enterprises move to the cloud", and at the same time allows GoodVision AI to move in close to the real workload - figuring out how customers use their computing power, where Tokens are spent, where latency is stuck, and where the compliance boundary is drawn. This is information that external suppliers can hardly get from PPTs. When enterprises start to introduce large models and Agents, this capability naturally extends to model calls, cross-cloud scheduling and privatized inference.
This judgment is completed by the Smart Routing Engine.
After an AI request enters the platform, the system first identifies the intent of the request, then comprehensively considers task complexity, Token cost, response speed, data sensitivity, regional compliance and node availability, and allocates resources in real time between different models and computing power resources.
In specific scenarios, complex reasoning and general knowledge tasks can be routed to flagship models of cloud services such as ChatGPT, Claude and Gemini; lightweight tasks such as translation, summarization, customer service response, and repetitive processing are handed over to low-cost small models; requests involving sensitive data or requiring extremely low latency are sent to enterprise private models or edge nodes to reduce the transmission of data between external clouds.
The Smart Routing Engine also supports Agent to Agent (A2A) calls. When an Agent processes a task, it often needs another Agent to retrieve, analyze, review or generate content. The system will arrange appropriate models and operation nodes for this call according to task objectives, capability matching, cost, latency and data permissions. In this way, multiple Agents can collaborate more stably to complete a job.
Multi-model routing answers "which model to call", and GoodVision AI wants to go one step further: where this model should run. The same model deployed in public clouds, enterprise computer rooms or edge nodes has completely different costs, latency and data boundaries behind it. Therefore, what the Smart Routing Engine schedules is not only the models, but also the computing power locations where the models are located - it not only provides a unified API entry, but also is responsible for resource scheduling of the enterprise AI system, that is: simple tasks use lightweight resources, and complex tasks retain necessary model capabilities; and when cloud platforms, model services or computing nodes fluctuate, requests can automatically switch to other available environments.
This logic has been verified by actual business.
GoodVision AI once served an image and video generation platform. The customer's monthly AI expenditure had exceeded 500,000 US dollars at that time, and Token consumption was still growing at a rate of about 50% per month. After migrating the relevant models to GoodVision AI Factory for privatized deployment, the customer's overall cost decreased by about 60%, and network latency decreased by about 50%.
Its way of "saving money" comes from the simultaneous rearrangement of models, computing power and deployment locations, rather than simply replacing a cheaper model.
Computing power is starting to move to places with large populations
After the request is selected, it must fall to a certain physical location. This brings out the heavier half of GoodVision AI's business.
David Wang has a judgment: the current AI infrastructure is still at a stage similar to "before the emergence of CDN". Centralized large clouds are still most suitable for training and complex reasoning, but as Agents begin to frequently process tasks such as translation, customer service, code, images and videos, more and more requests should actually be completed by nodes close to users, close to data, and close to business sites.
There are clear business reasons behind this. AI glasses, robots, autonomous driving, and smart factories are extremely sensitive to latency, and many scenarios simply do not allow requests to be sent thousands of kilometers away before waiting for results; data from industries such as finance, healthcare, and government affairs often cannot even take the step of "crossing borders".
GoodVision AI's AI Factory is a modular inference center oriented to local and regional demands. According to company documents, it uses immersion cooling to control PUE below 1.2, can deploy up to 1.5MW of inference computing power in a 200-square-meter space, and is connected to the Smart Routing Engine to become a local node in the global scheduling network. At present, relevant projects are being advanced in Japan, South Korea and the United States. Among them, the nodes in Japan are planned to be expanded in phases over the next three years, with a total scale of 100MW. The company's expansion path is to first build flagship nodes with local partners, then deploy edge nodes to surrounding areas, and replicate this infrastructure delivery model to more regions.
There is a point worth explaining: building data centers is famously heavy and slow, so how can a company founded in 2019 advance in three countries at the same time?
The answer lies in the last computing power cycle.
Traditional data centers, from land acquisition, construction application to power-on operation, are often calculated in years, while enterprises only need a few months to purchase GPU servers - this time difference is the most realistic gap in current AI infrastructure. GoodVision AI relies on existing mine and energy infrastructure resources, cooperates with North American infrastructure partners in Texas, Wisconsin and other places, and upgrades the existing facilities that already have power supply and sites into AI-ready inference infrastructure. According to company documents, the first phase of the Fukushima project in Japan is expected to be completed and put into operation within three months.
This is a quite meaningful industrial succession: the power and sites left over from the last computing power cycle have become the most scarce entry tickets for this round of AI inference implementation. What it brings is not so much a cost advantage as a time advantage - in a track where everyone is rushing for the construction period, this may be the more valuable one.
Just like the early oil industry competing for important oil fields, whoever can occupy energy and space resources in key regions in advance will have the ticket to enter the next round of AI infrastructure competition. In David's view, the competition for AI infrastructure is showing a logic similar to the energy industry.
Based on this judgment, GoodVision AI plans to connect AI Factory nodes distributed in different countries and cities, and eventually form a dynamically schedulable global "AI Compute Grid", covering more than 100 cities and more than 100 AI Factories in the long run, so that computing power, like Internet content, can be more efficiently distributed to users according to demand.
In the future, such power indicators, land resources and infrastructure nodes that can carry AI computing power will become scarce assets.
From selling resources, to selling judgment
Looking back, the three businesses of cloud services, Smart Routing Engine and AI Factory are connected to each other and expand along the same link.
Cloud services allow GoodVision AI to access enterprise systems and connect to scattered multi-cloud resources; the Smart Routing Engine is responsible for understanding each request and completing real-time allocation of models and computing power; AI Factory provides controllable, low-latency, locally deployable underlying computing power. The three together form a complete link from enterprise demands, to Token scheduling, to computing power delivery.
This link is translating into rapid growth. According to the latest disclosed data of the company, GoodVision AI achieved revenue of about 24 million US dollars in the first nine months of the 2026 fiscal year ending June 30, 2026, which is nearly 5 times that of about 4.82 million US dollars in the same period of the previous year, a year-on-year increase of 397.9%. It is more noteworthy that the company achieved revenue of about 13.45 million US dollars only in the three months from April to June 2026, not only a year-on-year increase of 544%, but even exceeding the total revenue of about 10.55 million US dollars in the previous six months, showing that its business growth is still accelerating significantly. According to the latest management forecast, GoodVision AI's full-year revenue for fiscal 2026 is expected to reach about 38.87 million US dollars, about 5 times that of the previous fiscal year's 7.74 million US dollars.
For an infrastructure company still in the expansion stage, the significance of this set of figures does not lie in the absolute scale, but in showing that the business model has been verified, and the company is moving from the verification stage to the large-scale stage.
It is worth noting that what GoodVision AI delivers is changing. Early customers purchased cloud resources, with the core demand of obtaining more appropriate computing power. As the Smart Routing Engine takes on more enterprise requests and AI Factories are put into operation one after another, the company needs to further judge which model each request should be assigned to, which node to run on, and at what cost and latency to complete.
The core that enterprises purchase has extended from computing power itself to the way of using computing power.
Affordability to buy, and affordability to use
Return to the scenario of online troubleshooting at the beginning.
This kind of cost problem usually comes from the mismatch between model selection and computing power scheduling. Millions of requests with different difficulties are processed in the same way. As the industry shifts from training to inference, how to make models and computing power adapt to specific tasks has become a required lesson that enterprises must make up for.
The electricity that truly changed the world began with the spread of the power grid.
This network of AI is just starting to be laid now.