HomeArticle

Megvii reconstructs physical space operations with "Intelligent Agents": Where does its confidence come from?

晓曦2026-09-03 18:23
The new physics AI competition has only just begun.

Founded in 2011, Megvii has gone through multiple changes including listing shift and business innovation, and finally chosen a path that few people take. They are no longer satisfied with enabling AI to "see" the physical world, but to "transform" the physical world.

According to Megvii's definition, physical space operation refers to continuously perceiving people, objects, sites, rules and processes in physical spaces, understanding their spatiotemporal relationships and operating states, identifying anomalies and opportunities, and promoting the closed loop of business decision-making and actions.

Megvii has spent 15 years completing this leap from "being able to see" to "being able to act".

From "Perception Algorithm" to "Space Operation Agent", An Evolved Dimension Upgrade

In the eyes of the public, Megvii seems to be in constant transformation, but every step it takes follows clear rules.

Especially when we trace back to 2011 and review Megvii's 15-year technology evolution context, we will find that its current positioning is never a random decision, but an inevitable result.

Many people still remember the event of "paying for commemorative stamps with facial recognition" that took place in Hanover, Germany in 2015. That was only four years after its establishment, Megvii, as the behind-the-scenes algorithm supporter, made China's facial recognition technology debut on the global stage. Later, Megvii won the championship of global computer vision competitions, was approved as the national AI Open Innovation Platform for "Image Perception", initially completed the original accumulation of state perception capabilities for "people, events and objects", and first touched the threshold of physical space.

After that, Megvii launched the world's first smart camera C1, built the AI productivity platform Brain++, and redefined the product logic of AIoT with the concept of "algorithm-defined hardware". It is not difficult to see that at this stage, Megvii has integrated algorithms into hardware, and taken perception capability as the anchor point to enter the physical space.

In the following years, Megvii delivered the smart venue project for Beijing Winter Olympics, built multiple city-level projects including Beijing, and deployed digital perception systems for many Fortune 500 companies, continuously proving its operation and service capabilities in physical spaces.

From 2024 to the present, Megvii has released the Taiyi multimodal large model, and subsequently launched the Huanfang Agent, the Magic Cube Agent Analysis Box, and the Agent Application and Management Platform, officially building POA OS (Physical Operation Agent Operating System). So far, its intelligent technology stack for space operation has been gradually upgraded.

At the same time, the wave of the industry is also sweeping in, and the whole industry is beginning to discuss a more fundamental question: What is the ultimate destination of AI? Almost all answers point to Physical AI.

NVIDIA CEO Jensen Huang once proposed that the next era of agent AI is Physical AI. According to Future Markets' forecast, the global Physical AI market size will grow from about 383 billion US dollars in 2026 to 3.26 trillion US dollars in 2040. Frost & Sullivan estimates that the market size of China's Physical AI simulation and data platform will reach 1806.1 billion yuan in 2030.

The track is very hot, but Megvii's judgment is very calm: The real value no longer lies in "perception" itself, but in the full-link closed loop of "perception-understanding-decision-execution". Physical space operation is the concentrated embodiment of this judgment.

This is not a glamorous story. It means that Megvii will continue to serve government and enterprise customers, deliver projects, and deal with dust, noise and vastly different customer demands in every specific physical space. But from another perspective, this may be the most feasible path for Megvii. The 15 years of accumulation is not the leading edge of a single point algorithm, but the engineering capability and industry know-how honed in the real physical space.

Engineering Closed Loop & R&D System, Megvii's Most Underestimated Moat

There has always been a consensus in the AI circle: From top conference papers to production line deployment, there is a "valley of death" in between. Many people regard this valley as the gap in technology maturity, but people who have actually worked in the industry know that it is more like a completely different set of survival logic. Papers pursue to refresh the accuracy on standard datasets, while production lines face workshops with flickering light, customer data with uneven annotation quality, business processes that change every three months, and strict fault tolerance rate that "you will be criticized to the point of doubting life once a false alarm occurs".

Coincidentally, Megvii has been honing in this valley for more than ten years. Therefore, Megvii's full-stack engineering closed-loop capability is also "grown" layer by layer along with customer demands.

At the algorithm layer, Megvii firmly promotes the research on the collaboration of Physical large and small brains. On the one hand, it continues to invest in the cutting-edge exploration of general large models; on the other hand, for vertical industry scenarios, it carries out collaborative deployment of industry large models and scenario small models. At the hardware layer, Megvii actively deploys terminal intelligence and edge intelligence, and has built a complete AIoT hardware-software integrated product system, covering the full-link hardware requirements of physical space data collection and edge inference. At the system layer, Megvii has independently developed POA OS, which is the epitome of Megvii's scientific research strength, and also the core base of its engineering practice and agent product innovation. At the application layer, Megvii encapsulates the above capabilities into agent solutions for five major space scenarios.

The four layers of capabilities are progressive and interlocked: Algorithm is the "brain", hardware is the "sensory organ", system is the "nerve center", and application is the "hands and feet". The advantage of this architecture is that each layer is a reusable, iterable and combinable module. Once it runs through in a certain scenario, it can be quickly replicated to similar scenarios.

More importantly, Megvii has already run through these closed loops in real projects in the energy, transportation, operator and other industries. It can be said that Megvii is still one of the very few companies in the industry that has full-stack delivery capabilities from algorithms, hardware, systems to applications.

According to 36Kr's understanding, Megvii's technical R&D personnel still account for more than 60%, far exceeding the industry average; the total number of model releases in the first half of 2026 increased by 180% compared with the same period in 2025, and both mass production quality and efficiency have been improved. Megvii's current technical R&D is also carried out around two main technical lines.

The first main line is how large models can more efficiently enter industry space scenarios. The real industry scenarios are not closed, clearly ruled, fully annotated standard datasets. Megvii uses efficient post-training and few-shot solutions to enable large models to adapt to the judgment standards and cost constraints of different scenarios at lower costs.

The second main line is that agents continue to improve their autonomous capabilities in real physical spaces. The real physical space is open, dynamic and uncertain. Agents need to understand objects, spaces, behaviors and feedback, and complete the autonomous upgrade from perception to action.

In May this year, E-ViC, a spatial perception optimization algorithm jointly developed by Megvii and Beijing Humanoid Robot Innovation Center, was accepted by the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026). In five spatial understanding benchmark tests, E-ViC improved by an average of 10.1% compared with the base model, and achieved the most significant breakthrough in tasks requiring fine positioning (such as target placement and spatial reference), even surpassing large models with four times its parameter size and commercial flagship models such as GPT-5. In July, the multimodal context learning work proposed by Megvii was accepted by the global top conference ICML 2026. Every acceptance by a top conference represents a technical capability that can solve practical problems in real scenarios — which is exactly where Megvii's confidence in moving from R&D to production lines lies.

Attendees came to Megvii's booth for exchanges

Why POA OS is the "Deterministic Just-in-Need Middle Layer" in the Physical World

Strategic dimension upgrade requires a new technical base. Therefore, Megvii launched POA OS, namely Physical Operation Agent OS, the operating system for physical space operation agents.

This is an intelligent base of "Space as a Service", which is responsible for the standardized adaptation and scheduling of algorithm capabilities and hardware devices, enabling agents to achieve autonomous perception, autonomous memory, autonomous decision-making, autonomous execution and autonomous evolution in the real environment. If Megvii in the past was selling individual "algorithm parts", what POA OS tries to do is to assemble these parts into an operation machine that can be replicated in batches.

The most special part of POA OS is that it emphasizes "deterministic physical space rules and safety bottom lines". AI in the digital world can make mistakes: imperfect conversations, inaccurate recommendations, and users will at most complain. But once AI in the physical world makes a mistake, it may directly lead to personal safety risks, production accidents and property losses. One of the core functions of POA OS is to solidify the rules, processes and safety boundaries of the physical world into "hard constraints" at the system layer, so that agents can make autonomous decisions within the rule framework instead of acting arbitrarily.

This actually answers a core problem in the implementation of Physical AI: How to balance the autonomy of agents and the certainty of the system.

Megvii has taken a different path from most players. It does not sell robot bodies, does not develop general large models, nor does it only sell algorithms. What Megvii provides is "space operation capability", enabling an office building, a factory, or an urban area to be autonomously perceived, decided and operated by agents.

It is reported that Megvii is expected to achieve full profitability in 2026. The significance of turning losses into profits goes far beyond the financial figures themselves. It means that the path that Megvii chose, which "few people take" — to become a physical space operation service provider, is being verified by the market as a feasible commercial path.

Looking back on Megvii's 15 years, every transformation is an extension and iteration of its core capabilities. Algorithm is the starting point, hardware is the tentacle, system is the skeleton, and agent is the final delivered value form.

The path of physical space operation is very difficult. It requires a company to have top algorithm R&D capabilities, solid hardware engineering capabilities, complex system integration capabilities, and in-depth understanding of industry scenarios at the same time — it is not easy to do any of these four things well, and there are very few companies that can do all four well.

Megvii happens to be one of them.

From perception to execution, from "seeing" to "transforming", for a company with strong technical strength and accumulated engineering capabilities, this dimension upgrade is not too late at all. The new Physical AI competition has just begun.