HomeArticle

The "Last Centimeter" of Industrial Embodiment: Strong Momentum Unleashed After Nine Years of Deep Accumulation

36氪产业创新2026-09-30 10:19
Address the commercial proposition of how to convert model capabilities into industrial-grade certainty.

If a company wants to train an industrial AI large model from scratch, the first step it takes may not be modeling, but waiting — waiting for production line failures to occur, waiting for defect reports, waiting for data to flow back. The process may take a month, or even a year. An industrial AI large model without real production line data is like a structure built on quicksand. Data scarcity has long been an insurmountable barrier for this industry.

Nowadays, this pain point can be solved by taking just a few photos.

In September this year, Srit Intelligent, an industrial AI service provider, publicly showcased for the first time its augmented industrial large model "DataWing" trained on 9 years of accumulated 48TB real-scenario data, as well as the HaiBot robot equipped with the embodied big-small brain model CogniMotor and the predictive world-action model SeerWAM. After enterprises access the system, they can call the corresponding Skill with just one sentence or one image, and the efficiency of data generation is more than 10 times higher than that of manual collection.

This company has stayed out of the spotlight for a long time, but it is by no means unknown in the automotive industry: it is the first local industrial AI brand in China to enter the supply chain of international luxury car enterprises, holding more than 93 intellectual property rights; it has nearly 100 industrial end customers, more than half of which are leading automotive enterprises. Meanwhile, Srit's industrial AI solutions cover all four core processes of stamping, welding, painting and final assembly. There are no competitors for this set of solutions among local industrial AI companies, and its three-year repurchase rate exceeds 60%.

At present, the boom in the two tracks of large models and embodied intelligence is cooling down, and both capital and customers are asking: beyond performance ranking and benchmark testing, who can bring real value? Srit just responds to the common proposition of their commercialization: how to transform model capabilities into industrial-grade certainty.

Nine years of sharpening a sword: while others burn money to collect data, it turns production lines into "data oil fields"

On the surface, the DataWing has a very low threshold. Behind its out-of-the-box availability is its solution to the long-standing contradiction plaguing industrial AI: enterprises are eager to improve efficiency with AI, but it is difficult for them to continuously complete high-quality multi-modal data collection, so that they cannot train adapted models.

Compared with C-end large models, industrial data is more scarce: factories have high confidentiality, and data is only open before customer acceptance; even leading robot companies can only enter the site after work stops at night to collect limited data. The continuous supply of industrial AI data essentially depends on the corporate reputation of the enterprise, which includes customer trust and engineering delivery capabilities. With the support of the two, enterprises can continuously get new projects to supplement first-hand working condition data.

"No matter how much capital the participants have, they need to start from scratch and collect data on site, so data is a very precious resource," Jia Chunying, Chairman of Srit, told 36Kr.

In response to this, Srit's solution is a two-wheel flywheel driven by real and synthetic data.

In fact, the absolute amount of data in the industry is not low, but what is really lacking is high-quality data. Such data needs to cover different scenarios, and the proportion of defect data is sufficiently low, so as to ensure the diversity and accuracy of machine learning.

Among the 48TB data collected by Srit, 17 types of scenarios are covered with defects accounting for less than 2%, which can continuously supply diverse and complete data raw materials to the model.

However, the problem with real data is that it needs to wait for a real scenario to actually occur, which will waste a lot of time. Therefore, Srit also introduces VR/AR teleoperation data for teaching, and then matches it with world model simulation synthesis to form a multi-dimensional digital twin as much as possible. Finally, this simulation scenario is fitted with the reflow data from real machine execution. The four modes of "vision, touch, force and action" are bound on the same space-time axis, and the twin data is optimized by production line data to realize that "the production line is a data factory". This is the core of its generative augmentation.

But the biggest challenge in this process lies in process understanding: whether a scratch constitutes a defect needs to be judged in combination with the position, size and customer quality standards — there is a gap of process experience between generating an image with a crack and generating a valid industrial training sample.

Srit has been immersed in automotive production lines since 2019, and gradually developed a methodology for digital processing of processes in the running-in process with customers, including understanding the principles behind the processes and the environment where defects occur. In the industry, automotive production lines are called "the pearl on the crown of industrial manufacturing", and the four major processes cover almost all manufacturing vertical categories. In this forging process of high intensity, high requirements and high standards, Srit has gradually grown into an expert in the four major automotive processes.

Nowadays, it only takes 10 real photos to start the large model to generate defect samples. Taking the integrated paint appearance inspection of "matte + glossy surface" implemented by leading automotive enterprises as an example, compared with similar solutions, its AI intelligent agent cuts the implementation time by half, and the recognition accuracy exceeds 98%.

No competition on humanoid form, focus on brain: widening the generation gap with "demo-focused" embodied robots

Similar to AI large models, embodied robots have also entered the deep water zone of commercialization. An industry insider told 36Kr that now investors no longer look at PPTs, but go to factories in batches to see robots working in the workshop, rather than performing pre-set processes in the exhibition room; the model of talking about vision and relying on demos to raise financing has come to an end.

Srit's robots have already been shuttling in customers' factories to load and unload goods.

Compared with the embodied robot companies on the current market, Srit does not compete on the body of robots, but focuses on the brain. It does not follow the trend of developing humanoid complete machines, but focuses on four types of functions: sorting, handling, assembly and quality inspection, concentrating R&D resources on data engines and embodied models, and cooperating with partners in hardware.

On the brain side, CogniMotor is responsible for task understanding, decomposition and planning; the world model SeerWAM rehearses before the action happens, judging how the environment may change and what interference may occur during execution.

In real implementation, all these actions must go through the execution gating of "scoring first, then releasing". The predictive decision-making of "thinking before doing" is the watershed between industrial implementation and laboratory demonstration.

However, the big-small brain and world model trained only with visual data cannot solve the "last centimeter" problem in industrial operations. After the robot makes physical contact, workpiece offset, contact deformation and assembly resistance often occur. These multi-dimensional physical quantities constitute the structural upper limit of the visual route.

In response to this, Srit's solution is "hand-eye-brain collaboration". In addition to visual data, Srit introduces the fusion perception of force-tactile and vision, so that the robot "has a sense of touch" at the moment of contact, and the force-tactile data after real machine operation finally flows back to the data engine.

This VTLA industrial force-tactile large model is the core of Srit's embodied model . Tactile sense is the only native perception that can directly map physical properties, which cannot be obtained from the Internet, and can only be collected from real production lines. Industry surveys show that there is still a 12-18 month window of opportunity in the direction of industrial force-tactile training and inference coupling. It can be said that force-tactile data is a blank category with zero supply in the industry, which is a missing lesson that pure visual players cannot make up for, and it is also the data asset of first-movers.

With the force-tactile data, combined with multi-modal data such as vision, language, sound and semantics, Srit finally builds a physical training and inference coupling link of "VLA action generation — world model rehearsal evaluation — execution gating — real machine execution", and vision, touch and force form a closed loop in real contact.

Behind the VTLA industrial force-tactile large model, it ultimately comes down to the density of high-quality customer scenarios. Srit has accumulated a scenario network of nearly 100 industrial terminals with more than half being leading automotive enterprises after 9 years of deep cultivation, and covers the four major processes of stamping, welding, painting and final assembly, with extremely high cellular scenario density. The effect has been reflected in the process indicators: the HaiBot has a dual-arm load of 20kg, and the accuracy of the industry's first set of four-modal intelligent operation deployed in leading automotive enterprises can reach 0.1mm.

In fact, before the product was launched, Srit concentrated its resources on the brain and tactile large model, which seemed a bit untimely. Because the sales of robot body is the most intuitive realization channel, while the large model requires a lot of time to train, and cannot see results in the short term.

Srit took the initiative to choose this more difficult path, and its confidence comes from the self-sustainment of its core detection business for many years. They do not rely on financing to survive, and their technology investment follows the industrial law rather than the capital rhythm.

For Srit, making the brain behind robots represents a greater ambition, that is, building a real industrial embodied ecosystem: instead of making a full-stack closed system around its own robot body, it opens up the models and capabilities, delivers them to different forms of robots and production scenarios, and plans to promote heterogeneous adaptation of different robot forms in 2029.

This is not a consumption war with the "general brain school", but a layered competition and cooperation. The general brain expands across industries horizontally, while Srit uses "VTLA + tactile world model" to realize the industry's first vertical industrial embodied brain with four-modal alignment, and finally releases the embodied capabilities to all walks of life.

In terms of business model, its five realization outlets including data services, intelligent workstations, model authorization, tactile simulation data and quality inspection share one data flywheel, so that every dollar of revenue is both income and data.

Four non-accelerable admission tickets

Industrial embodiment is actually not a new topic. Leading companies are competing to tell similar stories, and their robots have also completed loading and unloading operations at the MMIT station of mobile phone OEMs such as Longcheer.

But beyond demonstrations, industry is ultimately a long-term and high-potential track. Taking the automotive enterprise supply chain as an example, its certification cycle is generally as long as 2-3 years, and it is very cautious about partners entering the production line to collect data. Even if they enter the production line, it will take several years to form 48TB of real production line data and a 100,000-hour sample library like Srit. More importantly, the engineering implementation practice on the production line cannot be written into papers, which is an experiential asset that cannot be quickly acquired through paper learning. In addition, many companies lack the ability to sustain themselves, and in fact they have no ability to get the long-term admission tickets to industrial AI and embodied intelligence.

In the industry, some technical routes want to use simulation data to make up for the shortcomings, but the dimensions are single, and it is difficult to build a real multi-dimensional digital twin; while real machine data is expensive and low in migration, and the confidentiality of industrial scenarios compresses the collection channels. This also explains the dilemma of the industry: the industry is noisy, but real data is still scarce.

Srit is determined to take on the role of this enabler.

At present, open capabilities are becoming an important direction of physical AI. Nvidia is opening up the world model and training evaluation framework to the industry, which is the idea of platform-based technology supply. Srit also hopes to become such a core capability provider: transforming the process experience accumulated over a long period of time into a common base that empowers the industry.

More importantly, the market size ceiling of industrial AI and embodied intelligence is high enough. The automotive industry is the most stringent verification field — cycle time, accuracy, safety and certification all need to be screened layer by layer, and Srit's embodied POC for complex industrial scenarios is completed and accepted by leading automotive enterprises in only 7 days.

Graduating from the most difficult scenarios means getting a pass for the general industry. For Srit, vertical field is the entry point, not the boundary. The know-how of grabbing, gluing, handling and inspection accumulated in automotive production lines can be fine-tuned and augmented by the world model, and the adaptation cycle to new production lines can be reduced to several weeks. Srit's 9 years of industrial accumulation is being reorganized through large models, from one-time engineering delivery to reusable, licensable, cross-industry flowable intelligent assets.

For the industry, this is a more pragmatic intelligent path: AI grows from the place that best understands the process, turning the flexible manufacturing problem of multi-variety, small-batch and frequent line change into a deterministic task that can be undertaken on a large scale.

When these capabilities become the base for the intelligent leap of all walks of life, industrial embodiment is no longer the sales of robots, but the infrastructure upgrade of the manufacturing paradigm. The competition in the robot industry is shifting from "whose demo is more amazing" to "who can precipitate scenarios into sustainable compound interest assets".

What really defines the next decade is not the company that is best at telling stories, but the enterprise that precipitates time into barriers and achieves accumulated growth through long-termism. The hundred-billion-yuan deep vertical market of the automotive industry is just the beginning, followed by the trillion-yuan space of the general industry. Eventually, leading industrial embodied and AI large model companies will inevitably move towards an open general + vertical platform. Enterprises like Srit that have data assets, self-sustainment capabilities and technical know-how will eventually release their value intensively at the moment of industrial realization.