From Data Middle Platform to Knowledge Middle Platform: How Enterprise-level Ontology Enables AI to Understand and Manipulate Digital Production Environments
In 2026, policies promoting the deep integration of artificial intelligence and industries have been intensively introduced at the national level. In April, the Ministry of Industry and Information Technology and the National Data Administration jointly issued the Notice on the Joint Implementation of the 2026 "Modulus-Data Resonance" Action, focusing on 20 key industries to promote "guiding data with models and empowering models with data". In June, the Implementation Plan for Promoting the Action of Building High-Quality Industrial Datasets released by the National Data Administration explicitly proposed to strengthen the construction of datasets such as knowledge bases, knowledge graphs and ontologies for agent applications. As a result, ontology has evolved from a technical concept to an industrial infrastructure, and the implementation path of enterprise-level AI is also facing redefinition.
This means that the integration of AI and industries is entering a new stage from "competition of model capabilities" to "competition of data and knowledge infrastructure". For large models to truly enter enterprise production environments, a set of semantic infrastructure that enables AI to understand business semantics, associate complex calculations, and drive system operations is required.
In the process of carrying out enterprise-level ontology construction with industry customers including China Unicom, energy and power, oil and gas, and finance, Hai Zhi has repeatedly pondered two questions: What exactly is ontology? What problems can ontology solve in enterprises?
This article will expand from four dimensions: the gap between ontology and enterprise data, the upgrade from data middle platform to knowledge middle platform, the relationship between ontology and graph, and the form of AI-oriented multi-modal database, sharing Hai Zhi's thoughts and practices in enterprise-level ontology construction.
01 Why can't large models "understand" enterprises? The gap between ontology and enterprise data
The extension from communication networks to data networks, and then to semantic networks, is an inevitable trend of industrial digitalization.
At present, Chinese enterprises, especially large state-owned enterprises like China Unicom, have completed digital transformation. Digital transformation brings an essential change: all businesses of enterprises can be described by data, with a large data volume and high complexity of data calculation. The vast majority of these data exist in the form of two-dimensional relational tables (such as order tables and customer tables), rather than one-dimensional, linearly arranged text data in sequence.
When industrial large models try to enter enterprise production environments, the first problem they encounter is this gap in data form.
In the previous round of industrial large model construction, many companies used enterprise corpus texts for large model training, but the results were minimal. The reason is that large language models are mainly trained with text parameters, while the real business data of enterprises largely exists in relational tables and the calculations behind them.
Therefore, rather than saying that large models are "hallucinating", it is better to say that they cannot well understand a large number of relational tables and the actual calculations behind the relational tables.
If we want large models to understand these relational tables, a lot of text interpretation work similar to Prompt or Embedding (text vectorization, that is, converting text into a machine-computable digital matrix) is required. However, the cognitive sequence of human beings is that language was invented first, and then data — data is essentially a further abstraction and structuring of the reality described by language. It is basically impossible for large models to understand relational tables and complex calculations in reverse from text through this reverse engineering.
From a technical perspective, the root of this problem lies in: the core capability of large models is semantic understanding and generation, while the operation logic of the enterprise digital environment is built on deterministic computing. The association, aggregation, filtering and other operations between relational tables, as well as the business rules and constraints formed thereby, constitute a formal system parallel to natural language. There is an essential tension between the probability generation mechanism that large models are good at, and the accuracy and verifiability required by this formal system. Large models "generate" the most probable answers based on probability distribution, and this mechanism is inherently uncertain; while the formal computing systems of enterprises require results to be accurate, reproducible and verifiable. This is the fundamental reason why large models are prone to errors when directly connecting to enterprise data.
Therefore, enterprises must answer a core question: how to effectively organize massive amounts of data and underlying complex calculations so that AI can perceive them. Today, the modern management of all enterprises relies on digital operations, and the problems that digital operations can solve are very intuitive: all businesses can be quantified, analyzed, explained and traced. This is difficult to achieve with pure text methods.
The starting point for Hai Zhi to build ontology is to connect AI, on the basis of its existing semantic environment, with the digital environment built by enterprises earlier. Let AI recognize the digital production environment through these semantic understandings, and even take over and operate the digital production environment. This is what enterprises need most.
This also leads to the core positioning of ontology: one end of the ontology connection is people (who rely on natural language to express business intentions), and the other end is machine systems (which rely on structured data and deterministic computing to run). Enabling people and machines to reach a consensus on "what exactly the same thing is" is the core problem to be solved by ontology.
02 From data middle platform to knowledge middle platform: what is managed is not only data
When building the data middle platform in the past, one core problem to be solved was: enterprises have a large number of business systems, but it is impossible to reach a consensus between these business systems, nor is it possible for a programmer to fully understand the code of all business systems. Therefore, a solution is widely adopted in the industry: extract the data of the underlying business systems, build cross-system data consensus externally, and then reconstruct cross-system business services.
This was the logic of building the data middle platform in those days.
Today, this logic is undergoing an evolution: the enterprise data middle platform is being upgraded to a knowledge middle platform.
From the perspective of architecture evolution, the core task of the data middle platform is to solve the problems of "where the data is" and "data availability", and its basic units are tables and fields. On this basis, the knowledge middle platform further answers "what the data means", "what the business rules are" and "how to execute actions".
The essence of this upgrade is to superimpose a "semantic control layer" on the original data asset layer: the data middle platform is responsible for opening up the physical access of data, while the knowledge middle platform is responsible for opening up the business meaning and execution path of data, so that AI can convert natural language problems into verifiable, executable and traceable semantic tasks.
The task of the knowledge middle platform is also to manage a large number of digital production environments and systems of enterprises. But the content of management is no longer just data, but to clearly explain:
- What kind of data is in each system;
- What kind of calculations can be performed;
- What kind of business operations can be performed based on each calculation result;
- What relevant business regulations executed by the systems exist in the business texts mapped by all business systems.
- This set of content is the real knowledge of the enterprise's digital engine.
No matter it is called a knowledge middle platform or a knowledge base, its essence is the management platform that carries ontology in the future. It can not only fully manage all digital businesses of enterprises (including businesses in unstructured semantics), but also the management boundary has far exceeded the scope of traditional data middle platforms.
Only in this way can AI understand the entire digital production environment of the enterprise, and then it is possible to talk about connecting all systems based on AI Coding to complete tasks automatically.
At the engineering implementation level, the construction methods of structured systems and structured data in various industries have been fully studied at present. The biggest real problem encountered is: whether structured and unstructured data need to be connected, and how to connect them?
From Hai Zhi's practical experience, because Hai Zhi's goal is to let AI operate the digital environment after understanding semantics, the construction path of ontology should be derived from structured to unstructured, not the other way around. The reasons are: unified semantics must be unique and deterministic, and as business systems of enterprises, the uniqueness of their data is naturally better than the descriptive text scattered in documents.
Moreover, what we need to solve today is the problem of operating systems: if the system itself does not have this function, there is no need to understand it in the ontology. Therefore, we should prioritize building unified semantics in the structured environment, and then use limited ontology generation technology similar to F1A (First Principles Alignment, that is, the limited ontology generation technology oriented to first principles — deduce from the most basic and certain business facts, and then generate ontology step by step) to align the document content with it.
After unstructured data is incorporated, it has two main applicable usages in the ontology: first, use unstructured documents as a knowledge base for retrieval; second, these text information can provide effective corpus supplement when complex calculation analysis is carried out in the digital environment. This is also the necessity of incorporating unstructured data.
In addition, it is often difficult to accurately locate between the business side and the data side, but there are always two types of calculations in enterprises: one type is OLAP, Online Analytical Processing, that is, the calculation used for business data analysis and report generation, to answer "what happened"; the other type is OLTP, Online Transaction Processing, that is, the calculation used for daily business operations and transaction records, responsible for "what is happening and how to deal with it".
When building ontology, AI needs to take into account that in the initial stage, it is only based on the core analysis calculation of data, such as business operation analysis business; in the future, it will extend to transactional calculation of operation connection between systems, such as services including supply chain management and customer service management.
From the perspective of technical architecture, traditionally, OLAP and OLTP each maintain a set of data models — the data warehouse layered model (oriented to analysis scenarios, organizing historical data by theme and hierarchy) serves analysis scenarios, while the transaction model (oriented to business processes, emphasizing real-time read-write consistency) serves business processes, and there is a lack of a unified semantic bridge between the two.
The value of ontology is that it does not simply superimpose semantic labels on the data model, but supplements the previously missing behavior semantics (what this action means), dynamic process modeling (how this process evolves over time) and business rule constraints (under what conditions can/cannot be executed), so as to truly open up the path from OLAP to OLTP, and realize the penetration from data analysis to business execution.
But OLAP and OLTP are not contradictory in essence. The key is that the two must follow the same set of unified semantic specifications, and the remaining problems mostly belong to the engineering implementation level.
03 Ontology and Graph: Unified Semantics and Graph Engineering
Hai Zhi started with knowledge graphs, so customers often ask: What is the relationship between ontology and graphs?
From the construction method of OWL (Web Ontology Language, an international standard language for defining and constructing ontologies), it is no doubt that ontology is also based on graph theory. W3C (the international standards organization that formulates Web standards for the Internet) clearly states in the OWL standard document: Ontology defines the terms used to describe and characterize a certain knowledge domain, including computer-available definitions of basic concepts in the domain and the relationships between concepts, so that knowledge can be shared and reused across systems and applications.
When building ontology based on a single scenario, it is not necessary to use a graph database to carry and manage it. Relational tables are acceptable, and vector indexes are also acceptable. The selection depends on the query and calculation requirements of specific scenarios.
A complex problem faced by Chinese enterprises today is: when more and more scenario-specific ontologies are built, will the data become more and more scattered just like directly building ODS (Operational Data Store, which refers to the temporary storage layer that directly aggregates the original data of business systems) in the past, and eventually become an asset management problem? Asset management itself is not difficult. The real difficulty lies in how to avoid fragmentation — that is, different scenarios build their own ontologies, which cannot be mutually recognized or reused.
Therefore, a consensus has gradually formed: just as building the DWD/public layer in the data warehouse in those days (unified here as the "data warehouse public layer", that is, the shared layer formed by unified cleaning and standardization of original data in the data warehouse) is to make semantic constraints and consensus. Today, in ontology construction, we may also prioritize building the unified layer of "graph semantics".
In other words, no matter which mutually recognized business scenario, which type of mutually recognized business activity, or which specific business object the data, calculation and operation belong to, an agreement must be reached at these levels.
Only based on unified semantics can we anchor all the service knowledge of data, calculation and operation. This is the same as the logic of the data warehouse. Therefore, for enterprises oriented to agents, large models or artificial intelligence, the core is to turn enterprise knowledge organization and construction into a graph knowledge project, that is, Graph Engineering.
Previous articles have systematically explained Graph Engineering
Graph Engineering is a knowledge network architecture oriented to solutions. If Harness Engineering solves the problem of "how to harness models", then Graph Engineering solves the problem of "what the models should trust" — to build a credible, inferable and evolvable structured world model for the models.
In Hai Zhi's definition, the overall evolution of Graph Engineering and AI engineering (Prompt Engineering → Context Engineering → Harness Engineering) is not parallel, but vertically nested: it is the deepening of Context Engineering to graphs, and also the technical core of the "control layer" and "memory layer" in Harness Engineering.
In other words: knowledge graphs are good at describing "which facts are related to each other" (what it is), while enterprise ontology further defines "what these relationships mean and what behaviors can be triggered" (how to do it).
Because enterprises will eventually find that if they want AI to make good use of it, data, calculations, operations, processes and decisions all need to be associated; after forming Agents and expert models, Agents also need to be associated, and expert models also need to be associated.
The association of all layers from top to bottom only relies on one thing as the hub, that is, unified semantics.
04 AI-oriented database: Graph, vector and dynamic ontology
In addition to thinking about business scenarios, another key problem that Hai Zhi is currently promoting is: assuming that the internal knowledge association of enterprises is very important and exists in the form of graphs, the form of past graph databases and future AI-oriented databases will undergo essential changes.
Last year, Hai Zhi undertook the "multi-modal database" expert project of the Ministry of Industry and Information Technology (the next-generation database technology direction oriented to unified storage and query capabilities for multiple data types — relations, graphs, vectors, documents, time series, etc.). This project is actually answering how to design the core service forms of "storage" and "computing" for enterprise knowledge operation in the future.
The first level is how to connect graphs and vectors: you can either "query the graph with vectors" (perform semantic similarity retrieval first, and then expand along the association relationship of the graph), or "convert the graph into a vector" (encode the graph structure into a vector representation) to participate in the fine-tuning of the graph attention model — a neural network model that can automatically identify which related nodes in the graph are more important.
This direction is highly consistent with the current technological evolution in the database field
Gartner predicted in the Database Market Guide that by 2026, the proportion of new core systems adopting multi-modal architecture will exceed 60%. It is widely recognized in the industry that multi-modal architecture is becoming the mainstream choice. The core idea of the multi-modal fusion database is to realize unified storage and query of multiple data models such as relations, vectors, graphs,