What is the real level of AI data governance?
Last month, I visited an enterprise for an exchange. The presentation materials they prepared were very polished: the headline read "Comprehensive Upgrade of AI-Driven Data Governance", with a picture of a governance dashboard, and the PPT was as long as 40 pages.
I asked three questions: What is the coverage rate of metadata annotation? They answered, roughly 20% to 30%. How many quality check rules have been configured? They replied, dozens of them, and more are still being built. How many links in the governance process have AI been applied to? They paused for a moment and said, currently it is mainly used for demonstrations to leaders.
As you can see, there is a huge gap between the "comprehensive upgrade" claimed in the PPT and the actual situation on the ground.
This article will tear off the fancy packaging and show you the real level of AI data governance — no overhyping, no unnecessary criticism, just stating facts. Only by recognizing the real level can we know where to go next.
01 First, make it clear: AI data governance consists of two separate things
Many people talk about "AI data governance" in a vague and confusing way, because they fail to realize that it is essentially two different matters.
The first one: Data governance for AI. When large models are introduced into enterprises, they need to be fed with data first. However, the data requirements of large models are completely different from those of traditional reports: reports require structured, consistently calibrated figures, while large models need unstructured data such as documents, logs, work orders, as well as semantics, context and knowledge associations. Can your original governance system handle this?
The second one: Using AI to do data governance. Treat large models as a governance tool: let them read metadata, generate annotations, identify caliber conflicts, automatically perform classification and grading, and write quality check rules. This matter is now being exaggerated as a miracle, and the actual effect will be detailed later.
These two matters have two respective real levels. Only by analyzing them separately can we see the situation clearly.
02 Governance for AI: the real level is far from passing
Let's talk about the first matter first. By presenting several industry-recognized data points, you will know what the current status is.
KPMG's survey shows that 66% of enterprises list "weak data foundation" as the primary challenge for large model implementation in enterprises — it is not that the model is not good enough, nor that the computing power is insufficient, but that there is no firewood under the data pot. MIT's 2025 report is even more straightforward: about 95% of generative AI pilot projects have not produced measurable returns, and the top-ranked reason for failure is exactly that the data is not well prepared.
Specific to the governance level, look at the actual status of several key indicators: the metadata annotation coverage rate of most enterprises falls in the range of 20% to 50% — which means that no one can clearly explain what more than half of the data tables contain, where they come from, and whether they can be used.
What's more critical is unstructured data. More than 80% of the total enterprise data is unstructured — documents, emails, drawings, work orders, meeting minutes. But in the past two decades, enterprise data governance has almost only focused on the less than 20% of structured data, which is exactly the opposite of what large models need most. There is a growing consensus in the industry: unstructured data governance is changing from an "optional task" to a "mandatory task", but the vast majority of enterprises have not even started to work on this task.
There is another more hidden short board: semantics. When large models give irrelevant answers after being introduced into enterprises, in most cases it is not because the model is not smart, but because there is no "enterprise semantic map" in front of it — a customer is called "Party A" in one system, "cooperation partner" in another system, and "transaction entity" in financial vouchers, no one tells the model that these refer to the same thing. The lack of a semantic layer is the deepest pit in current AI data governance, and also the most difficult lesson to make up for.
03 Using AI for governance: it is really useful, but don't deify it
Let's talk about the second matter. Using AI to do data governance is not a scam, and there are indeed several links where the efficiency improvement is real.
Metadata annotation. In the past, it took half an hour for a person to fill in the description of a data table manually. Now large models can read the table structure, sample data, combine upstream and downstream data lineage, and automatically generate business meaning descriptions, which increases the annotation coverage rate from 20%-50% to more than 90%. This has already been put into practice in existing data warehouse projects.
Caliber conflict detection. Let AI compare multiple tables and multiple indicator definitions, and find out "same name with different meanings" and "same meaning with different names". In enterprise data warehouses with tens of thousands of tables, manual work is almost impossible. AI does the preliminary screening and humans do the confirmation, which improves the efficiency by an order of magnitude.
Classification and grading. Security compliance requires every table and every field to be graded. Large models read the field names and sample values to initially judge the sensitivity level, and humans review and adjust it. A task that used to take half a year can be compressed to just a few weeks.
But — pay attention to this but — these three links have one thing in common: AI only does the work of "identification" and "preliminary screening", and the right of judgment and revision still belongs to humans. AI is completely powerless when it comes to disputes over business calibers, disputes over departmental rights and responsibilities, and historical data disputes. The most difficult part of data governance has never been a technical task, but an organizational task. In the foreseeable future, AI cannot replace this part.
There is another counter-intuitive fact: using AI for governance itself requires a pre-existing governance foundation. The premise for AI to identify the meaning of fields is that the data lineage link has been connected; the premise for AI to generate quality rules is that the quality framework has been established. Enterprises with a poor governance foundation cannot make good use of AI tools no matter how many they buy, just like a clever housewife cannot cook a meal without rice. That's why the same AI governance product acts as an accelerator in enterprises with a good data foundation, but is just a useless decoration in enterprises with a poor foundation.
04 A self-assessment form: what level is your AI data governance at?
Combining the two matters mentioned above, the real level of AI data governance can be divided into five levels. Check against it and see which level your enterprise is at — don't look at promotional materials, look at the actual status.
What is the real distribution across the industry? In one sentence: The vast majority of enterprises are stuck at L2, a few leading enterprises have reached L3, very few have touched L4, and L5 is currently mostly a concept in vendors' PPTs. If someone tells you that they have already reached L5, you can ask them: What is the inclusion rate of your unstructured data? To what extent have you built the semantic layer? Most likely the conversation will come to an awkward end.
05 Three misconceptions that separate publicity from reality
Why does "AI data governance" sound so hot, but the actual level is generally at L2? Because there are three common misconceptions in between.
Misconception 1: Treating procurement as construction. After buying an AI governance platform, people think they already have AI governance capabilities. The tool is lying in the server room, while the process is not changed, the personnel are not adjusted, the assessment mechanism remains the same, and the coverage rate has not increased by even one percentage point. The moment the procurement is completed, the real construction has just begun — the statement should be corrected the other way around: the completion of procurement does not mean the start of construction.
Misconception 2: Treating demonstration as capability. Vendors' demos always look amazing — they run on carefully prepared datasets. When you return to your own enterprise's data swamp, the same function will produce a mess of results. There is only one standard to judge real capability: run your own scenarios on your own data, and get your own results.
Misconception 3: Treating tools as governance. Governance is a systematic project that integrates organization, system, process, platform and assessment. AI tools are only a small part of the platform section. Expecting a single tool to solve governance problems is just like expecting a good car to solve your commuting problems — if the road is blocked, you don't have a driver's license, and you don't refuel, no matter how good the car is, it is just a decoration.
06 From L2 to L4: a practical and feasible path
The significance of recognizing the real level is not to be discouraged, but to find the right track. The path to move up from L2 is actually very clear.
Step 1: Solidify the structured stock governance. Don't let AI mess up your pace. Metadata coverage, master data quality, and unified indicator calibers — these traditional priorities are still the barriers you need to pass in the AI era. And now AI can help you speed up the process: using AI to do metadata annotation and caliber comparison is the most cost-effective first stop, with small input, quick results, and the team can get familiar with the operation at the same time.
Step 2: Promote unstructured governance based on AI application scenarios. Don't do unstructured governance for the sake of governance, follow the AI scenarios: if you first launch an intelligent customer service system, you can first govern customer service work orders and knowledge bases; if you first launch contract review, you can first govern the contract library. Build scenario by scenario, and make every governance investment bring visible business returns.
Step 3: Make long-term investment around the semantic layer. The unified definition of business objects, business terms and indicator calibers is the foundation for AI to "understand human language". There is no shortcut for this work, it requires the data team and business departments to align with each other again and again. But it is also a moat: enterprises that have built a complete semantic layer will be one step faster than others for every AI application.
07 Final remarks
Back to the enterprise mentioned at the beginning. Three months after our last exchange, they made an appointment for another visit. This time we didn't look at the PPT, but went directly to check the platform: the metadata annotation coverage rate was increased to 85% through AI preliminary screening, the number of quality rules was increased from dozens to more than 400, and the first unstructured scenario they selected was equipment maintenance work orders — because the production department had been urging for this for a long time.
The person in charge said a very practical sentence: In the past, we regarded "AI data governance" as a word for reports, but now we regard it as a work procedure.
What is the real level of AI data governance? It is that most enterprises have just stepped out of the mud of L2, that AI has proved its value in the marginal links of governance, while the main body of governance still needs to be built brick by brick, that publicity is two years ahead of leading practices, and practices are two quarters ahead of statistical reports.
Recognizing this point is not humiliating. What is humiliating is replacing the real level with the level shown in PPT, and even believing it.
This article is from the WeChat official account "Data-Driven Intelligence" (ID: Data_0101), written by Wang Jianfeng, authorized to be released by 36Kr.