Three Laws of Data: Value is never inherent in the data itself.
Let me start with a scenario I have encountered repeatedly in the past two years.
An enterprise spent millions of yuan on data governance. Standards were established, platforms were built, and historical data was cleansed once. The PPT at the acceptance meeting looked perfect: the data completeness rate rose from 62% to 96%, and the consistency rate increased from 71% to 94%, and the project even won an industry award.
Six months later, during a return visit, the business department said: "The data is still inaccurate, we still pull data from Excel on our own."
Where is the problem?
If you only look at the quality indicators, this project is a success. But if you shift your perspective, you will find that the clean data is stored in the third layer of the data warehouse. For the business department to get it, they have to submit a work order, wait for scheduling, and get approval across two departments. By the time they get the data, that promotion has long been over.
The data is of high quality, but far from the business scenario, and the best time to use it has passed. Among the three variables, only one has been taken seriously.
The value of data is never an inherent attribute of data itself.
It is jointly determined by three variables: quality, distance, and time. I call the rules that govern these three variables the three major laws of data.
I. Law of Data Inertia: Quality can only be guaranteed at the source
Have you ever thought about this: The data you are processing right now is no longer the original data it used to be.
Add version, age, and life cycle attributes to the data, and you will find that it will age, split, be copied, and be rewritten. From the moment a piece of data is generated, once it starts to circulate, it is no longer the original data.
This is the Law of Data Inertia: Once low-quality data enters the circulation process, the impact it generates is untraceable and unavoidable. No matter how much you cleanse or model the data later, you are just paying repeatedly for the previous "cutting corners".
I often compare data quality management to childbirth. Spending a little more effort at the beginning to give birth to a healthy baby is thousands of times more cost-effective than making up for it little by little after the child is born.
In terms of formula, the quality variable is Q ∈ [0, 1], which measures the degree to which the data is "healthy from birth". Note its feature: source pollution is irreversible loss. It is a variable that can only go down and is very difficult to go up.
The misunderstanding broken by this law is: "Data can be completely cleansed afterwards."
The typical symptoms are easy to identify: the same field is cleansed repeatedly but never gets completely clean; every time a governance project ends, everything returns to its original state half a year later. This is because the pollution source has not been addressed, you are just constantly diluting the pollution downstream.
Action guidance: Move governance actions to the business source, not just to the moment when data enters the data lake. Where a field is first entered in the system, that is where governance should take place.
II. Law of Data Mutual Exclusion: Versatility and value are inherently mutually exclusive
The more centralized data is, the more valuable it is?
On the contrary — the further away it is from the business, the less valuable it is.
Every time data is one layer further away from the business site, its value decays once. The formula is V(d) = V₀ · e^(−λd): V₀ is the value of data the moment it is generated at the business site, d is the number of layers away from the business site (business system → collection → cleansing → modeling → report → decision-making), λ is the context decay coefficient.
Why does the value decay? Because versatility and value are inherently mutually exclusive.
The cleaner and more versatile you make the data, the more business context that makes it valuable you lose. When it is abstracted to be completely separated from the business context, what remains is its form, and what is lost is its meaning.
This law has four inferences, each of which can find corresponding phenomena in enterprises.
Inference 1: The Middle Platform Paradox. The data middle platform is built for reuse, but the layer with the highest reuse rate has the lowest single-point value. This is determined by its design goal, not by poor implementation.
Inference 2: Rigid Caliber. The caliber further away from the business, the less people dare to modify it, and the less people dare to use it. Eventually it becomes a set of numbers that everyone can recite, but no one believes.
Inference 3: Cost Amplification. The cost of remedying data quality downstream is equal to the cost saved upstream multiplied by the number of links it passes through. This is the two sides of the same coin as the first law.
Inference 4: The Last Mile. 90% of the value of data is generated in the last mile back to business actions.
Action guidance: Return the right to define indicators to the business side. It is not to let the business side review the indicators defined by the IT department, but to let the business side define, modify, and take responsibility for the indicators on their own.
III. Law of Time Value of Data: Two golden periods, with a fault zone in between
The newer the data, the more valuable it is?
Wrong. It actually has two golden periods.
The formula is V(t) = A · e^(−αt) + B · (1 − e^(−βt)), the sum of the two terms forms a double peak.
The first peak is at the moment of generation — at that time, it can still change the result. Once the decision window closes, the value drops off a cliff, α is the business rhythm, the faster the rhythm, the more drastic the decay. For a real-time risk control model, if the transaction is already completed when the result is calculated, no matter how high its accuracy is, it makes no sense.
The second peak comes after a sufficiently long period of time — when you accumulate enough long history, it can be used again for modeling, attribution, and trend judgment, and its value rises for the second time. β is the historical thickness, the degree to which the value can be released depends on the question you want to answer.
The trouble is the middle period. It is neither fresh enough nor old enough — this is the value fault zone.
In the data warehouses of most enterprises, 80% of the storage cost is occupied by this data with the lowest value.
This is the most common and least noticeable type of waste.
There is another easily overlooked inference, which I call Continuity Premium: The value of historical data lies in "continuity". If one year of history is missing, the loss is far more than 1/N, because the law inference directly fails — you cannot use a sequence with gaps to judge the cycle.
Action guidance: Tiered storage should not only be designed according to "access frequency", but according to the time value curve. Three other things can be done right now: For real-time investment, first check how long the decision window is, only businesses with a decision window counted in minutes are worth building a second-level link; the archiving strategy should reserve a low-cost but reactivable recovery channel; do not delete historical data to save storage fees, that is the most expensive way to save money.
IV. Why Multiplication is Necessary: This is the most important sentence in the whole system
Combining the three laws, the value function of data is as follows:
V = Q × D × T
The expanded form is V(d, t) = Q · V₀ · e^(−λd) · [A·e^(−αt) + B·(1−e^(−βt))].
Please pay attention to the symbol in the middle: It is a multiplication sign, not a plus sign. This point is crucial.
If it is addition, you can use the long board of one item to make up for the short board of another. But multiplication does not work that way.
No matter how good the quality is, if it is too far away from the business, the value still tends to zero — a perfectly clean table stored outside the third layer of the data warehouse, no one uses it, its value is zero. No matter how close the distance is, if the decision window has passed, the value returns to zero — no matter how accurate the real-time calculation is, by the time you get the result, the promotion is over. The data is neither dirty nor far away, but stuck in the fault zone, its value is the lowest — this is the most hidden and most costly situation.
This leads to an uncomfortable inference: Data governance cannot achieve "breakthrough at a single point".
This is where many teams fail. They get 90 points in quality governance, but the data is three layers away from the business, Q = 0.9, D = 0.1, the product is only 0.09. The input-output ratio is extremely poor, and then they draw the conclusion that "data governance is useless".
This is actually the mapping of the Law of Minimum Factor (Liebig's Law) in the data field: The lowest of the three variables determines the upper limit of the overall value. If you pour money into the other two variables, the ceiling will not move.
V. How to Apply: Diagnose first, then invest
Four practical suggestions, follow the order to implement them.
First, diagnose first, do not initiate a project immediately. Score the three variables Q / D / T separately, and find the lowest one in your team at present. This step takes no more than two weeks, but can save a whole year's budget.
Second, the lowest variable is the only direction worth investing in at the moment. It is not "focused investment", it is "the only investment direction". Keep the other two variables from deteriorating.
Third, do not work on the three variables at the same time. That is equivalent to spreading the budget thin across three tables, and none of the tables has enough resources.
Fourth, re-diagnose every six months. Because the lowest variable will change. After the quality is improved, the distance often becomes the new bottleneck — this is actually a good thing, which means that the previous stage of work was done correctly.
Quick reference for the three laws, it is recommended to save it:
Final Words
After talking about the three laws, what they actually describe is the same thing: The value of data is never inherent to data itself.
It is a function of three variables: quality, distance, and time. All the work of governance is to make the three variables simultaneously move in a favorable direction.
Back to the enterprise mentioned at the beginning. Its problem is not that it is not willing to invest, but that it only invested in one of the three variables, and then expected the overall value to triple. There is no such good thing in a multiplication-dominated world.
So next time when someone asks "Is data governance useful or not", I suggest asking back first: Which variable in your team is the lowest right now?
If you can't answer it, it means you haven't started yet.
Which of the three variables is the lowest in your team? Let's talk about it in the comment section.
By the way: I recorded a 40-second audio version for each of the three laws, which has been posted on the WeChat Channels. If you want to get a quick overview first, you can go there to listen; if you want to see the derivation and implementation actions, they are all included in this article.
This article is from the WeChat official account "Data Driven Intelligence" (ID: Data_0101), author: Wang Jianfeng, published with authorization from 36Kr.