Stop fixating solely on output quantity, performance evaluation in the AI era must take these three levels into account.
Companies are ramping up their AI investments aggressively, yet the vast majority of performance appraisals still adhere to old standards, giving rise to the performance paradox: employees who overuse AI to boost output get rewarded, while those who verify and correct errors end up at a disadvantage. This paper proposes a three-dimensional evaluation framework that distinguishes three types of indicators: human contribution indicators, AI system and agent indicators, and human-machine collaboration system indicators, helping enterprises redefine good performance in the AI era and clarify the attribution of rights and responsibilities.
When work outcomes are produced collaboratively by humans and AI, how should "good performance" be defined? How can managers ensure that the pursuit of speed and efficiency does not come at the expense of human judgment and sense of responsibility?
At present, these issues have become critically important. Enterprises are making massive investments in AI, and managers are under huge pressure to prove that AI investments can generate returns. The annual "Data and AI Leadership" survey jointly conducted by Randy Bean (one of the authors of this paper) and Tom Davenport released data earlier this year: 91% of enterprises are increasing their AI investment, and 99% of enterprises list AI investment as the top organizational priority. However, only 18% of enterprises stated that their AI investments have delivered highly quantifiable business value.
AI has improved processing speed and output quality in a large number of workflows, but the vast majority of enterprises have not redesigned their employee performance evaluation methods to adapt to this change. Combining the latest research and our practical experience in assisting enterprises to implement AI, this paper proposes a set of three-dimensional performance evaluation frameworks suitable for performance appraisal in the AI era, helping enterprises distinguish three types of indicators: human contribution indicators, AI system and agent indicators, and human-machine collaboration system indicators.
The Human Performance Paradox
At present, when evaluating employees' work using AI, enterprises still follow the traditional success measurement criteria: productivity, goal completion rate, and work efficiency. This gives rise to a paradox: employees who rely heavily on AI appear to be extremely productive; while other employees slow down to verify preset conditions, question AI output results, and correct errors, they appear to be inefficient precisely when they are creating the greatest value.
Under this evaluation model, managers are likely to reward employees for over-reliance on AI, but punish human judgment that can avoid major mistakes; they blindly pursue output quantity, ignore actual business results, and fail to see the core elements that truly drive efficient work.
We have already seen the real consequences of applying old performance indicators to a brand-new work model. When AI systems outperform humans on traditional indicators, some employees will develop a defensive mindset: a survey by an AI manufacturer shows that 10% of employees admit that they will tamper with data to deliberately lower the performance of AI. This phenomenon is not unexpected. When management mainly regards AI as a tool to cut manpower, employees will inevitably be under pressure to prove that they are better than AI on traditional key performance indicators, because they know their jobs are at stake.
But the mismatch of indicators is only half of the problem, even the relatively easy half to solve. The other dilemma is that the traditional performance management system itself has long failed. Gallup's 2024 survey shows that at a time when enterprises were just starting to adapt to generative AI, only 2% of chief human resources officers in Fortune 500 companies highly recognized that their company's performance management system can motivate employees to improve themselves; among the surveyed employees, only 20% thought performance evaluation is fair and transparent.
Up to now, the situation has not improved significantly. Deloitte released a report in January this year stating that despite high hopes for automation, 84% of enterprises still have not redesigned jobs, workflows and career development paths around AI; only 30% of enterprises have set performance incentives for employees to use AI reasonably. Most enterprises' talent strategies only focus on employee training. Enterprises train employees to adapt to new work models, but the evaluation methods remain unchanged. The problem with this approach is that it takes for granted that making employees do their original jobs faster is still the right goal.
The combination of wrong indicators and ineffective performance management is creating a trust gap. Gartner's data shows that only 1 in 50 AI investments can bring transformative value, and only 1 in 5 can generate measurable return on investment, but enterprise CEOs still have high expectations for AI-driven growth. In contrast, frontline employees can clearly feel the huge gap between management's slogans and actual implementation. The gap between the board's optimistic expectations and the frontline reality is the root cause of performance management failure.
What Constitutes Good Performance in the AI Era?
Traditional performance indicators are built on the following work characteristics: independent, repeatable tasks that are fully responsible for by individuals. AI can automatically complete a large number of such tasks: draft writing, content summarization, translation, sample code writing, and preliminary analysis. A large-scale field experiment covering more than 750 knowledge workers shows that after using OpenAI's GPT‑4, employees' work speed increased by more than 25%, and task completion rate increased by 12.2%; within the coverage of the model's capabilities, the quality of output solutions is also significantly higher.
However, this study also reveals that the evaluation method that only focuses on output quantity hides huge risks. When the task slightly exceeds the capability boundary of AI, the probability of subjects who can use AI to output correct answers is 19% lower than that of people who do not use AI. Researchers named this phenomenon "jagged technical boundary": AI performs outstandingly in some tasks, but may suddenly fail when facing related tasks that seem highly similar. If the assessment only focuses on output speed and does not verify the accuracy of the results, it will prompt employees to hand over work to AI, break through its capability boundary, and hope that AI will not make mistakes.
Other studies have proved that traditional indicators will misjudge the real source of value. One study found that AI assistance can increase productivity by about 14% on average, but the benefits are mainly concentrated on less experienced employees. After senior practitioners use AI, their efficiency only increases slightly, and the quality also declines slightly. Experts create value by diagnosing problems, identifying extremely special scenarios, verifying prerequisites, and guiding the team when not to trust tools; most managers will not actively weigh this cost.
There are also more hidden and long-term risks. Controlled experiments show that repeated exposure to biased AI output will amplify people's inherent biases in perception, emotion and social judgment. People are often unaware of the influence of AI on themselves and are more likely to be coerced by biases. A colleague who seems to "speak with data" may be subtly acquiring the biases inherent in the model. The standard performance system has no mechanism to discover such problems at all.
A major hidden danger related to this is the proliferation of "low-quality quick works": using AI to quickly produce a large number of low-quality works, which are full of redundant content and various errors, with extremely low actual information value. Such results will consume a lot of time for the recipients, reduce colleagues' evaluation of the producers, and damage team cohesion. Enterprises that only assess output quantity and not accuracy are not only misjudging performance, but also actively undermining performance. Many managers have reported that it is difficult to digest the massive content produced by employees: the work that used to be completed in a week can now be done in a few hours, but quality control still requires experienced managers with strong judgment to invest time.
AI agents with autonomous execution capabilities further amplify the performance evaluation problem. In the past, most enterprises regarded AI as a tool for employees. Although only 11% of enterprises have put AI agents into production, after agents are integrated into the workflow, it brings what we call the collaborative performance problem. When the final result comes from a human-machine hybrid system, the evaluation framework must answer three questions at the same time: How is the performance of humans? How is the performance of the AI system? What about the overall effect of human-machine collaborative cooperation?
Based on research related to algorithm management and control, we believe that AI agents can influence human behavior through a management-like approach: recommend or restrict work content, record and evaluate performance data, trigger rewards or personnel replacement decisions. If evaluation and governance mechanisms are not set up for agents and employees respectively, no party can be held accountable. Once the agent evaluation form and the employee appraisal form are mixed together, both evaluations will lose their authenticity.
New Performance Measurement Framework
Combining ongoing research and enterprise cooperation practices, we propose a three-dimensional evaluation framework to measure human performance, AI systems and agents, and human-machine collaborative output respectively. Each layer is set with a small number of core indicators, reviewed frequently (monthly or quarterly), and directly linked to actual management actions: coaching, staffing, promotion, and tool governance.
I. Human Contribution Indicators
To evaluate the real work performance of people, it is necessary to weaken the output indicators that AI can easily boost, and focus on the core capabilities that AI cannot replace. The vast majority of positions need to pay attention to the following three types of capabilities:
1. Boundary Judgment
Can employees reliably identify scenarios where AI exceeds its capability range and make appropriate disposals?
- Escalation and Reporting Accuracy: If subsequent review confirms that the AI output is indeed defective, incomplete, or beyond the reliable capability range, the report is judged as "reasonable".
- Traceability Score: Whether the deliverable completely and transparently records the data source, generation time, and model used. A sampling ratio needs to be set in advance, for example, 20% of all deliverables are selected for review every quarter.
- Correction Quality Index: Corrections for which employees leave written explanations (output errors, outdated data, lack of context, etc.) are judged as "reasonable corrections"; unfounded improper corrections will be deducted points.
2. Collaborative Coordination Capability
Can employees use AI tools to improve the overall output of the team, rather than just improve personal efficiency, that is, empower others.
- Team AI Penetration Rate, supplemented by depth-of-use stratification: Occasional use (less than 2 times a week), regular use (2-4 times a week), deep integration (cross-task use in daily work). Focus on the distribution of personnel at all levels.
- Process Optimization Contribution Index, assign weights to the results of employees optimizing work processes: 1.0 point for newly built processes; 0.5 point for optimizing existing processes; an additional 0.25 point for sorting out documents and sharing. Statistics are made on a quarterly basis.
- Per Capita Output Ratio, track time series data. Compared with the absolute value, the change trend is more referenceable: Under the premise that the team size remains unchanged or even shrinks, the rise of this indicator represents excellent overall empowerment effect.
3. Learning Iteration Speed
Facing the iteration of tools, processes and systems, can employees adapt quickly. Building a systematic AI capability training system by enterprises is an important signal: mastering AI is a career development path, not a job threat.
- Tool Implementation Lag Duration: The number of working days between the official release of the tool and the first recorded effective use by employees. The lower the value, the faster the learning speed, which can be counted as the average of individuals or teams.
- Training Implementation Conversion Rate: Track observable behavior changes. Within 30 days after employees complete the training, the employee himself or his manager records at least one actual implementation application case.
- Experiment Attempt Rate: "Experiment" is defined as a recorded attempt to test new usage methods, prompt word strategies, and tool combinations, regardless of success or failure. This indicator helps to create a learning culture that tolerates failure.
II. AI System and Agent Indicators
Relying only on model accuracy and running time is not enough to complete the evaluation, especially for autonomous agents. It is recommended to focus on three major dimensions:
1. Goal Achievement Rate
Does the AI agent complete the preset tasks within the established constraints?
- Task Completion Rate: The output meets the preset acceptance criteria, no manual modification and resubmission are required to be counted as task success. Continuously track different agents and different task types to identify performance degradation.
- Error Rate: Errors are classified by severity level (serious/medium/minor), and the root cause is recorded. A very low overall error rate but frequent serious errors is a major risk that must be reported separately.
- Goal Deviation Index: It is especially critical for agents that perform multi-step execution. Identify situations where agents blindly optimize quantifiable proxy indicators and deviate from real business goals.
2. Interpretability and Traceability
Can we restore the decision-making logic of AI and retain verifiable evidence?
- Output Traceability Rate: Each output must be associated with the input data used, model version, and retrieval context. Conduct sampling audits according to the established sample, for example, spot check 20% of the output in each period. In highly regulated industries, too low this indicator will directly bring compliance risks.
- Reproducibility Score: Extract part of the output and re-run it with the retained parameters. The same input can stably get the same output, which means the system is auditable. If the result deviation exceeds the set threshold (for example, >5%), a traceability check is initiated.
- Interpretation Adequacy Score: A review team composed of domain experts, compliance personnel or end users judges whether the reason description attached to the agent's output is clear enough according to the unified scoring standard to ensure the unified review standard.
3. Reporting and Disposal Quality
When the agent reaches its own processing boundary, can it make appropriate disposal. This is the most important dimension of autonomous agents, which determines whether manual supervision can be truly implemented.
- Extreme Scenario Transfer Accuracy: Define boundary scenarios in advance by risk category (high-risk decisions, new inputs, compliance triggers). Only when the case is escalated and transferred to the corresponding processing level within the specified response time limit can the transfer be judged as correct.
- False Reporting Rate: Excessive reporting by agents will increase operational burden and consume team trust. This indicator is used to balance the sensitivity and accuracy of the reporting mechanism.
- Manual Intervention Support Rate: Measure whether the system architecture truly supports manual intervention, not just paper functions. Regardless of the frequency of occurrence, failure or obstruction of manual intervention is a major system failure.
III. Human-Machine Collaboration System Indicators
Evaluate whether the combined output of human and machine is better than that of humans alone or AI alone. This dimension is of high value, but requires a large amount of data support. Enterprises with sufficient resources can use the following indicators:
- AI Takeover Rate: Count the proportion of work that is completely out of manual intervention. A higher takeover rate is not necessarily a bad thing; but if it is accompanied by a decline in quality and an increase in errors, it means that value creation is turning into value loss.
- Complementary Collaboration Index: Count the proportion of cases where human intervention really generates value, such as discovering errors, reconstructing problems, and supplementing business judgments. A high score represents real collaboration, not simple stamping and release; a low score may either represent extremely strong agent capabilities or passive coping by employees, and the two situations require completely different solutions.
- Value Attribution Ratio: Split the total output value into the contribution of AI execution, and the contribution of human judgment, screening, and correction. This indicator is difficult to implement, but of great strategic significance. As AI capabilities iterate, this ratio will continue to change; long-term tracking can clearly see whether enterprises are cultivating the core capabilities of people or gradually weakening human value.
Redesign the Performance Management System
Enterprises do not need to reconstruct all processes at once. Choose a business process that has been changed by AI as the entry point: customer service, sales plan writing, system drafting, software development, financial reporting, and perform the following four steps:
1. Sort out the whole process and mark the capability boundary of AI. Disassemble the process into different task types: conventional repetitive type, highly judgment-dependent type, high-risk type, and strong relationship attribute type. Locate the "jagged boundary": tasks where AI output seems reasonable but often makes mistakes. This is the premise of all work.
2. Reconstruct the indicators first, and then modify the assessment form. Replace the output quantity with a small number of leading indicators: traceability verification pass rate, reporting and disposal quality, rework reduction range, customer experience change, team empowerment effect.
3. Distinguish between development coaching conversations and compensation decisions. If the same set of AI-generated data is used to coach employees and calculate compensation, employees will regard evaluation as monitoring. When employees clearly know which data is collected,