HomeArticle

Does all the time saved by AI coding go to waste on debugging? Claude official: Let AI do its own testing and modification.

新智元2026-10-09 07:41
Return your time to the things that truly matter.

Just imagine:

You used Claude Code to build a small app from scratch, and integrated an AI customer service for it.

You originally wanted it to automatically answer basic user questions like "how to log in" and "what content is included in the membership service", so that you could completely leave it unattended.

As a result, the customer service can indeed chat, but you have become its full-time training partner:

If its answer is inaccurate, you have to add a prompt immediately; if its explanation is too verbose, you add another restriction.

What's even more frustrating is that every time you modify a prompt, you have to re-ask all common questions, for fear that it will overlook one aspect while fixing another, and the originally correct answers will be broken after you just solved one problem.

Thus, a lot of the time saved by using AI to write code is spent back on rounds of manual debugging.

To solve the trouble of repeated debugging, Anthropic officially released an optimization guide on September 28.

The core idea of this guide can be summed up in one sentence: Claude can also do the work of repeated testing and modification.

Let it take over more work, and you can stop being a miserable tester for many rounds, and devote more energy to supervising the AI:

Set rules for it, clarify what counts as a good job, and then check whether it really meets the requirements.

Turn "Don't Talk Nonsense" into Specific Test Questions

To let the AI correct mistakes in a targeted manner, you must first make it clear "what counts as a correct result".

Usually we always hope that the AI customer service is "more reliable", but this requirement is too abstract.

If you can align the requirements with the problems that users will actually encounter, the standards will be much more specific.

For example:

When the user cannot find the export button, can it provide the correct click path?

When the user asks about functions that have not been developed yet, will it seriously make up nonsense?

When encountering a refund problem, can it handle it according to clear rules, and not make random promises when it is not sure?

The first official entry provided is: /claude-api build-eval.

Simply put, it can help you turn these specific requirements into test questions, and build a set of repeatedly runnable acceptance test question banks for small apps connected to Claude.

You can submit the problems that often make the AI make mistakes to it, but the question bank should not only include "error scenarios", but also cover those most common and ordinary problems.

Otherwise, no matter how many partial and difficult questions you accumulate, you may not be able to accurately test the daily experience of ordinary users.

After selecting the questions, you also need to make it clear "how to score".

Requirements like "professional", "intelligent" and "friendly" are too vague, and you need to turn them into more specific measurement standards.

For example, what information must be clearly stated, what mistakes must not be made, and what level counts as passing.

Official email classification example: Test samples need to be confirmed by users.

Claude will help you draft the corresponding scoring rules, and you need to confirm whether the answers that get high scores according to this set of rules can really solve the user's problems.

Some results can be checked directly by programs, such as whether there are missing key fields; for open-ended problems with multiple reasonable answers, you can also let the model score according to clear rules.

But the scorer itself needs to pass the verification first.

Take several already scored answers, compare them with your own judgment, and see if there are situations where "correct answers are deducted points" or "problems not solved are marked as passed".

After confirming that the scoring is reliable, conduct regular spot checks.

For developers, the biggest benefit of this step is that you can finally turn the repeated complaint of "why did you answer wrong again" into a set of continuously runnable automated checks.

Every time you modify the prompt later, you can run it again to see what has been fixed this time and whether other functions have been broken.

Modify and Test by Itself

Roll Back If the Performance Gets Worse

After the question bank is built, the next step is to use the second official entry: /claude-api hillclimb.

Hillclimb can be understood as "climbing the hill", which allows the AI to try to optimize little by little to see if there is progress.

But this KPI also needs to be specific, and you need to clarify what it is allowed to modify.

For example, to make the customer service's answers more accurate, it is allowed to modify the prompts; or on the premise of maintaining the answer quality, try to save API costs, and it is allowed to adjust the model configuration.

After receiving the task, Claude will view the wrong questions used for iteration, propose one modification per round, and then re-run the evaluation:

Modifications will only be retained if there is indeed improvement under the established goals; if the performance regresses after modification, it will be rolled back.

According to the official documentation, both entries require Claude Code v2.1.259 or higher.

Prompts, skill files, tool descriptions and model configurations can all be optimization objects, but you need to define the scope of what exactly can be modified.

You may worry: Will the AI become a test-taking expert in exam-oriented education? That is, it only memorizes the few questions you give, and starts talking nonsense when the question is phrased differently?

The official has indeed considered this point.

This process will separately reserve a part of the questions as "acceptance test questions" that are not disclosed in advance.

The model responsible for proposing modifications cannot see the content of these questions. After each round of modification, these questions will be used for testing to see if the effect is really improved, or only works for the practiced questions.

If the score of the practiced questions increases, but the score of the "acceptance test questions" does not improve, "overfitting" may occur: it gets better and better at doing familiar questions, but makes no progress when facing a new batch of questions.

In this case, Claude will roll back the modification of this round; if the performance becomes worse after modification, it will also roll back.

Re-test after each round of modification, keep or roll back according to the result.

This can help detect the problem of "only being good at familiar questions", but it does not mean that you can rest easy from then on.

Real problems not covered by the question bank may still cause it to make mistakes.

Official Customer Service Case

Cost Reduced to About 1/5

Anthropic demonstrated the effect through an internal customer service evaluation.

On 14 reserved test tickets that did not participate in guiding the modification, the decision accuracy of the final configuration increased from 78.6% to 90.5%, an increase of 11.9 percentage points, and the model invocation cost was reduced to about 1/5 of the original.

Adjust model, thinking intensity and prompts to improve accuracy and reduce cost. The figure shows the performance of the iteration set.

However, this is a result under a specific configuration and a specific small sample, which does not mean that after connecting this process, your application can also save 80% of the cost.

The model invocation cost compared here cannot be regarded as the total cost of the entire project.

But it is still very attractive for individual developers who pay for API costs out of their own pockets.

After all, if a cheap model always gives irrelevant answers, forcing users to ask follow-up questions repeatedly, and finally requires your manual intervention, it may not be cost-effective.

If simple questions also use expensive configurations, and generate a large section of unnecessary explanations, you may also waste money.

Now you can test the answer quality and invocation cost together, and let the AI try more cost-effective solutions on the premise of meeting the quality requirements.

However, the optimization itself also consumes model invocation costs.

Set a reasonable budget for experimental costs first, and then decide how many rounds the AI can run, which will be more reassuring.

Give Time Back to the Truly Important Things

When this automated process is running, you can reduce repetitive prompt repairing, and focus back on the product itself.

First, gain insights into users.

The operation that you take for granted may be completely impossible for a first-time user to find the entry. What you need to do is to add these real confusions to the question bank, so that the next optimization of AI will be more practical.

Second, control the standards.

No matter how polite the answer is, if it does not solve the problem, can it be counted as passing? If the meaning is exactly the same but only expressed in different words, will the scorer misjudge?

If there are loopholes in the scoring standards themselves, the AI may work for a long time only to get higher scores, while the user's problems remain unsolved.

Email classification evaluation: View the score of each question on the left, review the model input and answer on the right to check whether the scoring is reasonable.

Finally, make trade-offs.

Will simplifying the answers omit necessary steps? Will pursuing lower cost sacrifice accuracy? These key trade-offs related to user experience still need you to make the final decision.

The original intention of integrating AI into the application is to save effort.

If it eventually evolves into you accompanying it to modify answers around the clock, it will be putting the cart before the horse.

When Claude can take over more repetitive debugging work, ordinary developers can spend more time polishing their small tools and small applications, and push those shelved ideas forward again.

References:

https://claude.dev/blog/automating-eval-design-and-hillclimbing/?utm_source=chatgpt.com

This article is from the WeChat official account "AI Era", Author: ASI Revelation, Editor: Yuan Yu, published with authorization from 36Kr.