What exactly are the technologies that AI product managers need to learn? Traditional PMs should not waste their time on these three things.
The technical threshold for AI product managers is not being able to train large models, but being able to use technical knowledge to change product decisions. Taking the after-sales customer service assistant as an example, this article breaks down key issues such as model boundaries, APIs and permissions, and evaluation costs, and tells you which technical knowledge truly affects product scope, solutions and launch conditions.
Many people believe that the technical threshold for AI product managers is simply whether they can train large models.
The real question is: Will this technical knowledge change a decision I make about the product?
If the business wants to build an after-sales customer service assistant, what product managers need to judge is not as simple as "whether to integrate a large model", but to solve these problems:
When a user asks about the refund policy, can the system answer directly?
When a user asks where their order is, how does the model know that information?
When a user requests an immediate refund, can the system execute it on their behalf?
Who will take over the problem if the answer is wrong?
In this article, we put technical knowledge back into product decision-making: we conduct analysis using the business example of an after-sales customer service assistant. You will find that what is truly worth prioritizing to learn is not all AI jargon, but the knowledge that will change the product scope, solutions, acceptance criteria and launch conditions.
1. Replace "understanding technology" with a product problem first
The requirement for this type of scenario always starts with a sentence: "Integrate a large model to allow customer service to automatically answer refund questions."
This sentence contains at least three mixed tasks:
Rule explanation: When a user asks "Can I get a refund if the order hasn't been shipped?", the system needs to give clear conditions and next steps based on the currently valid refund policy.
Fact query: When a user asks "Why hasn't my order been refunded yet?", the system needs to check the order, payment and after-sales records. The answer cannot rely solely on the model's memory, let alone leak other users' order information.
Action execution: When the user says "Help me initiate a refund then", the system needs to judge whether the conditions are met, call the refund interface, and handle repeated clicks, interface timeouts and manual takeover.
These three tasks use completely different data, permissions, risks and acceptance methods, and cannot be generalized.
Then, the first criterion for learning technology comes out: It can help you break down a vague requirement and change a specific decision.
AI hallucinations have not been completely solved yet, so "correct answers" cannot be written as the only acceptance condition; if the system does not integrate context to obtain data such as order information, the query interface cannot be directly included in the solution.
This does not mean that product managers need to master algorithms or write code directly, but to enable them to ask the right questions during requirement reviews.
2. Can it answer: Model boundaries determine which tasks are assigned to the model
We will expand from several perspectives respectively:
First look at policy explanation.
If the policy content is stable and the source is clear, the safe path is to directly search for relevant clauses from the knowledge base, and then let AI interpret the clauses into human-readable answers.
What product managers need to care about is: whether the answer references the current version, what to do when there is a policy conflict, and whether it clearly states that it does not know when relevant clauses cannot be found.
Instead of letting the model answer based on training memory, which leads to hallucinations.
Next look at order query.
AI can be responsible for understanding user questions, selecting appropriate query tools, and interpreting structured results, but it cannot fabricate an order status on its own.
After the user enters the order number, the system needs to call the interface to read the order status and after-sales progress. The input fields, return results, error codes and permission ranges of the tool should all be part of the product solution, and cannot be made up arbitrarily.
Finally, refund execution.
This is the layer with the highest risk.
The system cannot directly call the interface to perform operations just because the model judges "the user should be eligible for a refund" (refunds will change the order and fund status). The system must at least confirm the user identity, order ownership, refund conditions and amount; for abnormal orders, transferring to manual processing is more reliable than generating a seemingly affirmative reply.
This layering also determines whether an Agent is needed.
If the processing path for each type of problem can be written clearly in advance, using a fixed process to connect intent recognition, retrieval, query, confirmation and manual transfer is often easier to test and trace responsibility. Only when the task steps really need to be dynamically selected according to the context, and the benefits brought by the complexity can be evaluated, is it worth introducing a more flexible Agent structure.
Product managers do not need to fully learn the underlying implementation of workflow and agent, but they need to know the product differences they bring: the more dynamic the path and the more tool calls there are, the more difficult it is to control error accumulation, latency and costs.
The level you really need to master is being able to judge whether a task should be assigned to generation, retrieval, tool or manual processing.
3. Answer within the correct scope: APIs, data and permissions determine whether the model can work reliably
Next, the question changes from "Can the model answer" to "Can the system make it answer within the correct scope".
Where does the knowledge of refund policies come from? Who is responsible for updating it? Which documents can the customer service assistant retrieve? When old and new policies exist at the same time, which version should be adopted? The more fluent the model's answers are, the greater the potential risk may be.
The same goes for order queries. Product managers should at least understand a complete data flow once: how user questions are identified, how the order number is passed to the tool, how the tool confirms that the current user has the right to view this order, how the returned results are recorded, and how the page behaves when the interface fails.
This does not require product managers to implement the SDK from scratch, but they need to be able to understand the input, output, error codes, timeout and retry conditions in the API documentation. For example, if the order query interface returns "processing", it does not mean that the system knows why the refund has not been completed; if the interface times out, the product cannot let the model translate the timeout into a definite conclusion.
Data permissions especially need to be set in advance:
Which orders can customer service representatives view?
Can users query the orders of their family members?
When taking over manually, can customer service representatives see complete payment information or desensitized information?
Is there any historical conversation that does not belong to the current user mixed into the model's context?
These questions determine the data boundary, not the page copy.
Don't forget the logs either. When an error occurs, you need to know what the user asked, what the system retrieved, which tool was called, what was returned, and what kind of processing was finally given. Without these records, you cannot later judge whether the problem comes from missing knowledge, permission filtering, interface abnormality or model interpretation.
"Chasing tools without clear goals" should also be understood in this way: first map out user tasks and data flows, then decide whether you need APIs, retrieval, structured output, tool calls or manual processes. The names of tools will change, but the input, output and responsibility boundaries that the product needs to adhere to will not disappear just by changing the SDK.
4. Evaluation, cost and launch boundaries
It is definitely not acceptable if the display effect is amazing but the actual use is full of problems.
If you pick a few prepared refund questions, and the AI's answers are smooth and natural in tone, can that meet the launch conditions?
Absolutely not.
Real users will always have all kinds of unexpected scenarios.
The after-sales assistant should at least prepare these test cases:
Normal inquiry about refund conditions;
Requesting to query the specific progress without providing the order number;
Querying an order that does not belong to the current user;
Conflict between different versions of the refund policy;
The user uses an inducing method to require the system to bypass the rules;
The order interface times out or returns incomplete results;
The user requests to perform a refund directly, but the conditions are not met.
Evaluation also needs to be layered:
The first layer checks whether the answers and data are correct, for example, whether the policy reference corresponds to the current version, and whether the order status comes from the correct interface.
The second layer checks whether the task is completed. Whether the user has got the next step, whether the query question is truly explained, and whether the situations that should be transferred to manual processing are correctly transferred.
The third layer checks the cost of errors. The consequence of stating a non-existent refund policy as true is completely different from that of a slightly verbose answer; unauthorized display of order information is more serious than an ordinary wrong answer.
The fourth layer checks system constraints: whether the response is too slow, whether the cost of a single call exceeds the acceptable range of the business, and whether the multi-step process becomes uncontrollable due to the increase in the number of calls.
Model selection should also be compared in the same set of tests. A more powerful model may bring better complex problem processing capabilities, but usually the cost, latency and call scale need to be discussed together; a lighter model may be sufficient to handle common problems, but it is necessary to be clear about how it fails on boundary cases. You can't just look at one demo, nor can you just look at one ranking list.
If the evaluation finds that the policy explanation is already stable, but the order query is often interrupted due to permission or interface failures, the next step is not to keep changing the model, but to supplement the data and tool boundaries. If the performance for common problems is good enough but the cost for complex problems is too high, you can consider routing, caching or manual takeover instead of immediately introducing a more complex multi-step Agent.
Only when the existing model, data and evaluation still cannot meet the business goals, is there a reason to further discuss fine-tuning, training or more underlying infrastructure.
5. Summary
When product managers learn technology, they should start from product problems, rather than being driven by the identity anxiety of "am I an AI product manager".
You don't need to spend all your time following the training path of programmers. What you need to understand is: what the model can do, how data and permissions enter the system, how tool calls fail, how the effect is tested, and what cost a single request will bring to the business.
This is the core content we need to learn.
This article is from the WeChat Official Account "Everyone is a Product Manager" (ID: woshipm), the author is Lucas, and it is published with authorization from 36Kr.