AI has maximized innovation efficiency, but why are good ideas becoming fewer and fewer?
Even with the same large model, the innovation outputs of different teams can be vastly different. AI can greatly improve processing efficiency, but it tends to amplify inherent human cognitive biases and organizational prejudices. Innovation is not entirely a technical issue. Recognizing the capability boundaries of AI and maintaining the connection between humans and the real world is the key to breaking the deadlock of innovation.
Nowadays, almost all innovation teams have access to the same tools: identical large models, and prompt libraries that are largely similar. However, the innovation outputs they generate vary dramatically. Some teams are experiencing an explosion of creativity, while others are overwhelmed by a large number of homogeneous, unmemorable ideas that read as if they were produced by the same person. The problem does not lie in which model is selected, but in the fact that generative AI acts on the deep-seated inherent human bottlenecks in the innovation process. In the face of AI, these bottlenecks will show completely different, even diametrically opposite changes.
This is the core viewpoint of a working paper commissioned by the *Journal of International Marketing Research*, whose authors come from Harvard Business School, Wharton School, Northwestern University and Columbia University. The study does not dwell on "what generative AI can do", but raises a more long-term valuable question: what cognitive, social and organizational constraints have long hindered innovation, and how will AI interact with each type of constraint?
To answer this question, the study integrates recent empirical research on generative AI and innovation, covering directions such as creative diversity, AI-simulated consumers, algorithmic screening of ideas, and post-launch review analysis, while combining mature research results in the fields of creativity, decision-making under uncertainty, and consumer behavior accumulated over decades. For the four major innovation stages described below, the study uniformly asks three questions: What is the underlying human constraint? What impact will it have on this constraint when generative AI is used directly without deliberate design? If targeted process reconstruction is carried out, how should it be implemented?
This unified analytical framework allows us to use the same set of judgment logic to examine various bottlenecks with completely different appearances — for example, the bias of the review committee towards new things, or the situation where consumers cannot clearly state what they want before experiencing the product in person. Does AI eliminate this bottleneck, or does it make the problem worse imperceptibly?
Innovation managers have not yet fully realized the importance of this issue. Facts have proved that some constraints can indeed be truly resolved by generative AI; for others, AI seems to be helping, but in reality it exacerbates the problem; for even more constraints, even if more powerful models are iterated in the future, there is still nothing that can be done about them.
Taking it for granted that AI can fix innovation pain points
Is itself a misunderstanding
The vast majority of discussions about AI and innovation default that AI can eliminate frictions in the innovation process in all aspects. The common assumption is that AI can completely replace part of the work of human innovators, generate more ideas, complete tests faster, reduce research costs, and accelerate experience review.
The analytical framework in this article refutes this view. After sorting out various bottlenecks along the four major links of the entire innovation process — idea generation, idea screening, consumer insight, and market learning, the study finds that: In most links, if AI is used in a simple and crude way without deliberate process optimization, the default effect of AI is often to exacerbate the original bottleneck rather than eliminate it.
The underlying reason is: the vast majority of innovation bottlenecks are essentially human problems, not technical problems. People will stick to inherent cognitions; review teams will form circle preferences to determine what counts as "potential"; consumers will not know what they really want until they see the product with their own eyes. Generative AI uses massive amounts of data of past human outputs as training materials, so it is easy to replicate the inherent behavior patterns that cause these bottlenecks, only faster and on a larger scale.
Let's look at the specific situation below: how generative AI amplifies the inherent human bottlenecks in each link of innovation, and the corresponding solutions.
1. Idea Generation: Overly biased towards conventional ideas
Idea generation is the process of conceiving alternative solutions for new products, new services and new functions. An outstanding and extraordinary idea is far more valuable than dozens of ordinary and qualified ideas. Therefore, the goal of this link is not to pile up a large number of acceptable solutions, but to dig out truly breakthrough and alternative ideas. In essence, idea generation is an exploration of looking for exceptions, not an assembly line operation that pursues efficient output.
If people directly use large models during brainstorming without adding structured constraints, two problems will arise at the same time. The large model itself tends to output the most common answers in the statistical sense — after all, its training logic is to continue writing texts based on probability. When people read such conventional outputs, their own thinking will be anchored, making it difficult to come up with alternative ideas that could have sprung up originally. Thus a two-step thinking narrowing is formed: AI first gives relatively conservative ideas, and then further limits the boundary of human thinking.
More seriously, this will evolve into a systematic problem. The training datasets of most generative models are highly convergent. When independent teams of different companies use AI to carry out so-called "creative work", they will easily converge to a few types of design directions in the end. Individual production efficiency has improved, but the creative diversity at the entire market level is declining.
To solve this problem, what needs to be modified is the model, not human thinking. Using chain-of-thought prompts to guide the model to complete the first draft first, and then explicitly ask it to iterate in a bolder and more differentiated direction, has a very significant effect on AI itself. But this method cannot help humans improve their own thinking quality. Once human thinking is anchored by the initial ideas given by AI, simply telling yourself to "broaden your mind" is basically ineffective, and you will even be further coerced by the output of AI.
2. Idea Screening: Biases towards the sophisticated solutions output by large models
At the idea screening stage, enterprises need to decide which ideas get financial support and which are put on hold, which is usually completed by a stage review committee or a project portfolio review team. This link requires judgments made under high uncertainty: the market has not yet given feedback, and many ideas themselves have not been fully implemented. The conscious or subconscious evaluation criteria of the review team will ultimately determine all the products that the enterprise subsequently launches.
Reviewers have always had biases towards novel ideas: original ideas have a stronger sense of risk and it is difficult to predict their results, so reviewers will unconsciously give priority to choosing safe and familiar solutions. Generative AI amplifies this bias in a brand new way: the proposals generated by AI are naturally fluent in writing and complete in structure, and fluency is easily mistaken for the quality of the solution. A delicately expressed idea will get a higher score, but this has nothing to do with whether it is a better idea in itself. The corresponding solution should start from the process and system, not rely on technology: all submitted solutions are uniformly formatted and then sent to reviewers for scoring, and it is hidden whether the idea comes from a human or AI.
This is not the only trap where generative AI quietly amplifies bottlenecks in the screening link. There is also a more hidden risk: if you use the historical pass/veto decisions of the review committee to train a large screening model, you do not eliminate the process bias, but solidify the bias on a large scale and package it as an objective algorithm result. A recent field experiment has drawn a more alarming conclusion: when reviewers get the suggestions given by the large model, accompanied by the reasons written by the model, they are more likely to directly stamp and reject the proposal, but their ability to distinguish whether the AI judgment is right or wrong has not improved. The reasons given by the model do not assist human judgment, but directly replace human thinking. This phenomenon is particularly prominent in edge pending projects that most test human judgment.
The corresponding intervention idea is completely opposite to the system that most enterprises are building: remove the explanatory text given by the large model. The same experiment shows that retaining only the black-box recommendation result of AI without accompanying reasoning reasons leads to higher final decision quality than retaining the model's explanation.
3. Consumer Insight: Research gets faster, but real answers become fewer
After the ideas pass the screening, they will be polished into complete solutions with specific functions, pricing and positioning. Testing the solutions is essentially to find out the real demands of customers, which is the preference measurement stage: Enterprises have long relied on questionnaires, interviews and conjoint analysis research to transform product concepts into judgments that hold up to market demand.
For practitioners who want to use AI to simulate consumers, the conclusions in this part of the paper are very realistic. For example, if you ask a large model to play thousands of different consumer groups and simulate their responses to price changes, you will get a smooth downward demand curve, but this result has clear and serious distortion. The model is not trained based on real people. It will not treat price changes as experimental variables, but only treat price as a clue to interpret the context — high price will be interpreted as "high-end" rather than "the same product simply increases in price". As a result, the downward demand curve formed after the price increase of real consumers may become a horizontal line or even upward sloping in AI simulation.
AI digital humans (digital twins) built based on rich behavioral data of real users will perform better, but still lack a key link: real consumers will make a large number of irrational choices, and these choices largely determine whether new products can be successfully launched. People will be reluctant to abandon sunk costs and will over-amplify the pain of switching products. In contrast, simulated digital humans have a much higher probability of making rational choices than real humans, which means that they will systematically underestimate the real frictions that lead to the failure of new products.
There is no perfect solution to this problem yet. Digital humans are not useless, and some research questions can use them for initial screening; but when it comes to key behaviors that determine the success or failure of new products, simulated results cannot be trusted. Whenever decisions involve sunk costs, switching costs, and information narrative frameworks, a more secure way is to retain real human research: lead user groups, ethnographic field surveys, to observe the real behavior of users, rather than only looking at the choices predicted by the model.
4. Communication and Market Learning: Confirmation bias still exists, but the arguments are more convincing
Communication and market learning cover all links after the product is launched: promote the large-scale implementation of innovations, interpret market feedback, and guide the next round of decision-making. Product adoption is not only an individual behavior, but also a social process, driven by early users and word-of-mouth communication. Market feedback needs to be collected, prioritized, and correctly interpreted to be of reference value.
After the product goes online, massive amounts of feedback will emerge, far exceeding the processing capacity of the team: comments from various platforms, social discussions, customer service tickets, return records, forum posts. Generative AI can indeed exert value in this link: it can digest massive amounts of information, perform clustering and summarization, with a processing scale that is unreachable by analyst teams, and the quality of extracting structured customer needs is already close to the level of professional analysts. Among the four links, this is the only stage where bottlenecks are truly alleviated by AI.
But problems also follow. AI will extract hundreds of user demands every day. Which ones are worth implementing, and which ones are just noise from individual users? The most important signals for innovation, such as emerging usage scenarios and certain types of user groups that cannot accept the product all the time, are exactly the information that frequency statistical summaries are most likely to bury.
A deeper hidden danger arises after the feedback is delivered to the internal stakeholder teams. Enterprises have always tended to interpret market feedback from a perspective that is favorable to themselves. AI itself will not bear the performance pressure of last year's product projects. In theory, it can mark a large number of bad reviews and good reviews equally objectively. The fluent and confident summaries output by AI, while facilitating information integration, also make it convenient for teams to selectively intercept content to support their existing views; especially when the way of asking questions already implies the answers the team expects to get, AI will cater to human ideas. Whether AI's integrated feedback ultimately leads to a more objective review, or only makes the inherent rationalization more convincing, remains the most important unsolved problem under this analytical framework.
Core Conclusions
After sorting out the above four stages, a core principle can be extracted: It is not how advanced the AI model is that determines whether AI plays a boosting or destructive role, but the inherent mechanism of the bottleneck itself.
For those bottlenecks that are essentially information processing, massive information, and process-related (too many solutions to be reviewed, too many feedback texts to be read, and formats that hide real quality), as long as the system is deliberately designed, generative AI is very suitable for solving such problems.
In contrast, if the bottleneck is rooted in real experience, human irrational behavior, and organizational interest games, AI will be difficult to work. No matter how far the model iterates, it cannot solve these realities: for product categories that they have never been exposed to, consumers cannot clearly express their preferences; after managers have invested two years of political capital in a failed project, they will always interpret the data selectively, no matter how high the quality of the data itself is.
This brings managers a more valuable tool than the AI application list: before introducing any new tools or new workflows, conduct a diagnosis first. Is the obstacle to progress an information problem, a human judgment problem, or an organizational incentive mechanism problem? From this perspective, every innovation stage has clear implications:
1. Idea Generation: Unconstrained generative AI will compress the creative space rather than expand it. Optimization should target model prompts, not adjust people's mentality.
2. Idea Screening: AI cannot eliminate the review team's bias towards new ideas. It will either amplify the bias with the help of sophisticated outputs, or encode and solidify the bias on a large scale, unless the entire process is deliberately reconstructed.
3. Consumer Insight: Simulated consumers are an efficient initial screening tool, but they will underestimate the irrational real obstacles that truly determine the life and death of new products.
4. Communication and Market Learning: AI solves the problem of overload of feedback information after launch, but it can also provide the team with more gorgeous arguments to defend the decisions that have already been made.
The most alarming viewpoint of the paper is not aimed at a certain bottleneck, but focuses on the long-term evolution of the organization. When AI takes over more and more work of idea generation, screening, and consumer simulation, frontline employees lose the opportunity to practice personally and accumulate judgment. There will be a large number of people in the enterprise who are skilled in using AI output results, but lack the intuition to identify when AI outputs go wrong.
Spreading this situation to the entire innovation department will form an automation trap: every independent AI tool performs well when tested separately, but the complete feedback loop gradually breaks away from real users. Idea conception, screening, testing, and interviews are all handed over to the model, and the model training materials are the outputs of other models, which are no longer anchored to the real market. Everything gets faster, but is disconnected from the real foundation.
The countermeasure given in the paper is not to resist the implementation of AI, but to clarify which links must retain the direct contact between humans and the real world. In-depth contact with leading users, ethnographic field surveys, and letting people take responsibility for key decisions. These are not nostalgic remnants of the AI era, but load-bearing walls that ensure the entire automation system will not be disconnected from objective reality.
Technology will continue to iterate, but those bottlenecks — cognitive anchoring effect, evaluation based on status and identity, the gap between consumers' verbal expression and real behavior, have remained stable for decades. Building innovation strategies on underlying mechanisms rather than betting on model capabilities is more resistant to the test of time.
Julian De Freitas, Ayelet Israeli, Gideon Nave, Artem Timoshenko, Olivier Toubia | Text
Julian De Freitas is an Assistant Professor of Marketing at Harvard Business School. Ayelet Israeli is the Marvin Bower Associate Professor of Business Administration at Harvard Business School and co-founder of the Customer Intelligence Lab at the Digital Data Design Institute at Harvard Business School. Gideon Nave is an Associate Professor of Marketing at the Wharton School. Artem Timoshenko is an Assistant Professor of Marketing at the Kellogg School of Management, Northwestern University, whose research focuses on the application of AI in marketing analytics and consumer insight. Olivier Toubia is a Professor of Marketing at Columbia Business School and a recognized leading scholar in the field of quantitative marketing.