Three flaws of AI podcasts: fabricating unsubstantiated content, removing supporting evidence, and skipping lines during narration.
All content streaming platforms are pushing audio-enabled features, and AI tools are also focusing on the "one-click podcast generation" format. But the result of this "one-click generation" may not be as good as expected.
Producing in-depth articles relies on solid data, distinct stances, and logically coherent narratives that connect from start to finish. If AI-powered podcast conversion cannot even preserve these fundamentals, then after you press the one-click button, the podcast will no longer contain your original article content.
To verify this, I selected two pieces of material from our previous articles and handed them over to Coze to reconstruct the content and record the podcast.
The first piece of content is a financial analysis of less than 600 words, which I call Segment A. It is densely packed with hard figures, and the test targets the fidelity of AI to accurate information.
The other piece is more than 300 words long, which I call Segment B. It features an extreme stance and dense metaphors, and the test checks whether AI dares to retain sharp judgments.
If these two checks cannot be passed, then "one-click generation" is a false proposition.
To match different styles, we chose serious dialogue, entertainment crosstalk, and academic pretentiousness as test directions. We first use Segment A for a preliminary test to see how the same set of hard figures will be handled under the three styles respectively.
Segment B is only tested under the serious and crosstalk styles, to dig into whether the extreme stance can be retained in serious dialogues and whether it will be diluted into gags in entertainment crosstalk. The academic style was not tested on Segment B, because Segment B has few technical terms and cannot reflect the unique missing-word problem of the academic style.
We conducted multiple rounds of tests in total. On the first day, the audio voices of Segment A in crosstalk mode were completely messed up. On the second day, the voices of Segment B in crosstalk mode were almost all female voices and the content was automatically compressed, both days ended in failure. On the third day, I ran the test again with the same instructions to see if this style is usable at all.
The benchmark for comparison is the "Listen to Full Text" feature of WeChat articles. "Listen to Full Text" provides zero rewriting, just like news broadcasts with no emotions. This state of "no mistakes but no highlights" is the status quo we face every day. Using Coze to make podcasts was originally intended to break this mediocrity, but it turns out that things are not that simple.
Day 1: Three Styles for Segment A
I first tested the serious dialogue style. I created a new conversation, entered the instructions, and Coze output the script and audio at one go, which took only five minutes. The script was expanded from 570 words to nearly 1800 words, and all 19 hard figures were retained.
I cross-checked with the original text and found problems. The phrase "not even in the same magnitude" was changed to "the gap has widened to more than four times". "The growth rate dropped to 10.42%" was changed to "an almost halving decline". A more hidden problem is that after a sharp judgment, AI added a sentence on its own: "Finally, I'm not saying WPS will die, its user base is still solid." This sentence does not exist in the material, it is a buffer added by AI itself.
The pronunciation of numbers in the audio is accurate, and all problems lie in the script layer. This time, Coze changed the vague judgment to an unsubstantiated precise number, and added a buffer layer for the sharp conclusion.
I switched to the entertainment crosstalk instruction. Different from the serious version, this time Coze generated the script first, and produced the audio after confirmation.
The script smoothed out many precise numbers, the decimals were removed, and 15.78% was completely missing. Eight numbers were blurred, and one was lost directly.
But a more serious problem than numbers is the voice line. The script has two roles A and B, A is female voice and B is male voice. As a result, almost all lines of B in the audio are read by female voice, and lines of A are also mixed with male voice. The voices are completely messed up, I can't continue listening at all, and can only check word by word against the script. The two-person podcast turned into a chaotic monologue.
The gags written by AI ("comparing monthly salary with Musk", "small shop vs Walmart") are also too safe, without real talk show rhythm. The confrontation I required was not implemented, and the result became a Roast Show where A dominates and B plays the supporting funny role.
I switched to the academic pretentiousness instruction. Similarly, the script is generated first, and the audio is produced after confirmation. The script upgraded "writing documents and making spreadsheets" to "Document Authoring, Spreadsheet Manipulation", the density of technical terms is full, all 19 numbers are retained, even with decimals.
But after the audio was generated, I checked sentence by sentence and found missing words. The word "Formatting" in the script is completely missing in the audio. The sentence "The data flywheel spins, Model Iteration accelerates" also lost half of its content. When there are too many technical terms, words and sentences start to be omitted.
The listening experience is a mix of Chinese and English popping out, which does look pretentious, but you can't tell whether it's real professionalism or fake professionalism. What you hear is only fragments, and the complete sentences are long gone.
Day 2: Two Styles for Segment B
The crosstalk voice line of Segment A was completely chaotic on the first day. When testing Segment B on the second day, I added constraints to the instructions: "A is male voice, B is female voice, strictly assigned, no mixing allowed".
I first tested the serious version of Segment B. The material was changed to ".md is killing .wps".
The word "killing" was retained. But AI added a buffer in front: "This sounds harsh, but the fact is...". The keyword was not deleted, but a soft cushion was added first.
The confrontation quality is better than that of Segment A, with 12 rounds of head-on debates, B defended .wps throughout the process and did not defect. The voices are also correct, A is male and B is female, no cross-over.
Looking back: the serious style has the highest information fidelity, both numbers and stances are retained, but AI will secretly add extra content beside them.
I switched to the crosstalk instruction to test Segment B. The script was generated first, with three pages of content. Then Coze prompted "The script is too long, please simplify it", and generated the audio after automatic compression.
What was compressed? The content "Nine and a half out of China's one billion workers are using WPS" is gone. The content "In the past, you spent 200 yuan to buy WPS membership for typesetting capability. Now AI can finish typesetting for you in 3 seconds, is this 200 yuan still worth it?" is also gone.
The ending also changed from neither side giving in to "the winner is .ai". It completely became a muddy, non-committal conclusion.
What's more troublesome is that AI made this decision in the background without asking you at all.
In terms of voice lines, the same problem still exists. Almost all are female voices, only the sentence "Okay okay, let me put it another way..." is male voice. You can't tell at all that two people are having a conversation.
Day 3: Retest of Segment B in Crosstalk Style
The crosstalk tests failed for two consecutive days, on the third day I ran the Segment B crosstalk test again with exactly the same instructions. This time the script is 750 words, Coze did not prompt "too long", and generated the audio after confirmation. The voice lines are correctly assigned, A is male and B is female, no cross-over. The content is complete and not compressed.
This time it finally looks decent, you can tell two people are talking, it has crosstalk vibes, but it's not that funny.
Two out of three crosstalk tests failed. On the first day, no constraints were added for Segment A, the voices were completely messed up; on the second day, constraints were added for Segment B, but it was still almost all female voices; on the third day, after adding constraints, it finally worked normally. This style itself is unstable, and the third successful run is more like getting a rare drop in a gacha game.
Note: The academic pretentiousness style was only tested on Segment A, Segment B was not covered.
Summary
Putting all these together, the serious style has the highest fidelity, but AI will secretly add extra content, fill in blanks for vague descriptions, and soften sharp stances. The crosstalk style has the most diverse problems, numbers are smoothed out, arguments are compressed, and voice line assignment failed twice out of three tests. The academic style suffers from missing words and sentences as soon as there are too many technical terms.
This test only used Coze, but I saw the same revision tendency recurring. It translates "not in the same magnitude" into a specific multiple, adds "this sounds harsh" in front of "killing", and directly compresses one-third of the content when the script is too long. These three operations appeared in all six audio works in this test.
But in-depth content precisely requires "accuracy, sharpness, and clear stance".
So what can we do?
If you have to use AI to make podcasts, here are several practical suggestions based on actual tests:
First, separate the script and audio generation process. Lock the script first, manually check every sentence before generating the audio, never use the "one-click" method.
Second, it is best to generate long texts in segments. This can at least reduce the risk of the script being automatically compressed.
Third, single-voice reading plus manual proofreading is the solution with the highest fidelity at present. We sacrifice expressiveness for content completeness.
There are two directions not tested this time, but worthy of verification in the next round: whether adding more "prohibited items" in the instructions can reduce AI's arbitrary filling of content; whether writing key numbers in text form (such as "five point two nine billion") can avoid rounding errors.
In these six tests, generated podcasts are never just simple reading. The serious version creates pseudo-precision, the crosstalk version loses numbers and arguments, and the academic version has missing words and sentences.
So this is still a trade-off. Before choosing a style, first figure out which part of the content you want to retain most.