10,000-Word In-Depth Article: Where Should We as Humans Go Amid the AI Music Wave?
I
If you have used leading AI music platforms for creation recently, you may have been awestruck by their overpowering, disruptive intelligence.
Even at the current stage, AI music platforms can easily generate high-quality works with rich layers, delicate arrangements, stunning timbres and distinct styles, whose appeal and professional level are no less than mainstream commercial content polished meticulously by professional teams.
While some may argue that AI music still has various shortcomings of one kind or another.
But for one thing, the gap between its current state and the ideal form is only a temporary status quo limited by cost issues and development stages, rather than a technical boundary permanently locked and impossible to break through. For another, every coin has two sides, and evaluation depends entirely on the perspective of the observer. The perfectly accurate rhythm and extremely precise pitch not only make up for the imperfection of human physical control, but also reduce the auditory discomfort caused by errors.
Some so-called "shortcomings" of AI music that are sometimes criticized are precisely the "advantages" that only AI can achieve while humans cannot, not to mention that introducing fuzzy deviations and random perturbations is by no means difficult for computers.
Due to factors such as cognitive illusions, self-protection and dignity defense, people tend to develop a certain degree of resistance (such as fear, anger, distrust, etc.) to things that may change existing path dependencies or affect vested interests, and prefer to prioritize finding rational explanations that are easier for them to accept. "AI music lacks human warmth", "AI creation cannot truly express human emotions and feelings", "AI arrangement has no unique characteristics and memorable points of its own" — is this really the case?
Before Deep Blue defeated the world chess champion in 1997, professional chess players were confident that AI could not beat humans at chess; later before AlphaGo defeated the world Go champion in 2016, professional Go players were also confident that AI could not beat humans at least in Go. From homing pigeons to smartphones, from ox carts and horse-drawn sedan chairs to autonomous driving, from barter trading to digital currency, from knot-tying record-keeping to quantum computing, from ghost and god worship to particle physics, from shaman exorcism to life engineering.
Every human technological revolution that drives the leap of productivity and transforms production relations is always accompanied by cognitive rifts, limited imagination, ethical disputes and institutional disconnections at the initial stage, but the general trend of reforming the old and innovating the new has never been easily changed as a result.
Perhaps, deconstructing the craft of music creation from the perspective of extrapolating to the future can help us see the road ahead through the fog more clearly at this crossroads where the AI music track is booming rapidly.
What is the essence of music creation?
On the surface, we master some tools that can generate physical vibrations, and learn some application rules that make vibrations show a sense of rhythm as time changes. To abstractly design a continuous arrangement of vibrations through certain methods is what we call music creation.
Digging deeper, none of our organizational ideas for sound sequences can come out of nowhere. They either come from past individual experience accumulation, such as various sound materials we have heard intentionally or unintentionally; or come from logically explainable mathematical methods, such as disassembly, reorganization, connection, movement, reordering, analogy or randomness.
Digging even deeper, after meeting basic survival needs, the human brain will have excess energy and information that spill over naturally. Driven by the subconscious, they automatically combine and reconstruct, forming novel symbolic expressions. This is not only the comprehensive result of the joint action of genes and the environment, but also a certain advanced strategy that ultimately points to human survival and continuation.
No matter which level it is, machines can achieve it more comprehensively, faster and more extraordinarily than humans.
For humans, thinking out the permutation and combination of notes consumes certain mental resources; for machines, almost negligible computing power consumption can render inexhaustible and vast inspirations. For humans, the total amount of music information received is limited by the objective length of life and the depth of memory; for machines, little storage capacity is enough to store the entire global inventory of released music across all eras. For humans, there is always an upper limit of physiological function in playing musical instruments and controlling voices; for machines, the speed and precision that are far beyond human reach are nothing more than trivial byte operations.
"Possibility thinking" is regarded as the core of creativity, but under the First Principles, no matter how complex and unique the tuning system, chord progression, scale mode, rhythm and beat, musical form and structure that humans can conceive, machines can fully "think" of them through engineers' codes and prompters' descriptions. No matter how exquisite and tricky the dynamic fluctuation, element feature change, sound modification skill and humanized processing that humans can achieve, machines can reproduce them through methods such as data training. It is only a matter of whether the development priority of targeted algorithms is high enough, and whether the cost expenditure and resource allocation of vertical features are put on the agenda.
Humans can imitate, learn and remember; machines invented by humans can even better imitate, learn and remember. One may ask, in front of AI with high-speed computing, massive throughput and large-scale parallelism, for the sound frequency information integration and pattern development called "composition/arrangement", does the human brain with only 20W power still have any remaining advantages?
Some people say: AI music (still) cannot replace human creativity.
But what on earth is creativity?
It is most likely not a supernatural special function or a magical gift from an omniscient deity, but a global phenomenon formed by the joint action of multiple components in the complex human brain system. Modern neuroscience research has gradually revealed that creativity is a complex brain function that can be understood. It stems from the synergy of multiple networks in the brain, the dynamic balance of multiple neurotransmitters and the flexible conversion of multiple thinking modes, and can be artificially improved through systematic acquired training and environmental optimization.
Coincidentally, AI is also using a series of similar models and strategies to integrate and realize its own ingenious ideas, and may activate stronger autonomous creativity through new breakthroughs in the future. Creative ability is no longer the mysterious exclusive property of human beings, but a combinatorial generalization that can be generated by binary signals and game formulas. The "emergence" that arises from the interaction of a large number of components is no longer a mental activity unique to humans, but is becoming a normal capability of AI systems.
Others say: AI music (still) cannot replace human aesthetic judgment. People with stronger aesthetic ability and keener market sense will instead become more needed roles in the AI era.
But what is aesthetics?
Human aesthetic experience is a neurophysiological mechanism derived from the adaptive advantages of biological evolution. Human nerves have extremely high plasticity, so under the joint influence of external factors such as social concepts, cultural capital and power symbols, individual aesthetic tastes will gradually change with growth experience. Media and algorithms have actually subtly guided and reshaped everyone's aesthetic preferences, and various real interactive behaviors of people on the electronic network in turn provide AI with comprehensive and real-time aesthetic marking data.
In recent years, machine learning analysis of large-scale art data has shown that the aesthetic features perceived by humans can indeed be modeled and predicted through quantitative methods.
If you want to select the target content that best fits the current public aesthetics from massive information, or pick out the key features with the highest conversion rate in a specific vertical category, then with the support of big data, AI can easily complete the task by simply applying statistical methods. If you want to more timely capture the flowing trend of new outlets amid changes in fashion trends, select or even directly create works that are more likely to capture the future market, then the artificial intelligence brain with more sufficient samples, updated data and faster speed is fully capable of delivering a satisfactory answer.
Although AI may not be able to personally experience the complex emotions and aesthetic experiences of human beings, it does not need to directly feel your feelings — it only needs to be able to judge or predict your aesthetic choices now or in the future.
Some others say: AI (still) cannot replace human intuition.
But what is intuition?
When humans carry out a certain fast and automatic pattern recognition and prediction based on innate structure and acquired experience at the subconscious level, according to current situational clues, the so-called "intuition" is formed. If we regard intuition as an evaluation and decision-making model, then implicit learning of complex knowledge or rules in the environment, cognitive chunking that packages scattered information elements into a whole, and positive and negative feedback calibration after each decision-making action in the memory bank, all constitute important sources of personal intuition.
Deep neural networks of artificial intelligence are exactly designed to do these things. Generative AI for artistic creation is collecting and automatically refining countless experiences accumulated by humans for thousands of years that can be internalized and learned — especially the published-level finished products from top artists and gold producers — into a huge and sophisticated art intuition system, which can provide all-inclusive rapid association and construction capabilities at any time.
AI is instead the intelligent agent that maximizes the "human intuition", this probability tool.
There are still people who say: AI (still) cannot replace human common sense, which is even more untenable.
Although the currently repeatedly trained AI may still have a certain "hallucination" rate, we who have lived for so many years and experienced so many things are not necessarily accurate and self-evident in the consensus and judgment of all fields. Moreover, with technological progress and the change of social norms, the old "common sense base" will continue to change. The collective cognition of people at this moment is only a fleeting state in the vast historical process, not to mention that art has always encouraged innovation instead of sticking to conventions.
In the final analysis, the life machine called "human" is also just the result of iterative evolution of a series of material hardware and biochemical algorithms. There is no essential difference between us created by the universe and the AI created by us. No one can bypass the underlying physical laws, and cannot violate causal correlation, conditional connection and probabilistic linkage. Theoretically, all things that carbon-based human brains can do, silicon-based computers should also be able to do approximately.
In fact, our perception of the world around us is a process in which the brain uses a generative model to make internal simulations of external things, and then matches them with the sensory evidence we receive. Whether it is human imagination or the "direct" perception of the real world, it is a set of simulated reality called up by the neocortex of the brain using generative patterns.
The appearance of the world we see, hear and touch is not equivalent to the real and objective physical reality, but only the three-dimensional world inferred, predicted and reconstructed by the activated neural network in the human brain. This also explains various human brain functional phenomena that were once mysticized and theologized, including dreams and hallucinations.
The research results of brain science have successfully inspired one of the basic methods that have achieved amazing results in artificial intelligence today: AI learns to identify various real things including images, sounds and languages very effectively by repeatedly comparing the data it generates with the actual data.
A tiny cue can lead to a big picture. The clearer we understand the operating mechanism of intelligent beings and the laws of the world, the more we can replicate the known methods and mechanisms on computers; for the parts that we have not yet decrypted and are still exploring, there is no way to reasonably prove why future artificial intelligence cannot achieve them.
Therefore, in this contest between human brain and computer, we can say that there is no room to fully prove our own superiority, not to mention that AI obviously has absolute advantages that the current human body cannot reach, such as ultra-high mathematical accuracy, super strong learning speed, super large performance durability, perfect concentration, perfect memory, and perfect replication ability.
Science fiction has turned into reality so rapidly and turbulently that we have not had time to fully prepare. When AI music can already bring amazing and extraordinary creative standards to the whole world, and is unprecedentedly changing people's future expectations for the value of original art.
At this moment, for those of us who are already in or expect to join the music industry, where should we go tomorrow?
II
AI music has entered a blowout era and brought the inclusive benefit of "creation equality", which frankly speaking, has no major disadvantages for consumers (that is, people who listen to music, buy music products and watch music performances).
The music available to us is already large enough in quantity and wide enough in style. From music genres of all regions to musical instrument skills of playing various instruments, from emotional expression of joy, anger, sorrow and fear to auditory shaping of joys and sorrows, from life situations of ups and downs to philosophical thinking of love and hatred, almost all the expansion directions in musicality and literariness that people can think of have been competed and practiced by thousands of music creators.
For a single listener, there is already endless music to enjoy.
In this ocean of notes and words, AI pouring down from the clouds does not seem to bring qualitative changes to the content options of music listeners. However, as AI music squeezes the share of hand-made music, in this music industry that is already a red sea with strong head effect, the worries and concerns of creators (and practitioners in all links of the industrial chain) are increasingly emerging:
Since AI can already create works that sound as good as or even better than hand-made works, is it necessary to continue to create manually?
Will the audience's acceptance of AI music become higher or lower in the future?
What kind of future impact will AI music have on the income opportunities and ideal realization paths of creators/producers?
How useful will the skills of musical instruments, music theory, composition and arrangement accumulated through years of learning be in the future?
The key to answering these questions lies in that we need to sort out clearly: why people want to listen to music and why they have preference in music selection.
First of all, music that matches one's own auditory aesthetics will make people fond of it. Relaxation, happiness, surprise... Obtaining these positive state incentives seems to be the main motivation for us to listen to music.
If we dig deeper, human's music listening experience is closely related to the brain's prediction and reward mechanism. When organized sound information unfolds continuously over time, our brain will unconsciously predict the next content based on the sound rules we have learned before. If the prediction is correctly realized, or the prediction produces an error but is still within the understandable higher-level music framework, the brain will release different degrees of dopamine and bring pleasure.
This is why stable rhythm, periodic harmony and structured melody usually cater to people's general music preferences, while exquisite variations, unexpected modulations and subtle pitch differences can bring moderate intellectual excitement to people who listen carefully.
If we dig deeper, our bodies retain the oldest nature engraved in our DNA.
On the one hand, the ability of chordate animals to receive external frequency vibrations initially evolved to identify different sound sources to help judge the living environment. When sudden, loud, harsh and dissonant sounds appear, animals will immediately become alert and identify the spatial orientation of the sound source and the event. In music, the sense of release brought by the progression of chords from tension to resolution and the development of rhythm from change to unity is exactly the safety framework of the alert system at work. The ability of "automatically separating different sound sources", which originally belonged to survival skills, has also become part of the brain exercise play in our music appreciation process. On the other hand, the timbre of human voice is also the "auditory fingerprint" of the speaker's physical quality and reproductive potential. Sounds with specific characteristics emitted by people who conform to the listener's gender orientation will convey a certain sexual attraction.
It is worth noting that due to different listening experiences and music training backgrounds, everyone's music decoding ability and aesthetic expectations will also vary. This further leads to the interweaving of voices, color changes and structural designs in music, which are kind and comfortable for some brains; old and lack of novelty for some brains; and strange and incomprehensible for others.
Only when the material application and grammatical construction of various elements in music are exactly within the aesthetic ability range of the listener at that time, the feeling of "pleasant to hear" will arise spontaneously.
However, for modern people who are in complex social activities at all times and have advanced cognitive systems, things are not so simple. In addition to the pure "pleasant to hear" — that is, the direct reflection of auditory aesthetics, people's choice of music actually integrates many other factors. I classify them into three (overlapping) categories: functionality, cultural property and emotional property.
In different life scenarios, we prefer music that can bring corresponding positive benefits.
For example, we want to listen to bright and cheerful music on the way to school or work, listen to relaxing and decompressing music after school or work, listen to dynamic and passionate music when exercising, listen to tender and romantic music during intimate dates, listening to sad songs when sad makes it easier to quickly gain resonance, and then gradually transition to listening to happy songs which will really make people happy — all these benefit from the function of music to regulate human emotions.
Listening to soothing music can reduce cortisol levels and relieve stress, listening to slow-paced music can guide heart rate to drop and improve anxiety and sleep, listening to pleasant music can release endorphins to become a natural happy water and analgesic, and various positive effects on nerve function, immune function, endocrine function, and even gene expression are all clinically proven functional effects of music.
For children, music training and