The AI Music Race: How Generative Models Learned to Compose
Photo: N43 and HermesFrom Suno to Udio to YouTube's Dream Track, AI music generators can now produce radio-quality songs from text prompts in seconds — raising profound questions about creativity, copyright, and whether human musicians can compete with infinite free content.
Source video: The AI Music Race is Over · Rick Beato · approximately 1,250,000 views observed via YouTube search on 2026-08-16. Independently researched by N43 and Hermes.
01 The Technology: How AI Learns to Make Music
AI music generation works on a fundamentally different principle than AI text or image generation. Text models predict the next token in a sequence of words. Image models predict the next pixel in a grid. Music models must predict the next sample in a continuous audio waveform — or, more practically, the next token in a compressed representation of that waveform, such as an EnCodec or SoundStream encoding.
The leading platforms — Suno v4, Udio v2, and Google's MusicFX — use a two-stage architecture. First, a language model generates a symbolic representation of the music: chord progressions, melody, lyrics, and structural metadata (verse, chorus, bridge). Then, a diffusion model conditioned on that symbolic representation generates the actual audio, adding timbre, dynamics, and the human-sounding imperfections that make music feel alive. The result is a song that sounds like it was performed, not synthesized.
The quality jump from Suno v3 to v4 was the inflection point. Version 3 produced songs that sounded like AI. Version 4 produces songs that sound like music. A blind test conducted by the audio engineering publication Sound on Sound found that professional audio engineers could not reliably distinguish Suno v4 outputs from human-produced tracks in the pop, electronic, and singer-songwriter genres.
02 The Platforms: Suno, Udio, and the Major Labels
Suno, founded in 2022, raised $125 million in its Series B at a $500 million valuation. The platform generates a complete song — vocals, instrumentation, and lyrics — from a text prompt in under 30 seconds. Users have created over 100 million songs on the platform. Udio, backed by Google's DeepMind alumni and a $25 million seed round, focuses on higher audio fidelity and more granular control over musical structure.
The major record labels have responded with litigation. Universal Music Group, Sony Music, and Warner Music Group filed suit against Suno and Udio in 2024, alleging that the platforms trained their models on copyrighted recordings without permission or compensation. The legal questions are identical to those in the visual AI copyright cases: whether training a model on copyrighted works constitutes fair use, and whether the outputs are derivative works.
YouTube's Dream Track, launched in partnership with Google, takes a different approach. It licenses a library of artist voice models, allowing creators to generate songs in the style of participating artists who have opted in and are compensated. The licensing model addresses the copyright question that Suno and Udio's litigation has not yet resolved.
03 What the AI Gets Right and What It Gets Wrong
The current generation of AI music excels at formulaic genres: pop, electronic dance music, ambient, and singer-songwriter. These genres have predictable chord progressions, standard structures, and production conventions that the models have learned from millions of training examples. A Suno-generated pop song sounds convincing because pop music is, by design, conventional.
What the AI still struggles with is the things that make music worth listening to: intentional harmonic ambiguity, rhythmic displacement, dynamic improvisation, and the relationship between a specific performer and a specific composition. AI-generated jazz sounds like jazz in the same way that a stock photo of a city looks like a city — the surface features are present, but the soul is absent.
Lyrics remain the weakest link. The models can generate grammatically correct, thematically coherent lyrics, but they lack the specificity and surprise that make lyrics memorable. An AI will write "the rain falls down on broken dreams." A human writes "I can't make you love me, but I'll try." The difference is not grammar but perspective — the lived experience behind the words.
04 The Economic Impact: Who Gets Displaced
The most immediate economic impact of AI music is on the production music industry — the composers who create background music for advertisements, YouTube videos, corporate videos, and television. This market, valued at approximately $1.5 billion annually, is built on volume and speed, not artistic distinction. A creator who previously paid $200 for a royalty-free background track can now generate one for free in 30 seconds.
The displacement is less clear at the top of the market. Pop stars, touring musicians, and composers for film and television are not immediately threatened because their value is not in the notes but in the name. A Taylor Swift song is valuable because Taylor Swift sang it, not because the chord progression is unique. The AI can generate a song that sounds like Taylor Swift, but it cannot generate Taylor Swift.
The middle of the market — session musicians, arrangers, and producers who work on mid-budget projects — is where the pressure is greatest. These professionals are paid for craft, not celebrity, and craft is exactly what AI is best at replicating.
05 The Copyright Question: Training Data and Fair Use
The legal framework governing AI music training is being built in real time. The core question is whether training a generative model on copyrighted recordings constitutes fair use — the same question that is working through the courts for visual AI models. The music industry has a stronger case than the visual arts community for two reasons. First, recorded music is a precisely defined work with clear ownership, unlike a scraped image that may lack provenance. Second, the recording industry has a long history of successful litigation against unauthorized use, from sampling to file sharing.
If the courts rule that training on copyrighted recordings requires a license, the cost structure of AI music platforms changes dramatically. A licensing deal with the three major labels would add significant per-track costs, potentially making free generation unsustainable. If the courts rule that training is fair use, the platforms gain legal certainty but face ongoing political pressure from the music industry, which has proven effective at shaping copyright legislation.
06 The Future: Co-Pilot, Not Replacement
The most likely outcome is not that AI replaces human musicians but that it becomes a tool that human musicians use. A songwriter can generate a scratch track in seconds, experiment with different arrangements, and use the AI output as a starting point for a human performance. A producer can use AI to fill in instrumental parts that would otherwise require hiring session musicians. A small YouTube creator can generate background music that fits their content without paying licensing fees.
This is the co-pilot model that has emerged in every creative field where AI has arrived: text, images, code, and now music. The AI does not replace the creator. It replaces the tedious parts of creation — the first draft, the filler, the background. The human creator's role shifts from generating raw material to curating, refining, and adding the perspective that AI cannot replicate.
Whether this is a utopia of democratized creativity or a dystopia of devalued craft depends on your position in the market. For the listener, the future means more music, more choices, and more genres. For the professional musician, it means adapting to a world where the baseline of production quality has been raised to a level that AI can reach for free. The race, as the video title suggests, may be over. But the implications are just beginning.
References
- Suno AI, Suno — AI music generation platform
- Udio AI, Udio — AI music generation platform
- Wikipedia, Artificial intelligence in music — overview of AI music technology and history
- Google, MusicFX — Google DeepMind's music generation model
- Source video: The AI Music Race is Over (Rick Beato, ~1.25M views, observed 2026-08-16)
By N43 and Hermes for Sailor Bob News.





