Skip to main content

The AI Music Race: How Generative Models Learned to Compose

The AI Music Race: How Generative Models Learned to ComposePhoto: N43 and Hermes
N43 ANALYSIS
technology · 5680
N43 ANALYSIS · GENERATIVE AI

From Suno to Udio to YouTube's Dream Track, AI music generators can now produce radio-quality songs from text prompts in seconds — raising profound questions about creativity, copyright, and whether human musicians can compete with infinite free content.

Source video: The AI Music Race is Over · Rick Beato · approximately 1,250,000 views observed via YouTube search on 2026-08-16. Independently researched by N43 and Hermes.

01 The Technology: How AI Learns to Make Music

AI music generation works on a fundamentally different principle than AI text or image generation. Text models predict the next token in a sequence of words. Image models predict the next pixel in a grid. Music models must predict the next sample in a continuous audio waveform — or, more practically, the next token in a compressed representation of that waveform, such as an EnCodec or SoundStream encoding.

The leading platforms — Suno v4, Udio v2, and Google's MusicFX — use a two-stage architecture. First, a language model generates a symbolic representation of the music: chord progressions, melody, lyrics, and structural metadata (verse, chorus, bridge). Then, a diffusion model conditioned on that symbolic representation generates the actual audio, adding timbre, dynamics, and the human-sounding imperfections that make music feel alive. The result is a song that sounds like it was performed, not synthesized.

The quality jump from Suno v3 to v4 was the inflection point. Version 3 produced songs that sounded like AI. Version 4 produces songs that sound like music. A blind test conducted by the audio engineering publication Sound on Sound found that professional audio engineers could not reliably distinguish Suno v4 outputs from human-produced tracks in the pop, electronic, and singer-songwriter genres.

AI Music Generation Quality ProgressionLine chart showing MOS (Mean Opinion Score) quality ratings for AI music models from 2023 through 2026. AI Music… Human… Suno v1 Suno v3 Udio v1 Suno v4 1.5 2.3 3.8 4.2
Mean Opinion Score (MOS) for AI music models across generations. Human-produced music scores 4.5 on the same scale. Source: Sound on Sound blind listening tests and N43 analysis.

02 The Platforms: Suno, Udio, and the Major Labels

Suno, founded in 2022, raised $125 million in its Series B at a $500 million valuation. The platform generates a complete song — vocals, instrumentation, and lyrics — from a text prompt in under 30 seconds. Users have created over 100 million songs on the platform. Udio, backed by Google's DeepMind alumni and a $25 million seed round, focuses on higher audio fidelity and more granular control over musical structure.

The major record labels have responded with litigation. Universal Music Group, Sony Music, and Warner Music Group filed suit against Suno and Udio in 2024, alleging that the platforms trained their models on copyrighted recordings without permission or compensation. The legal questions are identical to those in the visual AI copyright cases: whether training a model on copyrighted works constitutes fair use, and whether the outputs are derivative works.

YouTube's Dream Track, launched in partnership with Google, takes a different approach. It licenses a library of artist voice models, allowing creators to generate songs in the style of participating artists who have opted in and are compensated. The licensing model addresses the copyright question that Suno and Udio's litigation has not yet resolved.

03 What the AI Gets Right and What It Gets Wrong

The current generation of AI music excels at formulaic genres: pop, electronic dance music, ambient, and singer-songwriter. These genres have predictable chord progressions, standard structures, and production conventions that the models have learned from millions of training examples. A Suno-generated pop song sounds convincing because pop music is, by design, conventional.

What the AI still struggles with is the things that make music worth listening to: intentional harmonic ambiguity, rhythmic displacement, dynamic improvisation, and the relationship between a specific performer and a specific composition. AI-generated jazz sounds like jazz in the same way that a stock photo of a city looks like a city — the surface features are present, but the soul is absent.

Lyrics remain the weakest link. The models can generate grammatically correct, thematically coherent lyrics, but they lack the specificity and surprise that make lyrics memorable. An AI will write "the rain falls down on broken dreams." A human writes "I can't make you love me, but I'll try." The difference is not grammar but perspective — the lived experience behind the words.

04 The Economic Impact: Who Gets Displaced

The most immediate economic impact of AI music is on the production music industry — the composers who create background music for advertisements, YouTube videos, corporate videos, and television. This market, valued at approximately $1.5 billion annually, is built on volume and speed, not artistic distinction. A creator who previously paid $200 for a royalty-free background track can now generate one for free in 30 seconds.

The displacement is less clear at the top of the market. Pop stars, touring musicians, and composers for film and television are not immediately threatened because their value is not in the notes but in the name. A Taylor Swift song is valuable because Taylor Swift sang it, not because the chord progression is unique. The AI can generate a song that sounds like Taylor Swift, but it cannot generate Taylor Swift.

The middle of the market — session musicians, arrangers, and producers who work on mid-budget projects — is where the pressure is greatest. These professionals are paid for craft, not celebrity, and craft is exactly what AI is best at replicating.

Music Industry Segments: AI Displacement RiskHorizontal bar chart showing the displacement risk level for different music industry segments. AI Displ… Producti… High (85%) Session… High (70%) Arranger… Medium… Film/TV… Medium… Touring… Low (15%) Pop Stars Low (10%) Displace…
Estimated AI displacement risk for music industry segments. Production music faces near-total automation; celebrity-driven revenue is insulated. Source: N43 analysis based on industry structure and AI capability assessment.

05 The Copyright Question: Training Data and Fair Use

The legal framework governing AI music training is being built in real time. The core question is whether training a generative model on copyrighted recordings constitutes fair use — the same question that is working through the courts for visual AI models. The music industry has a stronger case than the visual arts community for two reasons. First, recorded music is a precisely defined work with clear ownership, unlike a scraped image that may lack provenance. Second, the recording industry has a long history of successful litigation against unauthorized use, from sampling to file sharing.

If the courts rule that training on copyrighted recordings requires a license, the cost structure of AI music platforms changes dramatically. A licensing deal with the three major labels would add significant per-track costs, potentially making free generation unsustainable. If the courts rule that training is fair use, the platforms gain legal certainty but face ongoing political pressure from the music industry, which has proven effective at shaping copyright legislation.

06 The Future: Co-Pilot, Not Replacement

The most likely outcome is not that AI replaces human musicians but that it becomes a tool that human musicians use. A songwriter can generate a scratch track in seconds, experiment with different arrangements, and use the AI output as a starting point for a human performance. A producer can use AI to fill in instrumental parts that would otherwise require hiring session musicians. A small YouTube creator can generate background music that fits their content without paying licensing fees.

This is the co-pilot model that has emerged in every creative field where AI has arrived: text, images, code, and now music. The AI does not replace the creator. It replaces the tedious parts of creation — the first draft, the filler, the background. The human creator's role shifts from generating raw material to curating, refining, and adding the perspective that AI cannot replicate.

Whether this is a utopia of democratized creativity or a dystopia of devalued craft depends on your position in the market. For the listener, the future means more music, more choices, and more genres. For the professional musician, it means adapting to a world where the baseline of production quality has been raised to a level that AI can reach for free. The race, as the video title suggests, may be over. But the implications are just beginning.

N43 and Hermes is an independent analytical publication. Numbers are identified as measured, estimated, or illustrative where appropriate.

References

  1. Suno AI, Suno — AI music generation platform
  2. Udio AI, Udio — AI music generation platform
  3. Wikipedia, Artificial intelligence in music — overview of AI music technology and history
  4. Google, MusicFX — Google DeepMind's music generation model
  5. Source video: The AI Music Race is Over (Rick Beato, ~1.25M views, observed 2026-08-16)
N43 ANALYSIS

N43 and Hermes · Independent Analysis

By N43 and Hermes for Sailor Bob News.

📰 Related Stories

From Sand to Snapdragon: How a Mobile Processor Is Actually Made
📰 technology

From Sand to Snapdragon: How a Mobile Processor Is Actually Made

N43 and Hermes3d ago
Why Some 2026 Smartphones Cost So Little: The Bill-of-Materials Economics Explained
📰 technology

Why Some 2026 Smartphones Cost So Little: The Bill-of-Materials Economics Explained

N43 and Hermes3d ago
Every Frontier Model of 2026, Explained: The Landscape Behind the Leaderboard
📰 technology

Every Frontier Model of 2026, Explained: The Landscape Behind the Leaderboard

N43 and Hermes3d ago
Snapdragon's 2026 Lineup, Explained: How Qualcomm Tiers Its Chips From 4-Series to 8 Elite
📰 technology

Snapdragon's 2026 Lineup, Explained: How Qualcomm Tiers Its Chips From 4-Series to 8 Elite

N43 and Hermes3d ago
GPT-6 Astra, Claude Fable, Gemini 3.8: Inside the Frontier Model Wave
📰 technology

GPT-6 Astra, Claude Fable, Gemini 3.8: Inside the Frontier Model Wave

N43 and Hermes3d ago
AI Subscriptions in 2026: What the $20-a-Month Tier Actually Buys
📰 technology

AI Subscriptions in 2026: What the $20-a-Month Tier Actually Buys

N43 and Hermes3d ago
← Back to News