Models that talk, sing and make noise:
Generative AI for audio

Alex Peattie (alexpeattie.com / @alexpeattie)
Sr. Director of Engineering, Front


DSSAI 2026

Slides online at alexpeattie.com/talks/generative-audio

Agenda

  • Why is audio generation interesting?
  • What makes audio generation hard
  • And what actually makes it easy

Why is audio generation interesting?

Demo: text to speech

Did you hear about the Data Science Summit in Warsaw? Apparently it’s absolutely crazy.

Demo: music generation

A song about how audio generation is easy, in the style of 2000s Hip Hop

Demo: music style control

center

center

“If AI research is Star Wars and OpenAI is the death star, then without a doubt the rebels are building audio models. The best models for voice – TTS, STS, STT, and the like – are not coming from the big labs. Instead, they’re built by their underfunded, understaffed, and underhyped siblings,”

ElevenLabs

$11Bn company

$?? millions in training cost

OmniVoice

Small non-profit

~250 H800 hours
~$500 training cost

Goal

  • Point future audio rebels who want to experiment with their own models in the right direction.
  • Give an understanding of the key concepts used by the SoTA audio generation models in 2026.

What makes audio generation hard

(on paper)

  1. Audio is high-dimensional
  2. Audio generation requires linguistic and cultural context
  3. Our ears are unforgiving

Building blocks

"Hello DSSAI"
Tokens
"Hello DSSAI"
Pixels (naively)
Audio samples (naively)

center

16,000 - 48,000 samples per second

"Hello DSSAI"
3 tokens
20-50k samples
10,000x difference

Audio generation requires linguistic and cultural context

abc → 🔊

The LED dog lead was made of lead.

“A boppy, uptempo pop tune in the style of Sabrina Carpenter”

A subtler example

Little Lucy, seeing the dog, said “What a cute puppy!”

Human ears are unforgiving

Incredible range

Falling leaf 🍂 Jumbo jet takeoff 🛫
  • 1 trillion times difference between quietest and loudest sounds
  • For the quietest sounds, our eardrum moves one picometer (100x smaller than a hydrogen atom’s diameter)

We’re sensitive to…

  • Noise
  • Tinny or metallic speech
  • Crackle or hiss
  • Glitches

Audio is linear

  1. Audio is high-dimensional
  2. Audio generation requires linguistic and cultural context
  3. Our ears are unforgiving

What makes audio generation easy

(how we solve these problems in practice)

  1. Audio is high-dimensional → Neural audio codecs
  2. Audio generation requires linguistic and cultural context → Use pre-trained LLMs
  3. Our ears are unforgiving → Leverage diffusion techniques

Sampled audio

Baseline

Spectrogram

20x compression

?

?

??? compression

Let’s play a game…

🌴🏐 →
Castaway
💍💍💍💍⚰️ →
Four Weddings and a Funeral
🔇🐑🐑🐑 →
The Silence of the Lambs
👦🏻👦🏼💍🌋⏎👑 →
Lord of the Rings: The Return of the King

Observations

  • This is effectively a compression task
  • Consider 💍 - it’s learned, and contextual
  • Compressing efficiently required a powerful, expensive model (our brains 🧠!)
🜱🜲🜶🜻🝡🝊🝋🝃 🜏🜠🝤🜼🝣🝢🝌🝮 🜘🜉🜃🝊🜶🝋🜲🜻 🝣🜱🝤🜼🝢🜠🝃🝮
Codebook
200kb audio
compressed to
0.8kb tokens
via
3GB model
"Hello DSSAI"
3 tokens
20-50k samples
🜱🜲🜶🜻
100 tokens
500x-1000x effective compression
EnCodec DAC TiCodec SNAC WavTokenizer Stable-Codec SpeechTokenizer X-Codec
Mimi LinaCodec SemantiCodec FACodec LSCodec WavLM MOSS-Audio-Tokenizer ContentVec
UniCodec FuseCodec MagiCodec LongCat-Audio-Codec LayaCodec Kanade H-Codec-2.0 Fish Audio Codec

Using pre-trained LLMs

Here’s an approach that works remarkably well:

  • Take a pretrained, open-weights LLM (e.g. Llama 1B)
  • Take audio files with transcriptions (human or AI generated)
  • Encode the audio as neural audio codec tokens
  • Fine-tune the LLM!

(Cf. Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis)

Training

<text>It was the best of times, it was the worst of times</text>
<audio>🜃🝃🜶🜻🜼🝊🝋…</audio>

Inference

Prompt:

<text>Here's a brand new sentence</text>
<audio>

Continuation:

🝤🜱🝣🜲🝊🝢…</audio>

Voice cloning (via one-shot prompting)

Prompt:

<text>Get to the chopper!</text>
<audio>🝊🝢🜼🝣🝌🜠…</audio>

<text>I'll be back</text>
<audio>

Continuation:

🜘🝊🝣🜱🝋…</audio>
<low-quality>🜻🜶🜃🝊🜉🝋</low-quality>
<high-quality>🜼🜠🜲🜶🝊🝮</high-quality>

Audio enhancement

<audio>🜻🜉🜃🜏🝤🝣</audio>
<instruction>Replace hello with hi</instruction>
<audio>🜻🜉🝃🝤🝣</audio>

Speech editing

<speech>🝣🝋🝢🜼🝌🝮</speech>
<rap>🜲🜉🜼🜃🝌🜘</rap>

Speech-to-rap

The recipe

  • Assemble your dataset
  • Choose your LLM
    • LLASA = Llama 1B/8B, VibeVoice = Qwen 2.5 1B/7B, OmniVoice = Qwen 3 0.6B
  • Choose your neural audio codec
  • Choose your training method (full fine tune, LoRA, DoRA, adapter, …)
  • Order-of-magnitude training costs $100s - $1000s
    • Less if you fine-tune from a foundation audio model for your task

Leverage diffusion techniques

(or flow matching)

Denoising diffusion: the intuition

add noise
×

During training

Learn denoising step

"A dog"

During inference

Chain denoising steps to go from pure noise to novel output

"A cat"

Advantage: naturally tends to generate clean outputs

  • Diffusion models generate by denoising
  • So they tend to be good add producing clean, sharp outputs
  • Good for avoiding audio artifacts

Advantage: CFG for control

Low CFG
High CFG
Better quality
Better prompt
adherence

Diffusion for audio

  • Modify or mask (delete) audio tokens
  • Or, introduce noise elsewhere (into the waveform or spectrogram)
  • Then reconstruct clean audio

Example diffusion approach: masked diffusion

During training:
🜱🜲🜶🜻🝡
🜱🜲🜶🜻🝡
🜱🜲🜶🜻🝡

"Hello world"

During training:
🜏🜠🜱🜲🝡🜶
🜏🜠🜱🜲🝡🜶
🜏🜠🜱🜲🝡🜶
🜏🜠🜱🜲🝡🜶

"Goodbye world"

LLM vs. Diffusion

  • We can use pure diffusion
  • Or we can use pure LLM-backed generation
  • Most frontier models use a combination

The frontier recipe

  • Assemble your dataset
  • Choose your LLM backbone
  • Choose your neural audio codec
  • Choose your training method
  • Choose your diffusion approach

Recap

  • Audio generation has never been more exciting
  • And it’s still accessible to small labs and hobbyists
  1. Audio is high-dimensional → Neural audio codecs
  2. Audio generation requires linguistic and cultural context → Use pre-trained LLMs
  3. Our ears are unforgiving → Leverage diffusion techniques

Thank you!

alexpeattie.com/talks/generative-audio

Questions?

Edit slides with `npm run dev:talk generative-audio`