

Here’s an approach that works remarkably well:
(Cf. Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis)
<text>It was the best of times, it was the worst of times</text> <audio>🜃🝃🜶🜻🜼🝊🝋…</audio>
Prompt:
<text>Here's a brand new sentence</text> <audio>
Continuation:
🝤🜱🝣🜲🝊🝢…</audio>
Prompt:
<text>Get to the chopper!</text> <audio>🝊🝢🜼🝣🝌🜠…</audio> <text>I'll be back</text> <audio>
Continuation:
🜘🝊🝣🜱🝋…</audio>
<low-quality>🜻🜶🜃🝊🜉🝋</low-quality> <high-quality>🜼🜠🜲🜶🝊🝮</high-quality>
Audio enhancement
<audio>🜻🜉🜃🜏🝤🝣</audio> <instruction>Replace hello with hi</instruction> <audio>🜻🜉🝃🝤🝣</audio>
Speech editing
<speech>🝣🝋🝢🜼🝌🝮</speech> <rap>🜲🜉🜼🜃🝌🜘</rap>
Speech-to-rap

Learn denoising step
→
→ 
"A dog"
Chain denoising steps to go from pure noise to novel output
→
→
→
→ 
"A cat"
"Hello world"
"Goodbye world"
Questions?
Edit slides with `npm run dev:talk generative-audio`