All Tech Talks

Tech Talk / 46 min

Audio Generation Models – From WaveNet to Real-Time AI Music

Cloud Ambassadors Team30 Apr 2026
Generative AIAudio GenerationText-to-SpeechWaveNetMusicGenAudioLDMLyriaDeep LearningMachine LearningAI Audio

About this session

Discover the fascinating journey of AI audio generation, from early rule-based systems to today's advanced models capable of producing realistic speech, music, and vocals in real time. This session explores the milestones that shaped the field, including breakthroughs like WaveNet, Jukebox, and modern token-based audio generation techniques. Learn how AI transforms complex audio into tokens, the difference between auto-regressive and latent diffusion models, and the training process that enables models to generate high-quality audio while maintaining safety and reliability. The session also covers data preparation, supervised fine-tuning, human feedback, and safeguards designed to reduce harmful or copyrighted outputs. Finally, explore the ongoing debates surrounding AI-generated music, including copyright ownership, creative authenticity, consistency in long-form audio, and the growing challenge of distinguishing AI-generated content from human-created work. Whether you're interested in generative AI, music technology, or machine learning, this session offers a comprehensive introduction to the future of AI-powered audio generation.

Browse more sessions