Tech Talk / 46 min
Audio Generation Models – From WaveNet to Real-Time AI Music
About this session
Discover the fascinating journey of AI audio generation, from early rule-based systems to today's advanced models capable of producing realistic speech, music, and vocals in real time. This session explores the milestones that shaped the field, including breakthroughs like WaveNet, Jukebox, and modern token-based audio generation techniques. Learn how AI transforms complex audio into tokens, the difference between auto-regressive and latent diffusion models, and the training process that enables models to generate high-quality audio while maintaining safety and reliability. The session also covers data preparation, supervised fine-tuning, human feedback, and safeguards designed to reduce harmful or copyrighted outputs. Finally, explore the ongoing debates surrounding AI-generated music, including copyright ownership, creative authenticity, consistency in long-form audio, and the growing challenge of distinguishing AI-generated content from human-created work. Whether you're interested in generative AI, music technology, or machine learning, this session offers a comprehensive introduction to the future of AI-powered audio generation.