Tutorials
LLMs x Music: Scalable Language Modeling Unlocks MIR Applications
Presenters: Zachary Novack, Seungheon Doh, Huan Zhang, Yewon Kim, Keunwoo Choi
Abstract:
Large language models (LLMs) are rapidly becoming a foundation for user-facing systems, with the potential to transform how people discover, understand, create, produce, perform, and learn music. Moving beyond a surface-level treatment of LLM usage, this tutorial offers an in-depth exploration of the capabilities that make these models particularly powerful for musical applications, showing how they can enable meaningful experiences for listeners, creators, performers, and learners.
We first introduce the foundations of modern LLMs, from large-scale pretraining to instruction tuning and preference alignment. We then show how their core capabilities support agentic pipelines that reason, call specialized tools, and interact with environments. Next, we explore how LLMs connect information across audio, symbolic notation, score images, and textual metadata for music understanding. Building on these foundations, the tutorial emphasizes applications designed around the needs of music users: conversational recommendation for listeners; collaborative composition for creators; and personalized assessment, coaching, and interactive education for performers and learners.
Participants will learn how the core capabilities of LLMs translate into practical applications for music users. The tutorial concludes by examining current limitations and promising directions for more capable, grounded, and musically meaningful systems.
Presenter Bios:
Zachary Novack is a Research Scientist at Spotify in the Artist First AI-Music Lab. Graduating in 2026 with a PhD from UC San Diego, Zachary’s research has focused both on efficient interactive music generation and large scale evaluation and benchmarking for LLMs for musical reasoning. He has received numerous awards across diverse venues, including oral presentations at ICML, ICASSP, and Interspeech, spotlight papers at ICLR and NeurIPS, and recently Best Paper at ISMIR 2025 for work on diagnosing Music-LLM evaluation. He co-presented the tutorial “Connecting Music Audio and Natural Language” in ISMIR 2024 and organized the AI Music workshop at NeurIPS 2025. Outside of academics, Zachary is passionate about music education and teaches competitive indoor drumline in the southern California area.
Seungheon Doh is a Research Scientist at Sony. His research focuses on conversational music understanding, recommendation, and creation. He has published research on large language models for music applications at venues including ISMIR, ICLR, IEEE ICASSP, and IEEE TASLP. He co-presented a tutorial on "Connecting Music Audio and Natural Language" at ISMIR 2024 and has organized several workshops (NLP4Musa, LLM4MA) exploring the intersection of music and language. Seungheon has interned at Sony AI, Adobe Research, Chartmetric, NAVER, and ByteDance.
Huan Zhang is a Research Scientist at Clefer Inc., exploring how LLMs can understand, generate, and evaluate performance expression in music.
She has developed multimodal models such as LLM coach that provides formative feedback on performances, and framework for text-controlled expressive performance generation. Huan has collaborated with industry leaders such as Sony CSL and Yamaha, and is deeply committed to applying her research to music education. Huan also serve as the MIREX Task Captain of RenCon (Expressive Performance Rendering Competition) and Music Performance Difficulty.
Yewon Kim is a Ph.D. student in the Computer Science Department at Carnegie Mellon University, where she is advised by Chris Donahue. Yewon's research explores how generative AI can enhance human creativity in music, developing interactive systems and methods for human-AI music co-creation.
She has published across HCI (ACM CHI, ACM DIS) and AI (NeurIPS) venues, and was recognized with a Best Paper Award at ACM CHI. She has also interned at Adobe Research.
Keunwoo Choi is an AI Research Director at Upstage, where he contributed to Solar Open 100B and is currently training omni-modal LLMs. He is also an adjunct professor at Culture Technology, KAIST, South Korea. Recently, he led an R\&D team on LLMs for drug discovery and biomedical applications at Genentech. Previously, he earned a PhD from Queen Mary University of London, UK, and has worked on music, audio, and AI research in the industry, including roles at Spotify, ByteDance, and Gaudio Lab Inc.
Bridging Music AI Research and Music Creation: How to Deploy AI Models to Digital Audio Workstations Using HARP
Presenters: Frank Cwitkowitz, Zhiyao Duan, Saumya Pailwan, Bryan Pardo, Huiran Yu
Abstract:
An overarching goal of MIR research is to develop tools for enhancing various kinds of musical experiences. However, there is still considerable technical friction between such technologies and the musicians and producers seeking to use them within their creative process. Open-source research deliverables (e.g., model weights and inference code) often require some degree of technical literacy to use, and this is typically beyond the preparation of most musicians and producers. At the same time, deploying such technologies into a standard music production workflow, which involves Digital Audio Workstations (DAWs) and plug-ins, typically requires software development expertise and effort which goes beyond the scope of most research projects.
In recent years, we have been developing HARP, an open-source framework and application to streamline the process of deploying audio research models into the music production workflow. HARP integrates seamlessly with most DAWs and allows end-users to discover, access, and leverage AI for audio and music. Our companion package PyHARP can be used to wrap and deploy arbitrary Python inference code to HARP and supports a variety of models with any combination of audio, MIDI, and/or text inputs and output. This makes it possible for audio researchers to produce project deliverables that integrate into the music production workflow in a few simple steps, without the hassle of creating an entire DAW plugin from scratch.
This tutorial aims to introduce HARP to the broader ISMIR community through live demonstrations and hands-on experience in using and deploying models to HARP. It is designed for MIR enthusiasts with diverse backgrounds, including researchers who may be interested in deploying their own models, musicians who wish to use cutting-edge AI tools as part of their creative process, and everyone in between.
Presenter Bios:
Frank Cwitkowitz received his B.S. and M.S. in Computer Engineering from the Rochester Institute of Technology in 2019. He is currently pursuing a Ph.D. degree in the Department of Electrical and Computer Engineering at the University of Rochester, under the supervision of Professor Zhiyao Duan. His research interests lie primarily at the intersection of music information retrieval and machine learning, with a particular emphasis on problems surrounding automatic music transcription and building tools for musicians. He has completed research internships at Chordify, Yousician, and Sony, and also regularly serves as a reviewer for conferences and journals such as ISMIR, ICASSP, TISMIR, and TASLP.
Zhiyao Duan is a professor in Electrical and Computer Engineering, Computer Science and Data Science at the University of Rochester. He is also a co-founder of Violy, a music tech company for instrument learning. He received his B.S. in Automation and M.S. in Control Science and Engineering from Tsinghua University, China, in 2004 and 2008, respectively, and his Ph.D. in Computer Science from Northwestern University, USA, in 2013. His research interest is in computer audition and its connections with computer vision, natural language processing, and computational neuroscience. He received a best paper award at SMC 2017, a best paper runner-up at ISMIR 2017 and ISMIR 2024, and an NSF CAREER award. His research has been funded by NSF, NIH, NIJ, New York State Center of Excellence in Data Science, Adobe, ByteDance, IngenID, Kwai, Meta, Microsoft, and University of Rochester internal awards on AR/VR, health analytics, and data science. He served as Scientific Program Co-Chair of ISMIR 2021, associate editor of IEEE Open Journal of Signal Processing, guest editor of Transactions of the International Society for Music Information Retrieval, and senior area editor of IEEE Signal Processing Letters. He is a member of the IEEE Signal Processing Society Audio and Acoustic Signal Processing Technical Committee. He was the President of the ISMIR society between 2024 and 2025. He co-presented tutorials at ISMIR in 2015, 2019 and 2023.
Saumya Pailwan received her M.S. in Computer Science from Northwestern University in December 2025 and her B.Tech in Computer Science from NMIMS University, Mumbai, India. She is currently a researcher in the Interactive Audio Lab under the supervision of Professor Bryan Pardo. Her research interests span machine learning and generative models for audio, with a focus on improving symbolic music generation. Her recent work focuses on integrating machine learning into music tools and enhancing their usability.
Bryan Pardo is a Northwestern Computer Science Professor, who studies fundamental problems in generative modeling of music, speech and sound effects, computer audition, content-based audio search and also develops inclusive interfaces blind and low-vision users of audio production tools. He is head of Northwestern University’s Interactive Audio Lab and co-director of the Northwestern University Center for HCI+Design. He received a M. Mus. in Jazz Studies in 2001 and a Ph.D. in Computer Science in 2005, both from the University of Michigan. He has authored over 150 peer-reviewed publications. He has developed speech analysis software for the Speech and Hearing department of the Ohio State University, statistical software for SPSS and worked as a machine learning researcher for General Dynamics. His patented technologies have been productized by companies including Bose, Adobe, Lexi, and Ear Machine.
Huiran Yu received her M.S in Computer Science from Carnegie Mellon University in December 2022 and her B.Eng in Computer Science and Technology from Tsinghua University in 2020. She is pursuing a Ph.D. degree in the Department of Electrical and Computer Engineering at the University of Rochester, under the supervision of Professor Zhiyao Duan. Her research interests lie in music information retrieval, specifically music transcription, and symbolic music analysis and generation. She also works on static trait disentanglement in streaming voice conversion. She has completed internships at Tencent, TikTok, and Bosch, and published papers in venues such as ISMIR and Interspeech.
Introduction to Music Information Retrieval: Three Perspectives
Presenters: Meinard Müller, Christof Weiß
Abstract:
Music Information Retrieval (MIR) is an interdisciplinary field that develops computational methods for analyzing, organizing, retrieving, and understanding music. This tutorial introduces MIR from three complementary perspectives: signal processing, computational musicology, and education. From the signal processing perspective, we present the fundamental representations, concepts, and algorithms of MIR, covering both traditional signal processing techniques and modern machine learning methods. Topics include music synchronization, source separation, audio decomposition, and content-based retrieval, while highlighting connections between classical approaches and recent developments such as differentiable formulations and learning objectives based on spectral and hierarchical loss functions. From the computational musicology perspective, we demonstrate how MIR techniques can be used to investigate harmony, tonality, musical structure, style, and performance differences through cross-version comparison, large-scale corpus analysis, and musicological applications. From the educational perspective, we present MIR as a rich domain for teaching signal processing, machine learning, and data science, introducing recent educational resources, interactive notebooks, datasets, open-source software, and current research directions. The tutorial is designed for a broad audience, ranging from newcomers seeking an accessible introduction to MIR to researchers interested in connecting classical methods with recent advances in the field.
Presenter Bios:
Meinard Müller received the Diploma degree (1997) in mathematics and the Ph.D. degree (2001) in computer science from the University of Bonn, Germany. Since 2012, he has been a professor for Semantic Audio Signal Processing at the International Audio Laboratories Erlangen, a joint institution of the Friedrich-Alexander-Universität Erlangen-Nürnberg (FAU) and the Fraunhofer Institute for Integrated Circuits IIS. His research focuses on music information retrieval, audio signal processing, and music processing. He has served the research community in various roles, including as a member of the IEEE Audio and Acoustic Signal Processing Technical Committee, a member of the Senior Editorial Board of the IEEE Signal Processing Magazine, and a member of the Board of Directors of the International Society for Music Information Retrieval (ISMIR, 2009-2021), serving as its president in 2020/2021. He was elevated to IEEE Fellow in 2020 for contributions to music signal processing and currently serves as Editor-in-Chief of TISMIR. He has extensive experience in teaching and tutorials, with numerous presentations at ICASSP and ISMIR. He is the author of the monograph "Information Retrieval for Music and Motion" (Springer 2007) and the "Fundamentals of Music Processing" (Sprinber 2015), which is widely used in MIR education.
Christof Weiß received the Ph.D. degree (2017) in media technology from Technische Universität Ilmenau, Germany. He also holds degrees in physics and music (composition) from the University of Würzburg and the Würzburg University of Music. Since 2022, he has been a professor for Computational Humanities at the University of Würzburg, where he leads a DFG-funded Emmy Noether research group on computational music analysis. Previously, he worked at the International Audio Laboratories Erlangen and the Fraunhofer Institute for Digital Media Technology (IDMT) and was a visiting researcher at Télécom Paris. His research lies at the intersection of music information retrieval, audio signal processing, and computational musicology, with a focus on harmonic and tonal analysis and corpus-based approaches. He has received several awards, including the Baldwin and Inge Knauf Award (2026), a Best Paper Award at CHR (2023), and the KlarText award for science communication (2018). He served as a member of the Board of Directors of the International Society for Music Information Retrieval (ISMIR, 2022-2023) and has been a member of the Editorial Board of the ACM Journal on Computing and Cultural Heritage (JOCCH) since 2022. He has extensive experience in interdisciplinary teaching and has given numerous invited talks and tutorials in MIR and computational musicology.
Evaluating Music Foundation Models: From MIREX To Multimodal, Preference-Aware Benchmarking
Presenters: Yinghao Ma, Junyan Jiang, Yizhi Li, Weixiong Chen
Abstract:
Recent progress in music foundation models, including music representation models, audio language models, multi- modal instruction-following systems, and long-form music generation models, has created an urgent need for new evaluation methodologies. Traditional MIR evaluation has largely focused on task-specific metrics and benchmark datasets, while modern music foundation models require broader evaluation across representation quality, instruction following, multimodal alignment, musicality, and human preference. This tutorial introduces a unified view of music evaluation in the foundation model era. We begin by revisit- ing the role of evaluation in MIR through the lens of MIREX and related benchmark traditions, and discuss how these ideas can be extended to support faster iteration of music foundation models. We then cover evaluation for represen- tation learning, using MARBLE and its later developments as examples of benchmark design for universal music audio representations. Next, we examine supervised fine-tuning and instruction-following evaluation for audio-LLMs, fo- cusing on the design principles behind CMI-Bench 2.0 and the limitations of current generation-based evaluation for music. We further introduce evaluation metrics and reward- model benchmarks, including CMI-RewardBench, and dis- cuss how such frameworks may inspire new MIREX-style tasks for music generation and multimodal systems. Finally, we outline future directions on how evaluation can support reinforcement-learning-based post-training, preference op- timization, and tokenizer design for next-generation music models. This tutorial is intended for MIR researchers, ma- chine learning practitioners, and newcomers interested in how evaluation can serve as a bridge between traditional mu- sic information retrieval and emerging foundation-model- based music systems. It aligns strongly with the ISMIR 2026 theme of “Crossroads” by connecting established MIR evaluation culture with new paradigms in generative and multimodal AI.
Presenter Bios:
Junyan Jiang is a Ph.D. candidate in New York University Shanghai and a visiting student at MBZUAI advised by Prof. Gus Xia. His research topic is music information retrieval and music generation, with a focus on self-supervised learning. Junyan has served as MIREX competition organizer.
Yinghao Ma is a last-year PhD candidate in the AI and Music program at the Centre for Digital Music, Queen Mary University of London, supervised by Dr Emmanouil Benetos and Dr. Chris Donahue. His research topic is on foundation models for music understanding. His work has been accepted to leading conferences in the field such as NeurIPS, ICLR, ISMIR and ICASSP. He is one of the co-founders of the open-source community Multimedia Art Projection (M-A-P). He is also a Googe PhD fellow (2025) on machine perception track on music fountation models.
Yizhi Li is a last-year PhD computer science student at the University of Manchester, focusing on LLM reasoning and multimodal foundation models, supervised by Professor Chenghua Lin.
Yizhi co-founded the MAP research community to promote open-source multimodal research, where the models have achieved over 50K monthly downloads, with over 70 collaborators from academia and industry.
He has published extensively at top conferences like NeurIPS, ICLR, and EMNLP on spanning topics, and received the Best Student Pitch award at the MultimodalAI'23 workshop.
Through his research, Yizhi has built a strong foundation in large-scale LM training, published 10+ first/leading author papers and achieved 2K+ citations.
Weixiong Chen is a first year PhD student at the Centre for Digital Music (C4DM), Queen Mary University of London. His research interests include multimodal music understanding, controllable music generation. His current work focuses on developing foundation models for music that integrate audio, symbolic, and textual representations to enable fine-grained music understanding and controllable generation.
Attribution in Music AI: Concepts, Systems, and Implications
Presenters: Jongpil Lee, Fabio Morreale, Wonil Kim, Joan Serra, Yuki Mitsufuji
Abstract:
Generative AI is transforming how music is created, remixed, and consumed. While recent advances have greatly improved music generation, fundamental questions surrounding attribution, provenance, authorship, licensing, and revenue distribution remain unresolved. Existing approaches primarily rely on post-hoc analysis, including fingerprinting, version identification, sample identification, watermarking, and training-data attribution. Although effective in specific scenarios, these methods often provide only indirect evidence and have limited ability to explain increasingly complex generation processes involving remixing, style transfer, iterative editing, and AI agents.
This tutorial introduces Attribution in Music AI as an emerging research area at the intersection of Music Information Retrieval (MIR), generative AI, copyright, and content provenance. It presents a unified framework spanning post-hoc attribution, attribution-by-design, component-level attribution, and copyright and interoperability frameworks. Participants will learn how attribution can become an integral part of modern music generation systems, supporting transparency, licensing, and interoperable rights management.
The tutorial combines conceptual foundations with practical systems and demonstrations. Topics include audio fingerprinting, musical version identification, sample identification, training-data attribution, component-level attribution, AI music agents, multimodal attribution, DDEX, and current industry developments. Designed for researchers and practitioners in MIR, machine learning, audio processing, and generative AI, the tutorial provides an accessible introduction to attribution in music AI while highlighting future research directions and practical implications.
Presenter Bios:
Jongpil Lee is the Co-founder and CEO of Neutune, where he is building MixAudio, a music AI agent system integrating generation and remixing with built-in component-level attribution. Neutune is a member of DDEX, C2PA, and WIPO initiatives related to AI and content provenance. He received the Ph.D. and M.S. degrees from the Graduate School of Culture Technology at the Korea Advanced Institute of Science and Technology (KAIST), and the B.S. degree in electronic engineering from Hanyang University.
Fabio Morreale is a Staff Research Scientist at Sony AI. He has a PhD in Computer Science (University of Trento, Italy) and a MSc and BSc in Computer Science (University of Verona, Italy). He also worked as a Postdoctoral Researcher at the Augmented Instruments Laboratory at Queen Mary University of London and at the Interaction Lab of the University of Trento. His research is centered on the design of new technologies for music performance and on the politics of music technology.
Wonil Kim is the Co-founder and Head of Music Research Team at Neutune and a Ph.D. student at the Graduate School of Culture Technology at KAIST. At Neutune, he works on MixAudio, a music AI agent system for generation, remixing, production support, and attribution. He received the M.S. degree from KAIST Graduate School of Culture Technology and the B.S. degree in Electronic Production and Sound Design from Berklee College of Music. Drawing on his professional background in music production, he studies practical AI systems that support creative workflows for music producers and creators.
Yuki Mitsufuji received the B.S. and M.S. degrees in information science from Keio University, Minato, Japan, and the Ph.D. degree from the University of Tokyo, Bunkyo, Japan. He is currently a Lead Research Scientist and VP of AI Research with Sony, leading two departments (Creative AI Lab, Music Foundation Model Team), and a Visiting Research Professor with New York University, New York, NY, USA. He is on the IEEE AASP Technical Committee 2023–2026. He chaired multiple workshops on generative models for audio at ICASSP, NeurIPS, and ECCV.
Joan Serra is a Staff Research Scientist and Team Lead at Sony AI. His research focuses on machine learning for audio and multimedia analysis, synthesis, and retrieval. He received his M.Sc. and Ph.D. degrees in machine learning for audio from the Music Technology Group at Universitat Pompeu Fabra. He previously held research positions at IIIA-CSIC, Telefónica R\&D, and Dolby Laboratories, working on artificial intelligence and machine learning. He has also been a visiting researcher at the Max Planck Institute for the Physics of Complex Systems and the Max Planck Institute for Computer Science. His work includes over 150 publications and more than 20 patents, and he regularly serves as a reviewer and area chair for major conferences.
Music for Learning: From Experimental Design to Multimodal Measurement
Presenters: Ying Que, Xiao Hu, Fanjie Li, Ruilun Liu
Abstract:
Music for Learning is an interdisciplinary research area at the intersection of Music Information Retrieval (MIR), education, psychology, and human-computer interaction. While MIR has made remarkable progress in music analysis, recommendation, and generation, less attention has been devoted to understanding how music influences learning processes and how MIR technologies can be designed and evaluated for educational applications. This tutorial presents a methodological framework for conducting human-centered research on music for learning, spanning controlled laboratory experiments, naturalistic field studies, multimodal data collection, measurements, and analysis. Participants will learn how MIR techniques can be used to characterize musical stimuli, and how these representations can be integrated with behavioral, subjective, and physiological measurements, especially eye tracking, electrodermal activity, heart-rate variability, and electroencephalography, to investigate learners’ cognitive processes and affective responses. This tutorial further introduces practical considerations for experimental design, multimodal data collection, measurement, and modelling, illustrated through laboratory experiments, naturalistic field studies, and live Jupyter Notebook demonstrations. Designed for researchers and practitioners from both MIR and adjacent disciplines, the tutorial provides an accessible introduction to music for learning research while informing design implications for developing evidence-based, user-centered MIR technologies for the learning contexts.
Presenter Bios:
Ying Que is an Assistant Professor in the School of Psychology at South China Normal University. She received her Ph.D. from the Faculty of Education at The University of Hong Kong. Her research lies at the intersection of music psychology and multimodal learning analytics, with a particular focus on understanding how background music influences learners' cognition, emotion, and behavior. She has conducted a series of laboratory and real-world studies integrating behavioral measures, eye tracking, electroencephalography, electrodermal activity, and heart-rate variability to investigate music-facilitated learning. Her work has been published in leading journals and conferences, including Reading Research Quarterly, International Journal of Human-Computer Interaction, Interactive Learning Environments, Scientific Reports, ISMIR and Learning Analytics & Knowledge (LAK) conference. She is actively involved in developing multimodal methodologies that bridge MIR technologies with learning analytics and educational research.
Xiao Hu is an Associate Professor in the College of Information Science at the University of Arizona. She obtained her Ph.D. in Information Science from University of Illinois. Her research aims to improve people’s learning and well-being through intelligent systems, with interests spanning Music Information Retrieval, Learning Analytics, Artificial Intelligence in Education, and Human-Computer Interaction. She has published extensively in leading venues across MIR, Technology-enhanced learning, and HCI, and has served on the program committees of ISMIR and related conferences. Her recent research focuses on user-centered music recommendation, multimodal analytics, and music-supported learning, integrating MIR techniques with behavioral and physiological sensing to better understand users’ experiences in authentic learning contexts. She has extensive experience leading and supervising interdisciplinary research and delivering tutorials and invited talks on MIR and learning technologies.
Fanjie Li is a Ph.D. candidate in the Department of Teaching and Learning at Vanderbilt University. Her research combines human-centered informatics, learning sciences, generative AI, and participatory design to develop intelligent technologies that support learning and well-being. Before joining Vanderbilt University, she conducted research at The University of Hong Kong on user-centered, context-aware music recommendation using physiological sensing, music processing, and user experience methodologies. Her work focuses on understanding learners’ interactions with music in authentic settings and developing intelligent systems that adapt to users’ cognitive and affective needs.
Ruilun Liu received his Ph.D. from the Faculty of Education at The University of Hong Kong. His research interests include music information retrieval, physiological signal processing, natural language processing, and multimodal learning analytics. His work focuses on user-centered music recommendation and multimodal modeling of human responses to music by integrating MIR features with physiological and behavioral data. He has contributed to the development of multimodal data pipelines that bridge audio analysis, machine learning, and learning analytics, supporting research on music-facilitated learning and human-centered MIR applications.
