MUSE-VA: A Large-Scale Synthetic Multimodal Dataset for Music Emotion Understanding and Generation

Jiahao Mei1, Yixuan Yuan2, Haoyu Gu3, Shimeng Li1, Leyuan Jin1, Mengyue Wu1*, Yue Ding2*

1 X-LANCE Lab, Shanghai Jiao Tong University, China 2 Shanghai Mental Health Center, Shanghai Jiao Tong University School of Medicine, China 3 School of Future Technology, South China University of Technology, China

MUSE-VA is a large-scale, VA-balanced multimodal music emotion dataset built through a five-stage LLM-driven pipeline to support emotion understanding, controllable music generation, and cross-modal affective research.

Abstract

Music affective computing relies on high-quality, large-scale, and emotionally balanced data resources, yet existing music emotion datasets are often limited in scale, uneven in valence-arousal coverage, narrow in annotation granularity, and lacking in cross-modal affective pairings. MUSE-VA introduces a large-scale multimodal music emotion dataset constructed from uniformly sampled VA coordinates through an affective and musical knowledge-injected five-stage LLM agent pipeline. The dataset contains 6,254 complete music pieces totaling 321.1 hours, with continuous VA annotations, 12 discrete emotion labels, music metadata, text descriptions, and one-to-one music-image pairs. Objective experiments and a human listening study show that MUSE-VA supports music emotion regression, controllable affective music generation, and cross-modal affective modeling.

Contributions

  • Balanced multimodal dataset. MUSE-VA provides 6,254 complete music pieces and 321.1 hours of audio with VA values, discrete emotion labels, text descriptions, music metadata, and paired images.
  • VA-driven construction. The dataset starts from uniformly sampled VA coordinates and uses a five-stage LLM distillation pipeline to control emotion-space coverage by design.
  • Knowledge-guided reasoning. Affective and musical knowledge bases guide emotion mapping, theme selection, music association, captioning, visual imagery, and consistency verification.
  • Broad validation. MUSE-VA is evaluated through emotion regression, affective music generation, music-image matching, and subjective listening studies.

Architecture

VA-Driven Five-Stage LLM Agent Pipeline

Overview of the five-stage pipeline for constructing MUSE-VA. The pipeline starts from uniformly sampled coordinates in the VA space and progressively accumulates contextual information across five stages, producing multimodal data entries with audio, images, and structured annotations.

Overview

Full dataset

Valence-Arousal Map

Click a point or an emotion label to highlight matching cases and view the details.

The full MUSE-VA dataset contains 6,254 tracks. This interactive map visualizes a random subset of 157 examples; access the complete dataset on Hugging Face Dataset.