Abstract
Music affective computing relies on high-quality, large-scale, and emotionally balanced data resources, yet existing music emotion datasets are often limited in scale, uneven in valence-arousal coverage, narrow in annotation granularity, and lacking in cross-modal affective pairings. MUSE-VA introduces a large-scale multimodal music emotion dataset constructed from uniformly sampled VA coordinates through an affective and musical knowledge-injected five-stage LLM agent pipeline. The dataset contains 6,254 complete music pieces totaling 321.1 hours, with continuous VA annotations, 12 discrete emotion labels, music metadata, text descriptions, and one-to-one music-image pairs. Objective experiments and a human listening study show that MUSE-VA supports music emotion regression, controllable affective music generation, and cross-modal affective modeling.