← Search

Sungkyun Chang

4 accepted papers

2026

CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction

ICML 2026poster

While music generation models have evolved to handle complex multimodal inputs mixing text, lyrics, and reference audio, evaluation mechanisms have lagged behind, remaining fragmented and narrowly focused. In this paper, we bridge this critical gap by establishing a comprehensive ecosystem for Compo…

Cited by 0SourceScholar
2026

OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs

ICLR 2026poster

Recent advances in multimodal large language models (MLLMs) have demonstrated substantial potential in video understanding. However, existing benchmarks fail to comprehensively evaluate synergistic reasoning capabilities across audio and visual modalities, often neglecting either one of the modaliti…

Cited by 0SourcecodeScholar
2021

Neural Audio Fingerprint for High-Specific Audio Retrieval Based on Contrastive Learning

ICASSP 2021accepted

Most of existing audio fingerprinting systems have limitations to be used for high-specific audio retrieval at scale. In this work, we generate a low-dimensional representation from a short unit segment of audio, and couple this fingerprint with a fast maximum inner-product search. To this end, we p…

Cited by 0SourceScholar
2018

Cover Song Identification Using Song-to-Song Cross-Similarity Matrix with Convolutional Neural Network

ICASSP 2018accepted

In this paper, we propose a cover song identification algorithm using a convolutional neural network (CNN). We first train the CNN model to classify any non-/cover relationship, by feeding a cross-similarity matrix that is generated from a pair of songs as an input. Our main idea is to use the CNN o…

Cited by 0SourceScholar