← Search

Yuxin Song

14 accepted papers

2026

CoLoGen: Progressive Learning of Concept-Localization Duality for Unified Image Generation

CVPR 2026

Unified conditional image generation remains difficult because different tasks depend on fundamentally different internal representations. Some require conceptual understanding for semantic synthesis, while others rely on localization cues for spatial precision. Forcing these heterogeneous tasks to

Cited by 3SourcecodeScholar
2026

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders

ICLR 2026poster

Employing Multimodal Large Language Models (MLLMs) for long video understanding remains a challenging problem due to the dilemma between the substantial number of video frames (i.e., visual tokens) versus the limited context length of language models. Traditional uniform sampling often leads to sele…

Cited by 0SourcecodeScholar
2025

DistinctAD: Distinctive Audio Description Generation in Contexts

CVPR 2025highlight

Audio Descriptions (ADs) aim to provide a narration of a movie in text form, describing non-dialogue-related narratives, such as characters, actions, or scene establishment. Automatic generation of ADs remains challenging due to: i) the domain gap between movie-AD data and existing data used to trai…

Cited by 2SourcePDFScholar
2025

Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot Generalization

ICML 2025poster

Computer Vision (CV) has yet to fully achieve the zero-shot task generalization observed in Natural Language Processing (NLP), despite following many of the milestones established in NLP, such as large transformer models, extensive pre-training, and the auto-regression paradigm, among others. In thi…

2025

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI

ICCV 2025poster

Reasoning plays a crucial role in advancing Multimodal Large Language Models (MLLMs) toward Artificial General Intelligence.However, existing MLLM benchmarks often fall short in precisely and comprehensively evaluating long-chain reasoning abilities from three key aspects: (1) lack of difficulty and…

2025

Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

NeurIPS 2025spotlight

In this work, we aim to develop an MLLM that understands and solves questions by learning to create each intermediate step of the reasoning involved till the final answer. To this end, we propose Collective Monte Carlo Tree Search (CoMCTS), a new learning-to-reason method for MLLMs, which introduces…

Cited by 0SourcecodeScholar
2024

Automated Multi-level Preference for MLLMs

NeurIPS 2024poster

Current multimodal Large Language Models (MLLMs) suffer from ''hallucination'', occasionally generating responses that are not grounded in the input images. To tackle this challenge, one promising path is to utilize reinforcement learning from human feedback (RLHF), which steers MLLMs towards learni…

2024

Dense Connector for MLLMs

NeurIPS 2024poster

*Do we fully leverage the potential of visual encoder in Multimodal Large Language Models (MLLMs)?* The recent outstanding performance of MLLMs in multimodal understanding has garnered broad attention from both academia and industry. In the current MLLM rat race, the focus seems to be predominantly…

2024

MERG: Multi-Dimensional Edge Representation Generation Layer for Graph Neural Networks

ICASSP 2024accepted

Edges are essential in describing relationships among nodes. While existing graphs frequently use a single-value edge to describe association between each pair of node vectors, crucial relationships may be disregarded if they are not linearly correlated, which may limit graph analysis performance. A…

Cited by 0SourceScholar
2024

Multi-Level Graph Learning For Audio Event Classification And Human-Perceived Annoyance Rating Prediction

ICASSP 2024accepted

WHO’s report on environmental noise estimates that 22 M people suffer from chronic annoyance related to noise caused by audio events (AEs) from various sources. Annoyance may lead to health issues and adverse effects on metabolic and cognitive systems. In cities, monitoring noise levels does not pro…

Cited by 0SourceScholar
2024

Octopus: A Multi-modal LLM with Parallel Recognition and Sequential Understanding

NeurIPS 2024poster

A mainstream of Multi-modal Large Language Models (MLLMs) have two essential functions, i.e., visual recognition (e.g., grounding) and understanding (e.g., visual question answering). Presently, all these MLLMs integrate visual recognition and understanding in a same sequential manner in the LLM hea…

Cited by 1SourcePDFScholar
2023

UATVR: Uncertainty-Adaptive Text-Video Retrieval

ICCV 2023poster

With the explosive growth of web videos and emerging large-scale vision-language pre-training models, e.g., CLIP, retrieving videos of interest with text instructions has attracted increasing attention. A common practice is to transfer text-video pairs to the same embedding space and craft cross-mod…

Cited by 64PDFcodeScholar
2023

What Can Simple Arithmetic Operations Do for Temporal Modeling?

ICCV 2023poster

Temporal modeling plays a crucial role in understanding video content. To tackle this problem, previous studies built complicated temporal relations through time sequence thanks to the development of computationally powerful devices. In this work, we explore the potential of four simple arithmetic o…

Cited by 14PDFcodeScholar
2022

Medical Ultrasound Image Quality Assessment for Autonomous Robotic Screening

RA-L 2022

Autonomous ultrasound scanning robots have attracted the attention of researchers, and the real-time quality assessment of ultrasound images is the key technology of them. Existing robot systems usually use pixel-level feature statistical methods such as grayscale, confidence map, etc. However, in c

Cited by 18SourceScholar