← Search

Liang Yin

5 accepted papers

2026

VisuRiddles: Fine-grained Perception is a Primary Bottleneck for Multimodal Large Language Models in Abstract Visual Reasoning

ICLR 2026poster

Recent strides in multimodal large language models (MLLMs) have demonstrated significant progress in many reasoning tasks, but they still fail in Abstract Visual Reasoning (AVR) tasks. Our experimental findings indicate that the core bottleneck lies not only in the reasoning capabilities of MLLMs bu…

Cited by 0SourcecodeScholar
2025

A Novel Compressive Compound Word Encoding and Independent Word Attention for Symbolic Music Generation

ICASSP 2025accepted

Symbolic music generation involves using symbolic encoding to represent music pieces as token sequences and using neural sequence models to create music by generating sequences of tokens. Symbolic encodings are primarily categorized into two types: independent word encoding and compound word encodin…

Cited by 0SourceScholar
2025

ContraDiff: Planning Towards High Return States via Contrastive Learning

ICLR 2025poster

The performance of offline reinforcement learning (RL) is sensitive to the proportion of high-return trajectories in the offline dataset. However, in many simulation environments and real-world scenarios, there are large ratios of low-return trajectories rather than high-return trajectories, which m…

2025

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance

ICCV 2025poster

While large multi-modal models (LMMs) demonstrate promising capabilities in segmentation and comprehension, they still struggle with two limitations: inaccurate segmentation and hallucinated comprehension. These challenges stem primarily from constraints in weak visual comprehension and a lack of fi…

2025

MSTAR: Box-free Multi-query Scene Text Retrieval with Attention Recycling

NeurIPS 2025poster

Scene text retrieval has made significant progress with the assistance of accurate text localization. However, existing approaches typically require costly bounding box annotations for training. Besides, they mostly adopt a customized retrieval strategy but struggle to unify various types of querie…

Cited by 0SourcecodeScholar