← Search

Tingyu Weng

5 accepted papers

2026

UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing

ICLR 2026poster

In this paper, we propose UniLIP, a unified framework that adapts CLIP for multimodal understanding, generation and editing. Although CLIP excels at understanding, it lacks reconstruction abilities required to be a unified visual encoder. However, previous CLIP-based unified methods fail to balance…

Cited by 0SourcecodeScholar
2025

Aligned Better, Listen Better for Audio-Visual Large Language Models

ICLR 2025poster

Audio is essential for multimodal video understanding. On the one hand, video inherently contains audio, which supplies complementary information to vision. Besides, video large language models (Video-LLMs) can encounter many audio-centric settings. However, existing Video-LLMs and Audio-Visual Larg…

Cited by 2SourcePDFScholar
2025

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding

ICCV 2025poster

In recent years, the introduction of Multi-modal Large Language Models (MLLMs) into video understanding tasks has become increasingly prevalent. However, how to effectively integrate temporal information remains a critical research focus. Traditional approaches treat spatial and temporal information…

Cited by 0SourcePDFScholar
2025

UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language Interface

NeurIPS 2025spotlight

Generalist models have achieved remarkable success in both language and vision-language tasks, showcasing the potential of unified modeling. However, effectively integrating fine-grained perception tasks like detection and segmentation into these models remains a significant challenge. This is prima…

Cited by 0SourcecodeScholar
2023

Decompose Novel into Known: Part Concept Learning For 3D Novel Class Discovery

NeurIPS 2023poster

In this work, we address 3D novel class discovery (NCD) that discovers novel classes from an unlabeled dataset by leveraging the knowledge of disjoint known classes. The key challenge of 3D NCD is that learned features by known class recognition are heavily biased and hinder generalization to novel…

Cited by 1SourcePDFScholar