← Search

Yun Zheng

23 accepted papers

2026

ShowTable: Unlocking Creative Table Visualization with Collaborative Reflection and Refinement

CVPR 2026

While existing generation and unified models excel at general image generation, they struggle with tasks requiring deep reasoning, planning, and precise data-to-visual mapping abilities beyond general scenarios. To push beyond the existing limitations, we introduce a new and challenging task: creati

Cited by 0SourcecodeScholar
2026

UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing

ICLR 2026poster

In this paper, we propose UniLIP, a unified framework that adapts CLIP for multimodal understanding, generation and editing. Although CLIP excels at understanding, it lacks reconstruction abilities required to be a unified visual encoder. However, previous CLIP-based unified methods fail to balance…

Cited by 0SourcecodeScholar
2025

Aligned Better, Listen Better for Audio-Visual Large Language Models

ICLR 2025poster

Audio is essential for multimodal video understanding. On the one hand, video inherently contains audio, which supplies complementary information to vision. Besides, video large language models (Video-LLMs) can encounter many audio-centric settings. However, existing Video-LLMs and Audio-Visual Larg…

Cited by 2SourcePDFScholar
2025

CAPability: A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness

NeurIPS 2025poster

Visual captioning benchmarks have become outdated with the emergence of modern multimodal large language models (MLLMs), as the brief ground-truth sentences and traditional metrics fail to assess detailed captions effectively. While recent benchmarks attempt to address this by focusing on keyword ex…

Cited by 0SourceScholar
2025

ContextHOI: Spatial Context Learning for Human-Object Interaction Detection

AAAI 2025technical

Spatial contexts, such as the backgrounds and surroundings, are considered critical in Human-Object Interaction (HOI) recognition, especially when the instance-centric foreground is blurred or occluded. Recent advancements in HOI detectors are usually built upon detection transformer pipelines. Whil…

Cited by 1SourcePDFScholar
2025

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding

ICCV 2025poster

In recent years, the introduction of Multi-modal Large Language Models (MLLMs) into video understanding tasks has become increasingly prevalent. However, how to effectively integrate temporal information remains a critical research focus. Traditional approaches treat spatial and temporal information…

Cited by 0SourcePDFScholar
2025

Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models

CVPR 2025poster

Recent Multi-modal Large Language Models (MLLMs) have been challenged by the computational overhead resulting from massive video frames, often alleviated through compression strategies. However, the visual content is not equally contributed to user instructions, existing strategies (e.g., average po…

2025

Orchestrating the Symphony of Prompt Distribution Learning for Human-Object Interaction Detection

AAAI 2025technical

Human-object interaction (HOI) detectors with popular query-transformer architecture have achieved promising performance. However, accurately identifying uncommon visual patterns and distinguishing between ambiguous HOIs continue to be difficult for them. We observe that these difficulties may arise…

Cited by 1SourcePDFScholar
2025

UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language Interface

NeurIPS 2025spotlight

Generalist models have achieved remarkable success in both language and vision-language tasks, showcasing the potential of unified modeling. However, effectively integrating fine-grained perception tasks like detection and segmentation into these models remains a significant challenge. This is prima…

Cited by 0SourcecodeScholar
2024

CoReS: Orchestrating the Dance of Reasoning and Segmentation

ECCV 2024poster

"The reasoning segmentation task, which demands a nuanced comprehension of intricate queries to accurately pinpoint object regions, is attracting increasing attention. However, Multi-modal Large Language Models (MLLM) often find it difficult to accurately localize the objects described in complex re…

2024

CrossMAE: Cross-Modality Masked Autoencoders for Region-Aware Audio-Visual Pre-Training

CVPR 2024poster

Learning joint and coordinated features across modalities is essential for many audio-visual tasks. Existing pre-training methods primarily focus on global information neglecting fine-grained features and positions leading to suboptimal performance in dense prediction tasks. To address this issue we…

Cited by 5SourcePDFScholar
2024

FuseTeacher: Modality-fused Encoders are Strong Vision Supervisors

ECCV 2024poster

"Learning visual representation with image-text datasets attracts a lot of attention in recent years. Existing approaches primarily rely on cross-modality supervision, and incorporate intra-modality supervision if necessary. They overlook the potential benefits of modality-fused supervision. Since m…

2024

Relevant Intrinsic Feature Enhancement Network for Few-Shot Semantic Segmentation

AAAI 2024technical

For few-shot semantic segmentation, the primary task is to extract class-specific intrinsic information from limited labeled data. However, the semantic ambiguity and inter-class similarity of previous methods limit the accuracy of pixel-level foreground-background classification. To alleviate these…

Cited by 16SourcePDFScholar
2023

Dual Mean-Teacher: An Unbiased Semi-Supervised Framework for Audio-Visual Source Localization

NeurIPS 2023poster

Audio-Visual Source Localization (AVSL) aims to locate sounding objects within video frames given the paired audio clips. Existing methods predominantly rely on self-supervised contrastive learning of audio-visual correspondence. Without any bounding-box annotations, they struggle to achieve precise…

2023

MomentDiff: Generative Video Moment Retrieval from Random to Real

NeurIPS 2023poster

Video moment retrieval pursues an efficient and generalized solution to identify the specific temporal segments within an untrimmed video that correspond to a given language description. To achieve this goal, we provide a generative diffusion-based framework called MomentDiff, which simulates a typi…

2023

Progressive Spatio-Temporal Prototype Matching for Text-Video Retrieval

ICCV 2023oral

The performance of text-video retrieval has been significantly improved by vision-language cross-modal learning schemes. The typical solution is to directly align the global video-level and sentence-level features during learning, which would ignore the intrinsic video-text relations, i.e., a text…

Cited by 46PDFcodeScholar
2023

RA-CLIP: Retrieval Augmented Contrastive Language-Image Pre-Training

CVPR 2023poster

Contrastive Language-Image Pre-training (CLIP) is attracting increasing attention for its impressive zero-shot recognition performance on different down-stream tasks. However, training CLIP is data-hungry and requires lots of image-text pairs to memorize various semantic concepts. In this paper, we…

Cited by 38SourcePDFScholar
2023

RLEG: Vision-Language Representation Learning with Diffusion-based Embedding Generation

ICML 2023poster

Vision-language representation learning models (e.g., CLIP) have achieved state-of-the-art performance on various downstream tasks, which usually need large-scale training data to learn discriminative representation. Recent progress on generative diffusion models (e.g., DALL-E 2) has demonstrated th…

Cited by 12SourcePDFScholar
2021

Exploring Visual-Audio Composition Alignment Network for Quality Fashion Retrieval in Video

ICASSP 2021accepted

Fashion retrieval in video suffers from the issues of imperfect visual representation and low quality of search results under the E-commercial circumstance. Previous works generally focus on searching the identical images from visual perspective only, but lack of leveraging multi-modal information f…

Cited by 0SourceScholar
2021

Few-Shot Incremental Learning With Continually Evolved Classifiers

CVPR 2021poster

Few-shot class-incremental learning (FSCIL) aims to design machine learning algorithms that can continually learn new concepts from a few data points, without forgetting knowledge of old classes. The difficulty lies in that limited data from new classes not only lead to significant overfitting issue…

Cited by 395PDFScholar
2020

Weakly Supervised Learning with Side Information for Noisy Labeled Images

ECCV 2020poster

In many real-world datasets, like WebVision, the performance of DNN based classier is often limited by the noisy labeled data. To tackle this problem, some image related side information, such as captions and tags, often reveal underlying relationships across images. In this paper, we present an eff…

Cited by 60SourcePDFScholar