← Search

Yuxin Peng

61 accepted papers

2026

Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation

ICLR 2026poster

Recent advancements in Large Vision-Language Models (LVLMs) have demonstrated remarkable multimodal perception capabilities, garnering significant attention. While numerous evaluation studies have emerged, assessing LVLMs both holistically and on specialized tasks, fine-grained image tasks—fundament…

Cited by 0SourcecodeScholar
2026

CKDA: Cross-modality Knowledge Disentanglement and Alignment for Visible-Infrared Lifelong Person Re-identification

AAAI 2026technical

Lifelong person Re-IDentification (LReID) aims to match the same person employing continuously collected individual data from different scenarios. To achieve continuous all-day person matching across day and night, Visible-Infrared Lifelong person Re-IDentification (VI-LReID) focuses on sequential t

Cited by 0SourcePDFScholar
2026

Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning

ICLR 2026poster

Any entity in the visual world can be hierarchically grouped based on shared characteristics and mapped to fine-grained sub-categories. While Multi-modal Large Language Models (MLLMs) achieve strong performance on coarse-grained visual tasks, they often struggle with Fine-Grained Visual Recognition…

Cited by 0SourcecodeScholar
2026

LIMSSR: LLM-Driven Sequence-to-Score Reasoning under Training-Time Incomplete Multimodal Observations

ICML 2026spotlight

Real-world multimodal learning is often hindered by missing modalities. While Incomplete Multimodal Learning (IML) has gained traction, existing methods typically rely on the unrealistic assumption of full-modal availability during training to provide reconstruction supervision or cross-modal priors…

Cited by 0SourceScholar
2026

Learning Taxonomic Trees with Hierarchical Representation Regularization for Large Multimodal Models

ICML 2026poster

Taxonomies provide key information about the semantic relationships between concepts and the inherent organization of vision and language. Despite their impressive capabilities, large multimodal models (LMMs) often lack taxonomic knowledge, leading to low hierarchical visual recognition (HVR) consis…

Cited by 0SourceScholar
2026

Learning to Generate via Understanding: Understanding-Driven Intrinsic Rewarding for Unified Multimodal Models

CVPR 2026

Recently, unified multimodal models (UMMs) have made remarkable progress in integrating visual understanding and generation, demonstrating strong potential for complex text-to-image (T2I) tasks. Despite their theoretical promise, a persistent capability gap exists: UMMs typically exhibit superior vi

Cited by 0SourcecodeScholar
2026

OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding

CVPR 2026

Video Temporal Grounding (VTG), the task of localizing video segments from text queries, struggles in open-world settings due to limited dataset scale and semantic diversity, causing performance gaps between common and rare concepts. To overcome these limitations, we introduce OmniVTG, a new large-s

Cited by 0SourcecodeScholar
2026

Repurposing 3D Generative Model for Autoregressive Layout Generation

CVPR 2026

We introduce LaviGen, a framework that repurposes 3D generative models for 3D layout generation. Unlike previous methods that infer object layouts from textual descriptions, LaviGen operates directly in the native 3D space, formulating layout generation as an autoregressive process that explicitly m

Cited by 0SourcecodeScholar
2026

TaRO: Temporal-Aware Reasoning Optimization for Video Temporal Grounding

ICML 2026poster

Multi-modal Large Language Models (MLLMs) have achieved remarkable progress in video temporal grounding (VTG) with the introduction of reinforcement learning (RL) for generating reasoning paths. However, existing models often produce superficial reasoning, such as providing generic video description…

Cited by 0SourceScholar
2026

Taming I2V models for Image HOI Editing: A Cognitive Benchmark and Agentic Self-Correcting Framework

ICML 2026poster

Current image editing methods excels at static attributes but fails at complex Human-Object Interactions (HOI), a critical challenge unaddressed by existing benchmarks that conflate HOI with static attributes, relying on global metrics incapable of simultaneously assessing dynamic interaction validi…

Cited by 0SourceScholar
2026

Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models

CVPR 2026

A high-performing, general-purpose visual understanding model should map visual inputs to a taxonomic tree of labels, identify novel categories beyond the training set for which few or no publicly available images exist. Large Multimodal Models (LMMs) have achieved remarkable progress in fine-graine

Cited by 0SourcecodeScholar
2026

Venus: Benchmarking and Empowering Multimodal Large Language Models for Aesthetic Guidance and Cropping

CVPR 2026

The widespread use of smartphones has made photography ubiquitous, yet a clear gap remains between ordinary users and professional photographers, who can identify aesthetic issues and provide actionable shooting guidance during capture. We define this capability as aesthetic guidance (AG) --- an ess

Cited by 0SourcecodeScholar
2025

Analyzing and Boosting the Power of Fine-Grained Visual Recognition for Multi-modal Large Language Models

ICLR 2025poster

Multi-modal large language models (MLLMs) have shown remarkable abilities in various visual understanding tasks. However, MLLMs still struggle with fine-grained visual recognition (FGVR), which aims to identify subordinate-level categories from images. This can negatively impact more advanced capabi…

2025

Balancing Preservation and Modification: A Region and Semantic Aware Metric for Instruction-Based Image Editing

ICML 2025poster

Instruction-based image editing, which aims to modify the image faithfully towards instruction while preserving irrelevant content unchanged, has made advanced progresses. However, there still lacks a comprehensive metric for assessing the editing quality. Existing metrics either require high costs…

2025

ConMo: Controllable Motion Disentanglement and Recomposition for Zero-Shot Motion Transfer

CVPR 2025poster

The development of Text-to-Video (T2V) generation has made motion transfer possible, enabling the control of video motion based on existing footage. However, current methods have two limitations: 1) struggle to handle multi-subjects videos, failing to transfer specific subject motion; 2) struggle to…

2025

DASK: Distribution Rehearsing via Adaptive Style Kernel Learning for Exemplar-Free Lifelong Person Re-Identification

AAAI 2025technical

Lifelong person re-identification (LReID) is an important but challenging task that suffers from catastrophic forgetting due to significant domain gaps between training steps. Existing LReID approaches typically rely on data replay and knowledge distillation to mitigate this issue. However, data rep…

2025

DKC: Differentiated Knowledge Consolidation for Cloth-Hybrid Lifelong Person Re-identification

CVPR 2025poster

Lifelong person re-identification (LReID) aims to match the same person using sequentially collected data. However, due to the long-term nature of lifelong learning, the inevitable changes in human clothes prevent the model from relying on unified discriminative information (e.g., clothing style) to…

2025

DyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understanding

CVPR 2025highlight

Humans can effortlessly locate desired objects in cluttered environments, relying on a cognitive mechanism known as visual search to efficiently filter out irrelevant information and focus on task related regions. Inspired by this process, we propose DyFo (Dynamic Focus), a training-free dynamic foc…

2025

EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing

CVPR 2025highlight

Given a piece of text, a video clip, and a reference audio, the movie dubbing task aims to generate speech that aligns with the video while cloning the desired voice. The existing methods have two primary deficiencies: (1) They struggle to simultaneously hold audio-visual sync and achieve clear pron…

2025

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding

ICCV 2025poster

In this paper, we tackle the task of online video temporal grounding (OnVTG), which requires the model to locate events related to a given text query within a video stream. Unlike regular video temporal grounding, OnVTG requires the model to make predictions without observing future frames. As onlin…

2025

MAI: A Multi-turn Aggregation-Iteration Model for Composed Image Retrieval

ICLR 2025poster

Multi-Turn Composed Image Retrieval (MTCIR) addresses a real-world scenario where users iteratively refine retrieval results by providing additional information until a target meeting all their requirements is found. Existing methods primarily achieve MTCIR through a "multiple single-turn" paradigm,…

Cited by 0SourcePDFScholar
2025

Open-Vocabulary HOI Detection with Interaction-aware Prompt and Concept Calibration

ICCV 2025poster

Open Vocabulary Human-Object Interaction (HOI) detection aims to detect interactions between humans and objects while generalizing to novel interaction classes beyond the training set. Current methods often rely on Vision and Language Models (VLMs) but face challenges due to suboptimal image encoder…

2025

PosterO: Structuring Layout Trees to Enable Language Models in Generalized Content-Aware Layout Generation

CVPR 2025poster

In poster design, content-aware layout generation is crucial for automatically arranging visual-textual elements on the given image. With limited training data, existing work focused on image-centric enhancement. However, this neglects the diversity of layouts and fails to cope with shape-variant el…

Cited by 0SourcePDFScholar
2025

SCAP: Transductive Test-Time Adaptation via Supportive Clique-based Attribute Prompting

CVPR 2025poster

Vision-language models (VLMs) encounter considerable challenges when adapting to domain shifts stemming from changes in data distribution. Test-time adaptation (TTA) has emerged as a promising approach to enhance VLM performance under such conditions. In practice, test data often arrives in batches,…

2025

STOP: Integrated Spatial-Temporal Dynamic Prompting for Video Understanding

CVPR 2025poster

Pre-trained on tremendous image-text pairs, vision-language models like CLIP have demonstrated promising zero-shot generalization across numerous image-based tasks. However, extending these capabilities to video tasks remains challenging due to limited labeled video data and high training costs. Rec…

2025

Scan-and-Print: Patch-level Data Summarization and Augmentation for Content-aware Layout Generation in Poster Design

IJCAI 2025

In AI-empowered poster design, content-aware layout generation is crucial for the on-image arrangement of visual-textual elements, e.g., logo, text, and underlay. To perceive the background images, existing work demanded a high parameter count that far exceeds the size of available training data, wh

2025

Selective Visual Prompting in Vision Mamba

AAAI 2025technical

Pre-trained Vision Mamba~(Vim) models have demonstrated exceptional performance across various computer vision tasks in a computationally efficient manner, attributed to their unique design of selective state space models. To further extend their applicability to diverse downstream vision tasks, Vim…

2025

TRKT: Weakly Supervised Dynamic Scene Graph Generation with Temporal-enhanced Relation-aware Knowledge Transferring

ICCV 2025poster

Dynamic Scene Graph Generation (DSGG) aims to create a scene graph for each video frame by detecting objects and predicting their relationships. Weakly Supervised DSGG (WS-DSGG) reduces annotation workload by using an un- localized scene graph from a single frame per video for training. Existing WS-…

2025

UPP: Unified Point-Level Prompting for Robust Point Cloud Analysis

ICCV 2025poster

Pre-trained point cloud analysis models have shown promising advancements in various downstream tasks, yet their effectiveness is typically suffering from low-quality point cloud (i.e., noise and incompleteness), which is a common issue in real scenarios due to casual object occlusions and unsatisfa…

Cited by 0SourcePDFScholar
2024

Comprehensive Visual Grounding for Video Description

AAAI 2024technical

The grounding accuracy of existing video captioners is still behind the expectation. The majority of existing methods perform grounded video captioning on sparse entity annotations, whereas the captioning accuracy often suffers from degenerated object appearances on the annotated area such as motion…

Cited by 2SourcePDFScholar
2024

Continual Vision-Language Retrieval via Dynamic Knowledge Rectification

AAAI 2024technical

The recent large-scale pre-trained models like CLIP have aroused great concern in vision-language tasks. However, when required to match image-text data collected in a streaming manner, namely Continual Vision-Language Retrieval (CVRL), their performances are still limited due to the catastrophic fo…

2024

DART: Dual-Modal Adaptive Online Prompting and Knowledge Retention for Test-Time Adaptation

AAAI 2024technical

As an up-and-coming area, CLIP-based pre-trained vision-language models can readily facilitate downstream tasks through the zero-shot or few-shot fine-tuning manners. However, they still face critical challenges in test-time generalization due to the shifts between the training and test data distrib…

Cited by 10SourcePDFScholar
2024

Distribution-aware Knowledge Prototyping for Non-exemplar Lifelong Person Re-identification

CVPR 2024poster

Lifelong person re-identification (LReID) suffers from the catastrophic forgetting problem when learning from non-stationary data. Existing exemplar-based and knowledge distillation-based LReID methods encounter data privacy and limited acquisition capacity respectively. In this paper we instead int…

2024

Exploring Conditional Multi-Modal Prompts for Zero-shot HOI Detection

ECCV 2024poster

"Zero-shot Human-Object Interaction (HOI) detection has emerged as a frontier topic due to its capability to detect HOIs beyond a predefined set of categories. This task entails not only identifying the interactiveness of human-object pairs and localizing them but also recognizing both seen and unse…

2024

FCS: Feature Calibration and Separation for Non-Exemplar Class Incremental Learning

CVPR 2024poster

Non-Exemplar Class Incremental Learning (NECIL) involves learning a classification model on a sequence of data without access to exemplars from previously encountered old classes. Such a stringent constraint always leads to catastrophic forgetting of the learned knowledge. Currently existing methods…

2024

FashionERN: Enhance-and-Refine Network for Composed Fashion Image Retrieval

AAAI 2024technical

The goal of composed fashion image retrieval is to locate a target image based on a reference image and modified text. Recent methods utilize symmetric encoders (e.g., CLIP) pre-trained on large-scale non-fashion datasets. However, the input for this task exhibits an asymmetric nature, where the ref…

Cited by 5SourcePDFScholar
2024

FineFMPL: Fine-grained Feature Mining Prompt Learning for Few-Shot Class Incremental Learning

IJCAI 2024poster

Few-Shot Class Incremental Learning (FSCIL) aims to continually learn new classes with few training samples without forgetting already learned old classes. Existing FSCIL methods generally fix the backbone network in incremental sessions to achieve a balance between suppressing forgetting old classe…

2024

FinePOSE: Fine-Grained Prompt-Driven 3D Human Pose Estimation via Diffusion Models

CVPR 2024highlight

The 3D Human Pose Estimation (3D HPE) task uses 2D images or videos to predict human joint coordinates in 3D space. Despite recent advancements in deep learning-based methods they mostly ignore the capability of coupling accessible texts and naturally feasible knowledge of humans missing out on valu…

2024

FineParser: A Fine-grained Spatio-temporal Action Parser for Human-centric Action Quality Assessment

CVPR 2024poster

Existing action quality assessment (AQA) methods mainly learn deep representations at the video level for scoring diverse actions. Due to the lack of a fine-grained understanding of actions in videos they harshly suffer from low credibility and interpretability thus insufficient for stringent applic…

2024

FineSports: A Multi-person Hierarchical Sports Video Dataset for Fine-grained Action Understanding

CVPR 2024poster

Fine-grained action analysis in multi-person sports is complex due to athletes' quick movements and intense physical confrontations which result in severe visual obstructions in most scenes. In addition accessible multi-person sports video datasets lack fine-grained action annotations in both space…

2024

Learning Continual Compatible Representation for Re-indexing Free Lifelong Person Re-identification

CVPR 2024poster

Lifelong Person Re-identification (L-ReID) aims to learn from sequentially collected data to match a person across different scenes. Once an L-ReID model is updated using new data all historical images in the gallery are required to be re-calculated to obtain new features for testing known as "re-in…

2024

Training-free Video Temporal Grounding using Large-scale Pre-trained Models

ECCV 2024poster

"Video temporal grounding aims to identify video segments within untrimmed videos that are most relevant to a given natural language query. Existing video temporal localization models rely on specific datasets for training, with high data collection costs, but exhibit poor generalization capability…

2023

Confidence-aware Pseudo-label Learning for Weakly Supervised Visual Grounding

ICCV 2023poster

Visual grounding aims at localizing the target object in image which is most related to the given free-form natural language query. As labeling the position of target object is labor-intensive, the weakly supervised methods, where only image-sentence annotations are required during model training ha…

Cited by 12PDFcodeScholar
2023

Efficient Adaptive Human-Object Interaction Detection with Concept-guided Memory

ICCV 2023poster

Human Object Interaction (HOI) detection aims to localize and infer the relationships between a human and an object. Arguably, training supervised models for this task from scratch presents challenges due to the performance drop over rare classes and the high computational cost and time required to…

Cited by 24PDFcodeScholar
2023

Generating Structured Pseudo Labels for Noise-resistant Zero-shot Video Sentence Localization

ACL 2023long

Video sentence localization aims to locate moments in an unstructured video according to a given natural language query. A main challenge is the expensive annotation costs and the annotation bias. In this work, we study video sentence localization in a zero-shot setting, which learns with only video…

2023

Masked Retraining Teacher-Student Framework for Domain Adaptive Object Detection

ICCV 2023poster

Domain adaptive Object Detection (DAOD) leverages a labeled domain (source) to learn an object detector generalizing to a novel domain without annotation (target). Recent advances use a teacher-student framework, i.e., a student model is supervised by the pseudo labels from a teacher model. Though g…

Cited by 31PDFcodeScholar
2023

Phrase-Level Temporal Relationship Mining for Temporal Sentence Localization

AAAI 2023technical

In this paper, we address the problem of video temporal sentence localization, which aims to localize a target moment from videos according to a given language query. We observe that existing models suffer from a sheer performance drop when dealing with simple phrases contained in the sentence. It r…

2023

PosterLayout: A New Benchmark and Approach for Content-Aware Visual-Textual Presentation Layout

CVPR 2023poster

Content-aware visual-textual presentation layout aims at arranging spatial space on the given canvas for pre-defined elements, including text, logo, and underlay, which is a key to automatic template-free creative graphic design. In practical applications, e.g., poster designs, the canvas is origina…

2023

Scanning Only Once: An End-to-end Framework for Fast Temporal Grounding in Long Videos

ICCV 2023poster

Video temporal grounding aims to pinpoint a video segment that matches the query description. Despite the recent advance in short-form videos (e.g., in minutes), temporal grounding in long videos (e.g., in hours) is still at its early stage. To address this challenge, a common practice is to employ…

Cited by 18PDFcodeScholar
2022

An Embarrassingly Simple Approach to Semi-Supervised Few-Shot Learning

NeurIPS 2022accept

Semi-supervised few-shot learning consists in training a classifier to adapt to new tasks with limited labeled data and a fixed quantity of unlabeled data. Many sophisticated methods have been developed to address the challenges this problem comprises. In this paper, we propose a simple but quite ef…

Cited by 19SourcePDFScholar
2022

Weakly Supervised Temporal Sentence Grounding With Gaussian-Based Contrastive Proposal Learning

CVPR 2022poster

Temporal sentence grounding aims to detect the most salient moment corresponding to the natural language query from untrimmed videos. As labeling the temporal boundaries is labor-intensive and subjective, the weakly-supervised methods have recently received increasing attention. Most of the existing…

Cited by 109PDFcodeScholar
2015

The Application of Two-Level Attention Models in Deep Convolutional Neural Network for Fine-Grained Image Classification

CVPR 2015poster

Fine-grained classification is challenging because categories can only be discriminated by subtle and local differences. Variances in the pose, scale or rotation usually make the problem more difficult. Most fine-grained classification systems follow the pipeline of finding foreground object or obje…

Cited by 1090SourcePDFScholar