← Search

HAOYU CHEN

48 accepted papers

2026

Adaptive Action Chunking at Inference-time for Vision-Language-Action Models

CVPR 2026

In Vision-Language-Action (VLA) models, action chunking (i.e., executing a sequence of actions without intermediate replanning) is a key technique to improve robotic manipulation abilities. However, a large chunk size reduces the model's responsiveness to new information, while a small one increases

Cited by 0SourcecodeScholar
2026

Bias Is a Subspace, Not a Coordinate: A Geometric Rethinking of Post-hoc Debiasing in Vision-Language Models

CVPR 2026

Vision-Language Models (VLMs) have become indispensable for multimodal reasoning, yet their representations often encode and amplify demographic biases, resulting in biased associations and misaligned predictions in downstream tasks. Such behavior undermines fairness and distorts the intended alignm

Cited by 0SourcecodeScholar
2026

EmoPrefer: Can Large Language Models Understand Human Emotion Preferences?

ICLR 2026poster

Descriptive Multimodal Emotion Recognition (DMER) has garnered increasing research attention. Unlike traditional discriminative paradigms that rely on predefined emotion taxonomies, DMER aims to describe human emotional state using free-form natural language, enabling finer-grained and more interpre…

Cited by 0SourcecodeScholar
2026

GraphP-FL: Personalized Federated Graph Learning via Dynamic Structure Awareness and Fisher Information Elastic Alignment

ICML 2026poster

Federated Graph Learning (FGL) enables distributed clients to collaboratively train graph neural networks while strictly preserving data privacy.However, existing FGL methods implicitly assume the reliability of local graph structures and lack elastic awareness of parameter importance during model a…

Cited by 0SourceScholar
2026

M-Loss: Quantifying Model Merging Compatibility with Limited Unlabeled Data

AAAI 2026technical

Training of large-scale models is both computationally intensive and often constrained by the availability of labeled data. Model merging offers a compelling alternative by directly integrating the weights of multiple source models without requiring additional data or extensive training. However, co

Cited by 0SourcePDFScholar
2026

Multi-Object System Identification from Videos

ICLR 2026poster

We introduce the challenging problem of multi-object system identification from videos, for which prior methods are ill-suited due to their focus on single-object scenes or discrete material classification with a fixed set of material prototypes. To address this, we propose MOSIV, a new framework th…

Cited by 0SourceScholar
2026

PosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified Framework

ICLR 2026poster

Generating aesthetic posters is more challenging than simple design images: it requires not only precise text rendering but also the seamless integration of abstract artistic content, striking layouts, and overall stylistic harmony. To address this, we propose PosterCraft, a unified framework that a…

Cited by 0SourcecodeScholar
2026

SWIFT:A General Sensitive Weight Identification Framework for Fast Sensor-Transfer Pansharpening

AAAI 2026technical

Although deep learning-based methods have achieved promising performance in Pansharpening, they generally suffer from severe performance degradation when applied to data from unseen sensors. Existing cross-domain strategies, including retraining, fine-tuning, and zero-shot methods, fail to simultane

Cited by 0SourcePDFScholar
2025

3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark

ICCV 2025poster

3D spatial reasoning is the ability to analyze and interpret the positions, orientations, and spatial relationships of objects within the 3D space. This allows models to develop a comprehensive understanding of the 3D scene, enabling their applicability to a broader range of applications, such as au…

Cited by 0SourcePDFScholar
2025

AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models

ICML 2025oral

The emergence of multimodal large language models (MLLMs) advances multimodal emotion recognition (MER) to the next level—from naive discriminative tasks to complex emotion understanding with advanced video understanding abilities and natural language description. However, the current community suff…

2025

Beyond Generation: A Diffusion-based Low-level Feature Extractor for Detecting AI-generated Images

CVPR 2025poster

The prevalence of AI-generated images has evoked concerns regarding the potential misuse of image generation technologies. In response, numerous detection methods aim to identify AI-generated images by analyzing generative artifacts. Unfortunately, most detectors quickly become obsolete with the dev…

Cited by 0SourcePDFScholar
2025

CamEdit: Continuous Camera Parameter Control for Photorealistic Image Editing

NeurIPS 2025poster

Recent advances in diffusion models have substantially improved text-driven image editing. However, existing frameworks based on discrete textual tokens struggle to support continuous control over camera parameters and smooth transitions in visual effects. These limitations hinder their applications…

Cited by 0SourceScholar
2025

Deep Change Monitoring: A Hyperbolic Representative Learning Framework and a Dataset for Long-term Fine-grained Tree Change Detection

CVPR 2025highlight

In environmental protection, tree monitoring plays an essential role in maintaining and improving ecosystem health. However, precise monitoring is challenging because existing datasets fail to capture continuous fine-grained changes in trees due to low-resolution images and high acquisition costs. I…

2025

FreeNet: Liberating Depth-Wise Separable Operations for Building Faster Mobile Vision Architectures

AAAI 2025technical

In the pursuit of efficient vision architectures, substantial efforts have been devoted to optimizing operator efficiency. Depth-wise separable operators, such as DWConv, are found cheap in both FLOPs and parameters. As a result, they are increasingly incorporated into efficient backbones, trading f…

Cited by 0SourcePDFScholar
2025

From Laboratory to Real World: A New Benchmark Towards Privacy-Preserved Visible-Infrared Person Re-Identification

CVPR 2025poster

Aiming to match pedestrian images captured under varying lighting conditions, visible-infrared person re-identification (VI-ReID) has drawn intensive research attention and achieved promising results. However, in real-world surveillance contexts, data is distributed across multiple devices/entities,…

2025

GenHaze: Pioneering Controllable One-Step Realistic Haze Generation for Real-World Dehazing

ICCV 2025poster

Real-world image dehazing is crucial for enhancing visual quality in computer vision applications. However, existing physics-based haze generation paradigms struggle to model the complexities of real-world haze and lack controllability, limiting the performance of existing baselines on real-world im…

Cited by 0SourcePDFScholar
2025

Hiding Images in Diffusion Models by Editing Learned Score Functions

CVPR 2025poster

Hiding data using neural networks (i.e., neural steganography) has achieved remarkable success across both discriminative classifiers and generative adversarial networks. However, the potential of data hiding in diffusion models remains relatively unexplored. Current methods exhibit limitations in a…

2025

JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent

NeurIPS 2025poster

Photo retouching has become integral to contemporary visual storytelling, enabling users to capture aesthetics and express creativity. While professional tools such as Adobe Lightroom offer powerful capabilities, they demand substantial expertise and manual effort. In contrast, existing AI-based sol…

Cited by 0SourceScholar
2025

JarvisIR: Elevating Autonomous Driving Perception with Intelligent Image Restoration

CVPR 2025poster

Vision-centric perception systems often struggle with unpredictable and coupled weather degradations in the wild. Current solutions are often limited, as they either depend on specific degradation priors or suffer from significant domain gaps. To enable robust and autonomous operation in real-world…

Cited by 2SourcePDFScholar
2025

LCFed: An Efficient Clustered Federated Learning Framework for Heterogeneous Data

ICASSP 2025accepted

Clustered federated learning (CFL) addresses the performance challenges posed by data heterogeneity in federated learning (FL) by organizing edge devices with similar data distributions into clusters, enabling collaborative model training tailored to each group. However, existing CFL approaches stri…

Cited by 0SourceScholar
2025

Learning Binary-Antithetical Information Bottleneck for Generalizable Face Anti-Spoofing

ICASSP 2025accepted

We investigate generalizable face anti-spoofing (FAS) using information bottleneck theory. As generalizable FAS aims to detect spoofing in unseen scenarios, it has recently gained significant attention. Existing methods often use adversarial strategies or auxiliary modules to learn domain-invariant…

Cited by 0SourceScholar
2025

OV-MER: Towards Open-Vocabulary Multimodal Emotion Recognition

ICML 2025poster

Multimodal Emotion Recognition (MER) is a critical research area that seeks to decode human emotions from diverse data modalities. However, existing machine learning methods predominantly rely on predefined emotion taxonomies, which fail to capture the inherent complexity, subtlety, and multi-apprai…

2025

POSTA: A Go-to Framework for Customized Artistic Poster Generation

CVPR 2025poster

Poster design is a critical medium for visual communication. Prior work has explored automatic poster design using deep learning techniques, but these approaches lack text accuracy, user customization, and aesthetic appeal, limiting their applicability in artistic domains such as movies and exhibiti…

Cited by 4SourcePDFScholar
2025

PocketSR: The Super-Resolution Expert in Your Pocket Mobiles

NeurIPS 2025poster

Real-world image super-resolution (RealSR) aims to enhance the visual quality of in-the-wild images, such as those captured by mobile phones. While existing methods leveraging large generative models demonstrate impressive results, the high computational cost and latency make them impractical for ed…

Cited by 0SourceScholar
2025

PreGenie: An Agentic Framework for High-quality Visual Presentation Generation

EMNLP 2025

Visual presentations are vital for effective communication. Early attempts to automate their creation using deep learning often faced issues such as poorly organized layouts, inaccurate text summarization, and a lack of image understanding, leading to mismatched visuals and text. These limitations r

Cited by 0SourcePDFScholar
2025

PromptHaze: Prompting Real-world Dehazing via Depth Anything Model

AAAI 2025technical

Real-world image dehazing remains a challenging task due to the diverse nature of haze degradation and the lack of large-scale paired datasets. Existing methods based on hand-crafted priors or generative priors struggle to recover accurate backgrounds and fine details from dense haze regions. In thi…

Cited by 0SourcePDFScholar
2025

Toward Material-Agnostic System Identification from Videos

ICCV 2025poster

System identification from videos aims to recover object geometry and governing physical laws. Existing methods integrate differentiable rendering with simulation but rely on predefined material priors, limiting their ability to handle unknown ones. We introduce MASIV, the first vision-based framewo…

2025

Turbo2K: Towards Ultra-Efficient and High-Quality 2K Video Synthesis

ICCV 2025poster

Demand for 2K video synthesis is rising with increasing consumer expectations for ultra-clear visuals.While diffusion transformers (DiTs) have demonstrated remarkable capabilities in high-quality video generation, scaling them to 2K resolution remains computationally prohibitive due to quadratic gro…

Cited by 0SourcePDFScholar
2024

CoSeR: Bridging Image and Language for Cognitive Super-Resolution

CVPR 2024poster

Existing super-resolution (SR) models primarily focus on restoring local texture details often neglecting the global semantic information within the scene. This oversight can lead to the omission of crucial semantic details or the introduction of inaccurate textures during the recovery process. In o…

2024

Differentiable Auxiliary Learning for Sketch Re-Identification

AAAI 2024technical

Sketch re-identification (Re-ID) seeks to match pedestrians' photos from surveillance videos with corresponding sketches. However, we observe that existing works still have two critical limitations: (i) cross- and intra-modality discrepancies hinder the extraction of modality-shared features, (ii) s…

Cited by 10SourcePDFScholar
2024

Domain Shifting: A Generalized Solution for Heterogeneous Cross-Modality Person Re-Identification

ECCV 2024poster

"Cross-modality person re-identification (ReID) is a challenging task that aims to match cross-modality pedestrian images across multiple camera views. Existing methods are tailored to specific tasks and perform well for visible-infrared or visible-sketch ReID. However, the performance exhibits a no…

Cited by 6SourcePDFScholar
2024

Learned HDR Image Compression for Perceptually Optimal Storage and Display

ECCV 2024poster

"High dynamic range (HDR) capture and display have seen significant growth in popularity driven by the advancements in technology and increasing consumer demand for superior image quality. As a result, HDR image compression is crucial to fully realize the benefits of HDR imaging without suffering fr…

2024

Low-Res Leads the Way: Improving Generalization for Super-Resolution by Self-Supervised Learning

CVPR 2024poster

For image super-resolution (SR) bridging the gap between the performance on synthetic datasets and real-world degradation scenarios remains a challenge. This work introduces a novel "Low-Res Leads the Way" (LWay) training framework merging Supervised Pre-training with Self-supervised Learning to enh…

Cited by 15SourcePDFScholar
2024

RestoreAgent: Autonomous Image Restoration Agent via Multimodal Large Language Models

NeurIPS 2024poster

Natural images captured by mobile devices often suffer from multiple types of degradation, such as noise, blur, and low light. Traditional image restoration methods require manual selection of specific tasks, algorithms, and execution sequences, which is time-consuming and may yield suboptimal resul…

Cited by 6SourcePDFScholar
2024

Semi-Supervised Video Desnowing Network via Temporal Decoupling Experts and Distribution-Driven Contrastive Regularization

ECCV 2024poster

"Snow degradations present formidable challenges to the advancement of computer vision tasks by the undesirable corruption in outdoor scenarios. While current deep learning-based desnowing approaches achieve success on synthetic benchmark datasets, they struggle to restore out-of-distribution real-w…

2024

UltraPixel: Advancing Ultra High-Resolution Image Synthesis to New Peaks

NeurIPS 2024poster

Ultra-high-resolution image generation poses great challenges, such as increased semantic planning complexity and detail synthesis difficulties, alongside substantial training resource demands. We present UltraPixel, a novel architecture utilizing cascade diffusion models to generate high-quality im…

Cited by 17SourcePDFScholar
2023

Crafting Training Degradation Distribution for the Accuracy-Generalization Trade-off in Real-World Super-Resolution

ICML 2023poster

Super-resolution (SR) techniques designed for real-world applications commonly encounter two primary challenges: generalization performance and restoration accuracy. We demonstrate that when methods are trained using complex, large-range degradations to enhance generalization, a decline in accuracy…

Cited by 24SourcePDFScholar
2023

LART: Neural Correspondence Learning with Latent Regularization Transformer for 3D Motion Transfer

NeurIPS 2023poster

3D motion transfer aims at transferring the motion from a dynamic input sequence to a static 3D object and outputs an identical motion of the target with high-fidelity and realistic visual effects. In this work, we propose a novel 3D Transformer framework called LART for 3D motion transfer. With car…

2023

Learning a Deep Color Difference Metric for Photographic Images

CVPR 2023poster

Most well-established and widely used color difference (CD) metrics are handcrafted and subject-calibrated against uniformly colored patches, which do not generalize well to photographic images characterized by natural scene complexities. Constructing CD formulae for photographic images is still an…

2023

Masked Image Training for Generalizable Deep Image Denoising

CVPR 2023poster

When capturing and storing images, devices inevitably introduce noise. Reducing this noise is a critical task called image denoising. Deep learning has become the de facto method for image denoising, especially with the emergence of Transformer-based models that have achieved notable state-of-the-ar…

2023

Snow Removal in Video: A New Dataset and A Novel Method

ICCV 2023poster

Snowfall is a common weather phenomenon that can severely affect computer vision tasks by obscuring objects and scenes. However, existing deep learning-based snow removal methods are designed for single images only. In this paper, we target a more complex task -- video snow removal, which aims to re…

Cited by 22PDFcodeScholar
2022

Geometry-Contrastive Transformer for Generalized 3D Pose Transfer

AAAI 2022technical

We present a customized 3D mesh Transformer model for the pose transfer task. As the 3D pose transfer essentially is a deformation procedure dependent on the given meshes, the intuition of this work is to perceive the geometric inconsistency between the given meshes with the powerful self-attention…

2022

KEMP: Keyframe-Based Hierarchical End-to-End Deep Model for Long- Term Trajectory Prediction

ICRA 2022poster

Predicting future trajectories of road agents is a critical task for autonomous driving. Recent goal-based trajectory prediction methods, such as DenseTNT and PECNet [1], [2], have shown good performance on prediction tasks on public datasets. However, they usually require complicated goal-selection…

Cited by 20SourceScholar
2021

Intrinsic-Extrinsic Preserved GANs for Unsupervised 3D Pose Transfer

ICCV 2021poster

With the strength of deep generative models, 3D pose transfer regains intensive research interests in recent years. Existing methods mainly rely on a variety of constraints to achieve the pose transfer over 3D meshes, e.g., the need for manually encoding for shape and pose disentanglement. In this p…

Cited by 33PDFcodeScholar
2021

iMiGUE: An Identity-Free Video Dataset for Micro-Gesture Understanding and Emotion Analysis

CVPR 2021poster

We introduce a new dataset for the emotional artificial intelligence research: identity-free video dataset for micro-gesture understanding and emotion analysis (iMiGUE). Different from existing public datasets, iMiGUE focuses on nonverbal body gestures without using any identity information, while t…

Cited by 117PDFcodeScholar
2020

PIPAL: a Large-Scale Image Quality Assessment Dataset for Perceptual Image Restoration

ECCV 2020poster

Image quality assessment (IQA) is the key factor for the fast development of image restoration (IR) algorithms. The most recent IR methods based on Generative Adversarial Networks (GANs) have achieved significant improvement in visual performance, but also presented great challenges for quantitative…

Cited by 232SourcePDFScholar