← Search

Chao Liang

28 accepted papers

2026

BFMPF-Net: Bidirectional Frequency-Domain Modulation Progressive Fusion Network for Road Crack Segmentation

ICRA 2026poster

Recently, deep learning–based methods for road crack segmentation have achieved promising performance, particularly in robotic vision applications such as automated inspection and maintenance. However, most frequency-domain methods employ a decoupled processing strategy, overlooking the dynamic modu…

Cited by 0Scholar
2026

CitySeeker: How Do VLMs Explore Embodied Urban Navigation with Implicit Human Needs?

ICLR 2026poster

Vision-Language Models (VLMs) have made significant progress in explicit instruction-based navigation; however, their ability to interpret implicit human needs (e.g., ''I am thirsty'') in dynamic urban environments remains underexplored. This paper introduces CitySeeker, a novel benchmark designed t…

Cited by 0SourcecodeScholar
2026

Instilling an Active Mind in Avatars via Cognitive Simulation

ICLR 2026oral

Current video avatar models can generate fluid animations but struggle to capture a character's authentic essence, primarily synchronizing motion with low-level audio cues instead of understanding higher-level semantics like emotion or intent. To bridge this gap, we propose a novel framework for gen…

Cited by 0SourcecodeScholar
2026

InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions

ICLR 2026poster

End-to-end human animation with rich multi-modal conditions, e.g., text, image and audio has achieved remarkable advancements in recent years. However, most existing methods could only animate a single subject and inject conditions in a global manner, ignoring scenarios that multiple concepts could…

Cited by 0SourceScholar
2026

Leveraging Failed Samples: A Few-Shot and Training-Free Framework for Generalized Deepfake Detection

AAAI 2026technical

Recent deepfake detection studies often treat unseen sample detection as a ``zero-shot" task, training on images generated by known models but generalizing to unknown ones. A key real-world challenge arises when a model performs poorly on unknown samples, yet these samples remain available for analy

Cited by 0SourcePDFScholar
2026

Tutor-Student Reinforcement Learning: A Dynamic Curriculum for Robust Deepfake Detection

CVPR 2026

Standard supervised training for deepfake detection treats all samples with uniform importance, which can be suboptimal for learning robust and generalizable features. In this work, we propose a novel Tutor-Student Reinforcement Learning (TSRL) framework to dynamically optimize the training curricul

Cited by 0SourcecodeScholar
2026

VLM-Guided Group Preference Alignment for Diffusion-based Human Mesh Recovery

CVPR 2026

Human mesh recovery (HMR) from a single RGB image is inherently ambiguous, as multiple 3D poses can correspond to the same 2D observation. Recent diffusion-based methods tackle this by generating various hypotheses, but often sacrifice accuracy. They yield predictions that are either physically impl

Cited by 0SourceScholar
2025

CyberHost: A One-stage Diffusion Framework for Audio-driven Talking Body Generation

ICLR 2025oral

Diffusion-based video generation technology has advanced significantly, catalyzing a proliferation of research in human animation. While breakthroughs have been made in driving human animation through various modalities for portraits, most of current solutions for human body animation still focus on…

Cited by 0SourcePDFScholar
2025

FADA: Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG Distillation

CVPR 2025poster

Diffusion-based audio-driven talking avatar methods have recently gained attention for their high-fidelity, vivid, and expressive results. However, their slow inference speed limits practical applications. Despite the development of various distillation techniques for diffusion models, we found that…

2025

FriendsQA: A New Large-Scale Deep Video Understanding Dataset with Fine-grained Topic Categorization for Story Videos

AAAI 2025technical

Video question answering (VideoQA) aims to answer natural language questions according to the given videos. Although existing models perform well in the factoid VideoQA task, they still face challenges in deep video understanding (DVU) task, which focuses on story videos. Compared to factoid videos,…

2025

Link-based Contrastive Learning for One-Shot Unsupervised Domain Adaptation

CVPR 2025poster

Unsupervised domain adaptation (UDA) aims to learn discriminative features from a labeled source domain by supervised learning and to transfer the knowledge to an unlabeled target domain via distribution alignment. However, in some real-world scenarios, e.g., public safety or access control, it's di…

Cited by 0SourcePDFScholar
2025

Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency

ICLR 2025oral

With the introduction of video diffusion model, audio-conditioned human video generation has recently achieved significant breakthroughs in both the naturalness of motion and the synthesis of portrait details. Due to the limited control of audio signals in driving human motion, existing methods ofte…

2025

MobilePortrait: Real-Time One-Shot Neural Head Avatars on Mobile Devices

CVPR 2025poster

Existing neural head avatars methods have achieved significant progress in the image quality and motion range of portrait animation. However, these methods prioritize effectiveness over computational overhead. This paper presents MobilePortrait, a lightweight one-shot neural head avatars method that…

Cited by 8SourcePDFScholar
2025

OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

ICCV 2025poster

End-to-end human animation, such as audio-driven talking human generation, has undergone notable advancements in the recent few years. However, existing methods still struggle to scale up as large general video generation models, limiting their potential in real applications. In this paper, we propo…

Cited by 0SourcePDFScholar
2025

Rethinking the Adversarial Robustness of Multi-Exit Neural Networks in an Attack-Defense Game

CVPR 2025poster

Multi-exit neural networks represent a promising approach to enhancing model inference efficiency, yet like common neural networks, they suffer from significantly reduced robustness against adversarial attacks. While some defense methods have been raised to strengthen the adversarial robustness of m…

Cited by 0SourcePDFScholar
2025

Visual Relation Diffusion for Human-Object Interaction Detection

ICCV 2025poster

Human-object interaction (HOI) detection relies on fine-grained visual understanding to distinguish complex relationships between humans and objects. While recent generative diffusion models have demonstrated remarkable capability in learning detailed visual concepts through pixel-level generation,…

Cited by 0SourcePDFScholar
2024

CapHuman: Capture Your Moments in Parallel Universes

CVPR 2024poster

We concentrate on a novel human-centric image synthesis task that is given only one reference facial photograph it is expected to generate specific individual images with diverse head positions poses facial expressions and illuminations in different contexts. To accomplish this goal we argue that ou…

2024

QI-IRA: Quantum-Inspired Interactive Ranking Aggregation for Person Re-identification

AAAI 2024technical

Ranking aggregation (RA), the process of aggregating multiple rankings derived from multiple search strategies, has been proved effective in person re-identification (re-ID) because of a single re-ID method can not always achieve consistent superiority for different scenarios. Existing RA research m…

2023

TEPrompt: Task Enlightenment Prompt Learning for Implicit Discourse Relation Recognition

ACL 2023findings

Implicit Discourse Relation Recognition (IDRR) aims at classifying the relation sense between two arguments without an explicit connective. Recently, the ConnPrompt (Xiang et al., 2022) has leveraged the powerful prompt learning for IDRR based on the fusion of multi-prompt decisions from three diffe…

2022

One More Check: Making “Fake Background” Be Tracked Again

AAAI 2022technical

The one-shot multi-object tracking, which integrates object detection and ID embedding extraction into a unified network, has achieved groundbreaking results in recent years. However, current one-shot trackers solely rely on single-frame detections to predict candidate bounding boxes, which may be u…

2020

Crowdsourcing-Based Ranking Aggregation for Person Re-Identification

ICASSP 2020accepted

Person re-identification (re-ID) is widely applied in surveillance and criminal detection applications. The existing research focus on devising the stand-alone re-ID methods, ignoring their practical application in the multi-person collaboration scenario. To improve the search efficiency, a group of…

Cited by 0SourceScholar
2020

When Pedestrian Detection Meets Nighttime Surveillance: A New Benchmark

IJCAI 2020poster

Pedestrian detection at nighttime is a crucial and frontier problem in surveillance, but has not been well explored by the computer vision and artificial intelligence communities. Most of existing methods detect pedestrians under favorable lighting conditions (e.g. daytime) and achieve promising per…

2019

Cross-view Identical Part Area Alignment for Person Re-identification

ICASSP 2019accepted

Person re-identification aims to associate images captured by non-overlapping cameras. It is a challenging task because images are often in different conditions such as background clutter, illumination variation, viewpoint changes and different camera settings. Viewpoint changes and pose variations…

Cited by 0SourceScholar
2019

Rain Streak Removal via Multi-scale Mixture Exponential Power Model

ICASSP 2019accepted

Rain streaks severely hamper the visible performance of the outdoor surveillance videos, which becomes an attractive issue in recent computer vision research. Existing methods usually encode rain streaks into Gaussian Mixture Model (GM-M). However, the limited number of Gaussian components in the GM…

Cited by 0SourceScholar
2017

Transferring clothing parsing from fashion dataset to surveillance

ICASSP 2017accepted

In this paper we address the problem of automatic clothing parsing in surveillance video with the information from user-generated tags such as “jeans” and “T-shirt”. Although clothing parsing has achieved great success in fashion clothing, it is quite challenging to parse clothing in practical surve…

Cited by 0SourceScholar