← Search

Wen-Huang Cheng

24 accepted papers

2026

MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied Agents

CVPR 2026

Theory of Mind (ToM) refers to the ability to infer others' mental states, such as beliefs, desires, and intentions. Current vision-language embodied agents lack ToM-based decision-making, and existing benchmarks focus solely on human mental states while ignoring the agent's own perspective, hinderi

Cited by 0SourcecodeScholar
2026

TriDF: Evaluating Perception, Detection, and Hallucination for Interpretable DeepFake Detection

CVPR 2026

Advances in generative modeling have made it increasingly easy to fabricate realistic portrayals of individuals, creating serious risks for security, communication, and public trust. Detecting such person-driven manipulations requires systems that not only distinguish altered content from authentic

Cited by 0SourceScholar
2025

From Prompt to Progression: Taming Video Diffusion Models for Seamless Attribute Transition

ICCV 2025poster

Existing models often struggle with complex temporal changes, particularly when generating videos with gradual attribute transitions.The most common prompt interpolation approach for motion transitions often fails to handle gradual attribute transitions, where inconsistencies tend to become more pro…

Cited by 0SourcePDFScholar
2025

Future Sight and Tough Fights: Revolutionizing Sequential Recommendation with FENRec

AAAI 2025technical

Sequential recommendation (SR) systems predict user preferences by analyzing time-ordered interaction sequences. A common challenge for SR is data sparsity, as users typically interact with only a limited number of items. While contrastive learning has been employed in previous approaches to addres…

2025

Memory-Augmented Re-Completion for 3D Semantic Scene Completion

AAAI 2025technical

Semantic Scene Completion (SSC) aims to reconstruct a 3D voxel representation occupied by semantic classes based on ordinary inputs such as 2D RGB images, depth maps, or point clouds. Given the cost-effective and promising applications in autonomous driving, camera-based SSC has attracted considerab…

2025

MonoTAKD: Teaching Assistant Knowledge Distillation for Monocular 3D Object Detection

CVPR 2025poster

Monocular 3D object detection (Mono3D) holds noteworthy promise for autonomous driving applications owing to the cost-effectiveness and rich visual context of monocular camera sensors. However, depth ambiguity poses a significant challenge, as it requires extracting precise 3D scene geometry from a…

2025

Personalized Lip Reading: Adapting to Your Unique Lip Movements with Vision and Language

AAAI 2025technical

Lip reading aims to predict spoken language by analyzing lip movements. Despite advancements in lip reading technologies, performance degrades when models are applied to unseen speakers due to their sensitivity to variations in visual information such as lip appearances. To address this challenge, s…

2025

Perspective-Aware Teaching: Adapting Knowledge for Heterogeneous Distillation

ICCV 2025poster

Knowledge distillation (KD) involves transferring knowledge from a pre-trained heavy teacher model to a lighter student model, thereby reducing the inference cost while maintaining comparable effectiveness. Prior KD techniques typically assume homogeneity between the teacher and student models. Howe…

2025

Training-Free Industrial Defect Generation with Diffusion Models

ICCV 2025poster

Anomaly generation has become essential in addressing the scarcity of defective samples in industrial anomaly inspection. However, existing training-based methods fail to handle complex anomalies and multiple defects simultaneously, especially when only a single anomaly sample is available per defec…

2025

When Anchors Meet Cold Diffusion: A Multi-Stage Approach to Lane Detection

ICCV 2025poster

Accurate and stable lane detection is crucial for the reliability of autonomous driving systems. A core challenge lies in predicting lane positions in complex scenarios, such as curved roads or when markings are ambiguous or absent.Conventional approaches leverage deep learning techniques to extract…

2024

DQ-DETR: DETR with Dynamic Query for Tiny Object Detection

ECCV 2024poster

"Despite previous DETR-like methods having performed successfully in generic object detection, tiny object detection is still a challenging task for them since the positional information of object queries is not customized for detecting tiny objects, whose scale is extraordinarily smaller than gener…

2024

Distraction is All You Need: Memory-Efficient Image Immunization against Diffusion-Based Image Editing

CVPR 2024highlight

Recent text-to-image (T2I) diffusion models have revolutionized image editing by empowering users to control outcomes using natural language. However the ease of image manipulation has raised ethical concerns with the potential for malicious use in generating deceptive or harmful content. To address…

Cited by 5SourcePDFScholar
2024

EmoVIT: Revolutionizing Emotion Insights with Visual Instruction Tuning

CVPR 2024poster

Visual Instruction Tuning represents a novel learning paradigm involving the fine-tuning of pre-trained language models using task-specific instructions. This paradigm shows promising zero-shot results in various natural language processing tasks but is still unexplored in vision emotion understandi…

2024

Representation and Boundary Enhancement for Action Segmentation Using Transformer

ICASSP 2024accepted

In the task of action segmentation, the goal is to partition a lengthy, untrimmed video into a series of action segments. Recently, Transformer-based methods have outperformed the previous temporal convolutional networks (TCNs) in terms of overall performance. However, both TCNs and Transformers enc…

Cited by 0SourceScholar
2024

TrajPrompt: Aligning Color Trajectory with Vision-Language Representations

ECCV 2024poster

"Cross-modal learning shows promising potential to overcome the limitations of single-modality tasks. However, without proper design for representation alignment between different data sources, the external modality cannot fully exhibit its value. For example, recent trajectory prediction approaches…

2023

Most Important Person-Guided Dual-Branch Cross-Patch Attention for Group Affect Recognition

ICCV 2023poster

Group affect refers to the subjective emotion that is evoked by an external stimulus in a group, which is an important factor that shapes group behavior and outcomes. Recognizing group affect involves identifying important individuals and salient objects among a crowd that can evoke emotions. Howeve…

Cited by 10PDFScholar
2023

Size Does Matter: Size-aware Virtual Try-on via Clothing-oriented Transformation Try-on Network

ICCV 2023poster

Virtual try-on tasks aim at synthesizing realistic try-on results by trying target clothes on humans. Most previous works relied on the Thin Plate Spline or appearance flows to warp clothes to fit human body shapes. However, both approaches cannot handle complex warping, leading to over distortion o…

Cited by 28PDFcodeScholar
2023

Zero-Shot Face-Based Voice Conversion: Bottleneck-Free Speech Disentanglement in the Real-World Scenario

AAAI 2023technical

Often a face has a voice. Appearance sometimes has a strong relationship with one's voice. In this work, we study how a face can be converted to a voice, which is a face-based voice conversion. Since there is no clean dataset that contains face and speech, voice conversion faces difficult learning a…

Cited by 5SourcePDFScholar
2022

Social-SSL: Self-Supervised Cross-Sequence Representation Learning Based on Transformers for Multi-agent Trajectory Prediction

ECCV 2022poster

"Earlier trajectory prediction approaches focus on ways of capturing sequential structures among pedestrians by using recurrent networks, which is known to have some limitations in capturing long sequence structures. To address this limitation, some recent works proposed Transformer-based architectu…

2021

FashionMirror: Co-Attention Feature-Remapping Virtual Try-On With Sequential Template Poses

ICCV 2021poster

Virtual try-on tasks have drawn increased attention. Prior arts focus on tackling this task via warping clothes and fusing the information at the pixel level with the help of semantic segmentation. However, conducting semantic segmentation is time-consuming and easily causes error accumulation over…

Cited by 30PDFcodeScholar
2021

Naturalistic Physical Adversarial Patch for Object Detectors

ICCV 2021poster

Most prior works on physical adversarial attacks mainly focus on the attack performance but seldom enforce any restrictions over the appearance of the generated adversarial patches. This leads to conspicuous and attention-grabbing patterns for the generated patches which can be easily identified by…

Cited by 192PDFcodeScholar
2021

Re-Attention Is All You Need: Memory-Efficient Scene Text Detection via Re-Attention on Uncertain Regions

IROS 2021poster

Scene text detection plays an important role on vision-based robot navigation to many potential landmarks such as nameplates, information signs, floor button in the elevators. Recently, scene text detection with segmentation-based methods has been receiving more and more attention. The segmentation…

Cited by 5SourceScholar
2019

BeautyGlow: On-Demand Makeup Transfer Framework With Reversible Generative Network

CVPR 2019poster

As makeup has been widely-adopted for beautification, finding suitable makeup by virtual makeup applications becomes popular. Therefore, a recent line of studies proposes to transfer the makeup from a given reference makeup image to the source non-makeup one. However, it is still challenging due to…

Cited by 125PDFScholar
2019

Fuzzy Personalized Scoring Model for Recommendation System

ICASSP 2019accepted

In this research, we aim to propose a data preprocessing framework particularly for financial sector to generate the rating data as input to the collaborative system. First, clustering technique is applied to cluster all users based on their demographic information which might be able to differentia…

Cited by 0SourceScholar