← Search

YunHong Wang

78 accepted papers

2026

DocOS: A Benchmark for Proactive Document-Guided Actions in GUI Agents

ICML 2026poster

While Graphical User Interface (GUI) agents have shown promising performance in automated device interaction, they primarily depend on static parametric knowledge from pre-training or instruction tuning. This reliance fundamentally limits their ability to handle long-tailed tasks that require explic…

Cited by 0SourceScholar
2026

LSP Framework: A Compensatory Model for Defeating Trigger Reverse Engineering via Label Smoothing Poisoning

ICASSP 2026oral

Deep neural networks are vulnerable to backdoor attacks. Among the existing backdoor defense methods, trigger reverse engineering based approaches, which reconstruct the backdoor triggers via optimizations, are the most versatile and effective ones compared to other types of methods. In this paper,…

Cited by 0SourcePDFScholar
2026

MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied Navigation

CVPR 2026

Embodied navigation is a fundamental capability for robotic agents operating. Real-world deployment requires open vocabulary generalization and low training overhead, motivating zero-shot methods rather than task-specific RL training. However, existing zero-shot methods that build explicit 3D scene

Cited by 0SourceScholar
2026

Memory-Efficient Transfer Learning with Fading Side Networks via Masked Dual Path Distillation

CVPR 2026

Memory-efficient transfer learning (METL) approaches have recently achieved promising performance in adapting pre-trained models to downstream tasks. They avoid applying gradient backpropagation in large backbones, thus significantly reducing the number of trainable parameters and high memory consum

Cited by 0SourcecodeScholar
2026

Parameter-Efficient Adaptation for MLLMs via Implicit Modality Decomposition

CVPR 2026

Parameter-efficient fine-tuning (PEFT) has become a compelling approach for adapting large language models (LLMs) into multimodal large language models (MLLMs), enabling them to handle diverse modalities with substantially lower memory and computational costs. However, most existing PEFT methods neg

Cited by 0SourcecodeScholar
2026

Reasoning-Driven Anomaly Detection and Localization with Image-Level Supervision

CVPR 2026

Multimodal large language models (MLLMs) have recently demonstrated remarkable reasoning and perceptual abilities for anomaly detection. However, most approaches remain confined to image-level anomaly detection and textual reasoning, while pixel-level localization still relies on external vision mod

Cited by 0SourcecodeScholar
2026

ResWorld: Temporal Residual World Model for End-to-End Autonomous Driving

ICLR 2026poster

The comprehensive understanding capabilities of world models for driving scenarios have significantly improved the planning accuracy of end-to-end autonomous driving frameworks. However, the redundant modeling of static regions and the lack of deep interaction with trajectories hinder world models f…

Cited by 0SourcecodeScholar
2026

Semantic-Aware Motion Encoding for Topology-Agnostic Character Animation

ICML 2026poster

Generalizing motion representation across diverse characters remains challenging due to significant topological variations in skeletal structures across datasets and species, which hinders the development of scalable generative models. To bridge this gap, we propose a Semantic-Aware Topology-Agnosti…

Cited by 0SourceScholar
2026

WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling

ICML 2026poster

This paper presents WorldPlay, a streaming video diffusion model that enables real-time, interactive world modeling with long-term geometric consistency, resolving the trade-off between speed and memory that limits current methods. WorldPlay draws power from three key innovations. 1) We use a Dual A…

Cited by 0SourceScholar
2025

APHQ-ViT: Post-Training Quantization with Average Perturbation Hessian Based Reconstruction for Vision Transformers

CVPR 2025poster

Vision Transformers (ViTs) have become one of the most commonly used backbones for vision tasks. Despite their remarkable performance, they often suffer significant accuracy drop when quantized for practical deployment, particularly by post-training quantization (PTQ) under ultra-low bits. Recently,…

2025

FIMA-Q: Post-Training Quantization for Vision Transformers by Fisher Information Matrix Approximation

CVPR 2025highlight

Post-training quantization (PTQ) has stood out as a cost-effective and promising model compression approach over recent years, as it eliminates the need for retraining on the entire dataset. Unfortunately, most existing PTQ methods for Vision Transformers (ViTs) exhibit a notable drop in accuracy, e…

2025

GODBench: A Benchmark for Multimodal Large Language Models in Video Comment Art

ACL 2025long

***Video Comment Art*** enhances user engagement by providing creative content that conveys humor, satire, or emotional resonance, requiring a nuanced and comprehensive grasp of cultural and contextual subtleties. Although Multimodal Large Language Models (MLLMs) and Chain-of-Thought (CoT) have demo…

2025

Generating Editable Head Avatars with 3D Gaussian GANs

ICASSP 2025accepted

Generating animatable and editable 3D head avatars is essential for various applications in computer vision and graphics. Traditional 3D-aware generative adversarial networks (GANs), often using implicit fields like Neural Radiance Fields (NeRF), achieve photo-realistic and view-consistent 3D head s…

Cited by 0SourceScholar
2025

GeoBEV: Learning Geometric BEV Representation for Multi-view 3D Object Detection

AAAI 2025technical

Bird's-Eye-View (BEV) representation has emerged as a mainstream paradigm for multi-view 3D object detection, demonstrating impressive perceptual capabilities. However, existing methods overlook the geometric quality of BEV representation, leaving it in a low-resolution state and failing to restore…

2025

KwaiChat: A Large-Scale Video-Driven Multilingual Mixed-Type Dialogue Corpus

NAACL 2025findings

Video-based dialogue systems have compelling application value, such as education assistants, thereby garnering growing interest. However, the current video-based dialogue systems are limited by their reliance on a single dialogue type, which hinders their versatility in practical applications acros…

2025

Multi-modal Deepfake Detection via Multi-task Audio-Visual Prompt Learning

AAAI 2025technical

With the malicious use and dissemination of multi-modal deepfake videos, researchers start to investigate multi-modal deepfake detection. Unfortunately, most of the existing methods tune all the parameters of the deep network with limited speech video datasets and are trained under coarse-grained co…

Cited by 0SourcePDFScholar
2025

OpenRSD: Towards Open-prompts for Object Detection in Remote Sensing Images

ICCV 2025poster

Remote sensing object detection has made significant progress, but most studies still focus on closed-set detection, limiting generalization across diverse datasets. Open-vocabulary object detection (OVD) provides a solution by leveraging multimodal associations between text prompts and visual featu…

2025

RETAIL: Towards Real-world Travel Planning for Large Language Models

EMNLP 2025

Although large language models have enhanced automated travel planning abilities, current systems remain misaligned with real-world scenarios. First, they assume users provide explicit queries, while in reality requirements are often implicit. Second, existing solutions ignore diverse environmental

Cited by 0SourcePDFScholar
2025

RepoDebug: Repository-Level Multi-Task and Multi-Language Debugging Evaluation of Large Language Models

EMNLP 2025

Large Language Models (LLMs) have exhibited significant proficiency in code debugging, especially in automatic program repair, which may substantially reduce the time consumption of developers and enhance their efficiency. Significant advancements in debugging datasets have been made to promote the

2025

SPMTrack: Spatio-Temporal Parameter-Efficient Fine-Tuning with Mixture of Experts for Scalable Visual Tracking

CVPR 2025poster

Most state-of-the-art trackers adopt one-stream paradigm, using a single Vision Transformer for joint feature extraction and relation modeling of template and search region images. However, relation modeling between different image patches exhibits significant variations. For instance, background re…

2025

SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding

CVPR 2025poster

With the rapid development of Multi-modal Large Language Models (MLLMs), an increasing number of benchmarks have been established to evaluate the video understanding capabilities of these models. However, these benchmarks focus on standalone videos and only assess "visual elements" like human action…

2025

SkeletonMix: A Mixup-Based Data Augmentation Framework for Skeleton-Based Action Recognition

ICASSP 2025accepted

Skeleton-based human action recognition has received widespread attention for its robustness to changes in the background and appearance of actors compared to the RGB modality. Data augmentation is widely used to explicitly regularize the model to prevent overfitting, especially when the number of l…

Cited by 0SourceScholar
2025

TCAQ-DM: Timestep-Channel Adaptive Quantization for Diffusion Models

AAAI 2025technical

Diffusion models have achieved remarkable success in the image and video generation tasks. Nevertheless, they often require a large amount of memory and time overhead during inference, due to the complex network architecture and considerable number of timesteps for iterative diffusion. Recently, the…

Cited by 1SourcePDFScholar
2025

ToolSpectrum: Towards Personalized Tool Utilization for Large Language Models

ACL 2025finding

While integrating external tools into large language models (LLMs) enhances their ability to access real-time information and domain-specific services, existing approaches focus narrowly on functional tool selection following user instructions while overlooking the critical role of context-aware per…

Cited by 0SourcePDFScholar
2025

TransBench: Breaking Barriers for Transferable Graphical User Interface Agents in Dynamic Digital Environments

ACL 2025finding

Graphical User Interface (GUI) agents, which autonomously operate on digital interfaces through natural language instructions, hold transformative potential for accessibility, automation, and user experience. A critical aspect of their functionality is grounding — the ability to map linguistic inten…

2025

Unified Knowledge Maintenance Pruning and Progressive Recovery with Weight Recalling for Large Vision-Language Models

AAAI 2025technical

Large Vision-Language Model (LVLM), leveraging Large Language Model (LLM) as the cognitive core, has recently become one of the most representative multimodal model paradigms. However, with the expansion of unimodal branches, \emph{i.e.} visual encoder and LLM, the storage and computational burdens…

2025

Weak2Wise: An Automated, Lightweight Framework for Weak-LLM-Friendly Reasoning Synthesis

EMNLP 2025

Recent advances in large language model (LLM) fine‐tuning have shown that training data augmented with high-quality reasoning traces can remarkably improve downstream performance. However, existing approaches usually rely on expensive manual annotations or auxiliary models, and fail to address the u

2024

4Diffusion: Multi-view Video Diffusion Model for 4D Generation

NeurIPS 2024poster

Current 4D generation methods have achieved noteworthy efficacy with the aid of advanced diffusion generative models. However, these methods lack multi-view spatial-temporal modeling and encounter challenges in integrating diverse prior knowledge from multiple diffusion models, resulting in inconsis…

Cited by 27SourcePDFScholar
2024

AGS: Affordable and Generalizable Substitute Training for Transferable Adversarial Attack

AAAI 2024technical

In practical black-box attack scenarios, most of the existing transfer-based attacks employ pretrained models (e.g. ResNet50) as the substitute models. Unfortunately, these substitute models are not always appropriate for transfer-based attacks. Firstly, these models are usually trained on a largesc…

2024

ActiveDC: Distribution Calibration for Active Finetuning

CVPR 2024poster

The pretraining-finetuning paradigm has gained popularity in various computer vision tasks. In this paradigm the emergence of active finetuning arises due to the abundance of large-scale data and costly annotation requirements. Active finetuning involves selecting a subset of data from an unlabeled…

Cited by 2SourcePDFScholar
2024

AdaLog: Post-Training Quantization for Vision Transformers with Adaptive Logarithm Quantizer

ECCV 2024poster

"Vision Transformer (ViT) has become one of the most prevailing fundamental backbone networks in the computer vision community. Despite the high accuracy, deploying it in real applications raises critical challenges including the high computational cost and inference latency. Recently, the post-trai…

2024

DSD-DA: Distillation-based Source Debiasing for Domain Adaptive Object Detection

ICML 2024poster

Though feature-alignment based Domain Adaptive Object Detection (DAOD) methods have achieved remarkable progress, they ignore the source bias issue, i.e., the detector tends to acquire more source-specific knowledge, impeding its generalization capabilities in the target domain. Furthermore, these m…

Cited by 2SourcePDFScholar
2024

FSD-BEV: Foreground Self-Distillation for Multi-view 3D Object Detection

ECCV 2024poster

"Although multi-view 3D object detection based on the Bird’s-Eye-View (BEV) paradigm has garnered widespread attention as an economical and deployment-friendly perception solution for autonomous driving, there is still a performance gap compared to LiDAR-based methods. In recent years, several cross…

2024

Leveraging Predicate and Triplet Learning for Scene Graph Generation

CVPR 2024poster

Scene Graph Generation (SGG) aims to identify entities and predict the relationship triplets <subject predicate object> in visual scenes. Given the prevalence of large visual variations of subject-object pairs even in the same predicate it can be quite challenging to model and refine predicate repre…

2024

MutDet: Mutually Optimizing Pre-training for Remote Sensing Object Detection

ECCV 2024poster

"Detection pre-training methods for the DETR series detector have been extensively studied in natural scenes, e.g., DETReg. However, the detection pre-training remains unexplored in remote sensing scenes. In existing pre-training methods, alignment between object embeddings extracted from a pre-trai…

2024

Read, Spell and Repeat: Scene Text Recognition with Vision-Language Circular Refinement

ICASSP 2024accepted

Scene Text Recognition (STR) has long been considered an important yet challenging task in the field of computer vision. Recent works have demonstrated that utilizing language information is effective for the visually difficult images, like ones with occultation or blurring. However, the use of lang…

Cited by 0SourceScholar
2024

Rotation Has Two Sides: Evaluating Data Augmentation for Deep One-class Classification

ICLR 2024spotlight

One-class classification (OCC) involves predicting whether a new data is normal or anomalous based solely on the data from a single class during training. Various attempts have been made to learn suitable representations for OCC within a self-supervised framework. Notably, discriminative methods tha…

Cited by 3SourcePDFScholar
2024

Transforming Vision Transformer: Towards Efficient Multi-Task Asynchronous Learner

NeurIPS 2024poster

Multi-Task Learning (MTL) for Vision Transformer aims at enhancing the model capability by tackling multiple tasks simultaneously. Most recent works have predominantly focused on designing Mixture-of-Experts (MoE) structures and integrating Low-Rank Adaptation (LoRA) to efficiently perform multi-tas…

Cited by 1SourcePDFScholar
2023

BISVP: Building Footprint Extraction Via Bidirectional Serialized Vertex Prediction

ICASSP 2023accepted

Extracting building footprints from remote sensing images has been attracting extensive attention recently. Dominant approaches address this challenging problem by generating vectorized building masks with cumbersome refinement stages, which limits the application of such methods. In this paper, we…

Cited by 0SourceScholar
2023

Deepfake Video Detection via Facial Action Dependencies Estimation

AAAI 2023technical

Deepfake video detection has drawn significant attention from researchers due to the security issues induced by deepfake videos. Unfortunately, most of the existing deepfake detection approaches have not competently modeled the natural structures and movements of human faces. In this paper, we formu…

Cited by 16SourcePDFScholar
2023

Denoising Diffusion Autoencoders are Unified Self-supervised Learners

ICCV 2023oral

Inspired by recent advances in diffusion models, which are reminiscent of denoising autoencoders, we investigate whether they can acquire discriminative representations for classification via generative pre-training. This paper shows that the networks in diffusion models, namely denoising diffusion…

Cited by 72PDFcodeScholar
2023

Global-Local Characteristic Excited Cross-Modal Attacks from Images to Videos

AAAI 2023technical

The transferability of adversarial examples is the key property in practical black-box scenarios. Currently, numerous methods improve the transferability across different models trained on the same modality of data. The investigation of generating video adversarial examples with imagebased substitut…

2023

Learning Discriminative Representations for Skeleton Based Action Recognition

CVPR 2023poster

Human action recognition aims at classifying the category of human action from a segment of a video. Recently, people have dived into designing GCN-based models to extract features from skeletons for performing this task, because skeleton representations are much more efficient and robust than other…

2023

SA-BEV: Generating Semantic-Aware Bird's-Eye-View Feature for Multi-view 3D Object Detection

ICCV 2023poster

Recently, the pure camera-based Bird's-Eye-View (BEV) perception provides a feasible solution for economical autonomous driving. However, the existing BEV-based multi-view 3D detectors generally transform all image features into BEV features, without considering the problem that the large proportion…

Cited by 34PDFcodeScholar
2023

Unilaterally Aggregated Contrastive Learning with Hierarchical Augmentation for Anomaly Detection

ICCV 2023poster

Anomaly detection (AD), aiming to find samples that deviate from the training distribution, is essential in safety-critical applications. Though recent self-supervised learning based attempts achieve promising results by creating virtual outliers, their training objectives are less faithful to AD wh…

Cited by 10PDFScholar
2022

Lagrange Motion Analysis and View Embeddings for Improved Gait Recognition

CVPR 2022poster

Gait is considered the walking pattern of human body, which includes both shape and motion cues. However, the main-stream appearance-based methods for gait recognition rely on the shape of silhouette. It is unclear whether motion can be explicitly represented in the gait sequence modeling. In this p…

Cited by 78PDFcodeScholar
2022

PACE: Predictive and Contrastive Embedding for Unsupervised Action Segmentation

IJCAI 2022poster

Action segmentation, inferring temporal positions of human actions in an untrimmed video, is an important prerequisite for various video understanding tasks. Recently, unsupervised action segmentation (UAS) has emerged as a more challenging task due to the unavailability of frame-level annotations.…

Cited by 0SourcePDFScholar
2022

SparseTT: Visual Tracking with Sparse Transformers

IJCAI 2022poster

Transformers have been successfully applied to the visual tracking task and significantly promote tracking performance. The self-attention mechanism designed to model long-range dependencies is the key to the success of Transformers. However, self-attention lacks focusing on the most relevant inform…

2022

Video Anomaly Detection by Solving Decoupled Spatio-Temporal Jigsaw Puzzles

ECCV 2022poster

"Video Anomaly Detection (VAD) is an important topic in computer vision. Motivated by the recent advances in self-supervised learning, this paper addresses VAD by solving an intuitive yet challenging pretext task, i.e., spatio-temporal jigsaw puzzles, which is cast as a multi-label fine-grained clas…

2021

MIEHDR CNN: Main Image Enhancement based Ghost-Free High Dynamic Range Imaging using Dual-Lens Systems

AAAI 2021technical

We study the High Dynamic Range (HDR) imaging problem using two Low Dynamic Range (LDR) images that are shot from dual-lens systems in a single shot time with different exposures. In most of the related HDR imaging methods, the problem is usually solved by Multiple Images Merging, i.e. the final HDR…

Cited by 8SourcePDFScholar
2021

PC-RGNN: Point Cloud Completion and Graph Neural Network for 3D Object Detection

AAAI 2021technical

LiDAR-based 3D object detection is an important task for autonomous driving and current approaches suffer from sparse and partial point clouds caused by distant and occluded objects. In this paper, we propose a novel two-stage framework, namely PC-RGNN, which deals with these challenges by two speci…

Cited by 104SourcePDFScholar
2021

Path-BN: Towards effective batch normalization in the Path Space for ReLU networks

UAI 2021poster

Neural networks with ReLU activation functions (abbrev. ReLU Networks), have demonstrated their success in many applications. Recently, researchers noticed that ReLU networks are positively scale-invariant (PSI) while the weights are not. This mismatch may lead to undesirable behaviors in the optimi…

Cited by 1SourcePDFScholar
2021

STMTrack: Template-Free Visual Tracking With Space-Time Memory Networks

CVPR 2021poster

Boosting performance of the offline trained siamese trackers is getting harder nowadays since the fixed information of the template cropped from the first frame has been almost thoroughly mined, but they are poorly capable of resisting target appearance changes. Existing trackers with template updat…

Cited by 364PDFcodeScholar
2020

Cross-domain Object Detection through Coarse-to-Fine Feature Adaptation

CVPR 2020poster

Recent years have witnessed great progress in deep learning based object detection. However, due to the domain shift problem, applying off-the-shelf detectors to an unseen domain leads to significant performance drop. To address such an issue, this paper proposes a novel coarse-to-fine feature adapt…

Cited by 270PDFScholar
2020

I4R: Promoting Deep Reinforcement Learning by the Indicator for Expressive Representations

IJCAI 2020poster

Learning expressive representations is always crucial for well-performed policies in deep reinforcement learning (DRL). Different from supervised learning, in DRL, accurate targets are not always available, and some inputs with different actions only have tiny differences, which stimulates the deman…

2020

Multi-Scale Positive Sample Refinement for Few-Shot Object Detection

ECCV 2020poster

Few-shot object detection (FSOD) helps detectors adapt to unseen classes with few training instances, and is useful when manual annotation is time-consuming or data acquisition is limited. Unlike previous attempts that exploit few-shot classification techniques to facilitate FSOD, this work highligh…

2019

Attentive Relational Networks for Mapping Images to Scene Graphs

CVPR 2019poster

Scene graph generation refers to the task of automatically mapping an image into a semantic structural graph, which requires correctly labeling each extracted object and their interaction relationships. Despite the recent success in object detection using deep learning techniques, inferring complex…

Cited by 199PDFScholar
2019

KE-GAN: Knowledge Embedded Generative Adversarial Networks for Semi-Supervised Scene Parsing

CVPR 2019poster

In recent years, scene parsing has captured increasing attention in computer vision. Previous works have demonstrated promising performance in this task. However, they mainly utilize holistic features, whilst neglecting the rich semantic knowledge and inter-object relationships in the scene. In addi…

Cited by 59PDFScholar
2019

Led3D: A Lightweight and Efficient Deep Approach to Recognizing Low-Quality 3D Faces

CVPR 2019poster

Due to the intrinsic invariance to pose and illumination changes, 3D Face Recognition (FR) has a promising potential in the real world. 3D FR using high-quality faces, which are of high resolutions and with smooth surfaces, have been widely studied. However, research on that with low-quality input i…

Cited by 70PDFScholar
2018

Hierarchical Attention and Context Modeling for Group Activity Recognition

ICASSP 2018accepted

Group activity recognition in videos is a challenging task, with two major issues, i.e. attending to those persons and their body parts that contribute significantly to the activity, and modeling contextual person structures in the group. Most previous approaches fail to provide a practical solution…

Cited by 0SourceScholar
2018

Hough Transform Guided Deep Feature Extraction for Dense Building Detection in Remote Sensing Images

ICASSP 2018accepted

Detecting dense buildings without elevation information is an important and challenging task in remote sensing applications. In this paper, we present a novel cascaded deep neural network architecture, incorporating multi -stage region proposal detection and Hough transform to obtain better mid-leve…

Cited by 0SourceScholar
2018

Learning Face Age Progression: A Pyramid Architecture of GANs

CVPR 2018poster

The two underlying requirements of face age progression, i.e. aging accuracy and identity permanence, are not well studied in the literature. In this paper, we present a novel generative adversarial network based approach. It separately models the constraints for the intrinsic subject-specific chara…

Cited by 224SourcePDFScholar
2018

stagNet: An Attentive Semantic RNN for Group Activity Recognition

ECCV 2018poster

Group activity recognition plays a fundamental role in a variety of applications, e.g. sports video analysis and intelligent surveillance. How to model the spatio-temporal contextual information in a scene still remains a crucial yet challenging issue. We propose a novel attentive semantic recurrent…

Cited by 180SourcePDFScholar
2017

Binary Coding for Partial Action Analysis With Limited Observation Ratios

CVPR 2017poster

Traditional action recognition methods aim to recognize actions with complete observations/executions. However, it is often difficult to capture fully executed actions due to occlusions, interruptions, etc. Meanwhile, action prediction/recognition in advance based on partial observations is essentia…

Cited by 34PDFScholar
2017

Fast Person Re-Identification via Cross-Camera Semantic Binary Transformation

CVPR 2017poster

Numerous methods have been proposed for person re-identification, most of which however neglect the matching efficiency. Recently, several hashing based approaches have been developed to make re-identification more scalable for large-scale gallery sets. Despite their efficiency, these works ignore c…

Cited by 94PDFScholar
2017

Zero-Shot Action Recognition With Error-Correcting Output Codes

CVPR 2017poster

Recently, zero-shot action recognition (ZSAR) has emerged with the explosive growth of action categories. In this paper, we explore ZSAR from a novel perspective by adopting the Error-Correcting Output Codes (dubbed ZSECOC). Our ZSECOC equips the conventional ECOC with the additional capability of Z…

Cited by 186PDFScholar