← Search

Yabiao Wang

61 accepted papers

2026

Asynchronous Matching with Dynamic Sampling for Multimodal Dataset Distillation

ICLR 2026poster

Multimodal Dataset Distillation (MDD) has emerged as a vital paradigm for enabling efficient training of vision-language models (VLMs) in the era of multimodal data proliferation. Unlike traditional dataset distillation methods that focus on single-modal tasks, MDD presents distinct challenges: (i)…

Cited by 0SourceScholar
2026

LLM-Oriented Token-Adaptive Knowledge Distillation

AAAI 2026technical

Knowledge Distillation (KD) is a key technique for compressing Large-scale Language Models (LLMs), but prevailing logit-based methods employ static strategies misaligned with the student’s dynamic learning process. By treating all tokens indiscriminately with a fixed temperature, these methods resul

Cited by 0SourcePDFScholar
2026

Open the Motion Door: Atomic Motion Decomposition and Recomposition for Open-Vocabulary Motion Generation

CVPR 2026

Text-to-motion generation is a fundamental task in computer vision, aiming to synthesize 3D human motion sequences from natural language descriptions. However, due to the limited scale and diversity of existing datasets, models trained to directly map raw text to motion often struggle to generalize

Cited by 0SourceScholar
2026

Reasoning to Edit: Hypothetical Instruction-Based Image Editing with Visual Reasoning

ICML 2026poster

Instruction-based image editing (IIE) has advanced rapidly with the success of diffusion models. However, existing efforts primarily focus on simple and explicit instructions to execute editing operations such as adding, deleting, moving, or swapping objects. They struggle to handle more complex imp…

Cited by 0SourceScholar
2026

SwiftVideo: A Unified Framework for Few-Step Video Generation Through Trajectory-Distribution Alignment

AAAI 2026technical

Diffusion-based or flow-based models have achieved significant progress in video synthesis but require multiple iterative sampling steps, which incurs substantial computational overhead. While many distillation methods that are solely based on trajectory-preserving or distribution-matching have been

Cited by 0SourcePDFScholar
2026

TGPO: Efficient Policy Optimization through Sequence Anchor and Information Gating

ICML 2026poster

Reinforcement learning from verifiable rewards (RLVR) has become an important paradigm for enhancing the reasoning capabilities of large language models, while it also involves a persistent tradeoff between optimization stability and learning efficiency. Token-level importance weighting supports fin…

Cited by 0SourceScholar
2026

The devil is in the details: Enhancing Video Virtual Try-On via Keyframe-Driven Details Injection

CVPR 2026

Although diffusion transformer (DiT)-based video virtual try-on (VVT) has made significant progress in synthesizing realistic videos, existing methods still struggle to capture fine-grained garment dynamics and preserve background integrity across video frames. They also incur high computational cos

Cited by 0SourceScholar
2026

Transform Trained Transformer for Accelerating Native 4K Video Generation

ICML 2026poster

Native 4K (2176$\times$3840) video generation remains a critical challenge due to the quadratic computational explosion of full-attention as spatiotemporal resolution increases, making it difficult for models to strike a balance between efficiency and quality. This paper proposes a novel Transformer…

Cited by 0SourceScholar
2025

Dual-Interrelated Diffusion Model for Few-Shot Anomaly Image Generation

CVPR 2025poster

The performance of anomaly inspection in industrial manufacturing is constrained by the scarcity of anomaly data. To overcome this challenge, researchers have started employing anomaly generation approaches to augment the anomaly dataset. However, existing anomaly generation methods suffer from limi…

2025

Improving Autoregressive Visual Generation with Cluster-Oriented Token Prediction

CVPR 2025poster

Employing LLMs for visual generation has recently become a research focus. However, the existing methods primarily transfer the LLM architecture to visual generation but rarely investigate the fundamental differences between language and vision. This oversight may lead to suboptimal utilization of v…

2025

Mamba-YOLO-World: Marrying YOLO-World with Mamba for Open-Vocabulary Detection

ICASSP 2025accepted

Open-vocabulary detection (OVD) aims to detect objects beyond a predefined set of categories. As a pioneering model incorporating the YOLO series into OVD, YOLO-World is well-suited for scenarios prioritizing speed and efficiency. However, its performance is hindered by its neck feature fusion mecha…

Cited by 0SourceScholar
2025

MobileMamba: Lightweight Multi-Receptive Visual Mamba Network

CVPR 2025poster

Previous research on lightweight models has primarily focused on CNNs and Transformer-based designs. CNNs, with their local receptive fields, struggle to capture long-range dependencies, while Transformers, despite their global modeling capabilities, are limited by quadratic computational complexity…

2025

OSV: One Step is Enough for High-Quality Image to Video Generation

CVPR 2025poster

Video diffusion models have shown great potential in generating high-quality videos, making them an increasingly popular focus. However, their inherent iterative nature leads to substantial computational and time costs. Although techniques such as consistency distillation and adversarial training ha…

Cited by 10SourcePDFScholar
2025

PointRWKV: Efficient RWKV-Like Model for Hierarchical Point Cloud Learning

AAAI 2025technical

Transformers have revolutionized the point cloud learning task, but the quadratic complexity hinders its extension to long sequence and makes a burden on limited computational resources. The recent advent of RWKV, a fresh breed of deep sequence models, has shown immense potential for sequence modeli…

Cited by 14SourcePDFScholar
2025

SaRA: High-Efficient Diffusion Model Fine-tuning with Progressive Sparse Low-Rank Adaptation

ICLR 2025poster

The development of diffusion models has led to significant progress in image and video generation tasks, with pre-trained models like the Stable Diffusion series playing a crucial role. However, a key challenge remains in downstream task applications: how to effectively and efficiently adapt pre-tra…

2025

TIMotion: Temporal and Interactive Framework for Efficient Human-Human Motion Generation

CVPR 2025poster

Human-human motion generation is essential for understanding humans as social beings. Current methods fall into two main categories: single-person-based methods and separate modeling-based methods. To delve into this field, we abstract the overall generation process into a general framework MetaMoti…

Cited by 0SourcePDFScholar
2025

Towards Universal Dataset Distillation via Task-Driven Diffusion

CVPR 2025poster

Dataset distillation (DD) condenses key information from large-scale datasets into smaller synthetic datasets, reducing storage and computational costs for training networks. However, recent research has primarily focused on image classification tasks, with limited expansion to detection and segment…

Cited by 0SourcePDFScholar
2025

UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions

NeurIPS 2025poster

The quality of the video dataset (image quality, resolution, and fine-grained caption) greatly influences the performance of the video generation model. % The growing demand for video applications sets higher requirements for high-quality video generation models. % For example, the generation of m…

Cited by 0SourcecodeScholar
2025

UniCombine: Unified Multi-Conditional Combination with Diffusion Transformer

ICCV 2025poster

With the rapid development of diffusion models in image generation, the demand for more powerful and flexible controllable frameworks is increasing. Although existing methods can guide generation beyond text prompts, the challenge of effectively combining multiple conditional inputs while maintainin…

2025

WaveAR: Wavelet-Aware Continuous Autoregressive Diffusion for Accurate Human Motion Prediction

NeurIPS 2025poster

This work tackles a challenging problem: stochastic human motion prediction (SHMP), which aims to forecast diverse and physically plausible future pose sequences based on a short history of observed motion. While autoregressive sequence models have excelled in related generation tasks, their relianc…

Cited by 0SourceScholar
2024

A Diffusion-Based Framework for Multi-Class Anomaly Detection

AAAI 2024technical

Reconstruction-based approaches have achieved remarkable outcomes in anomaly detection. The exceptional image reconstruction capabilities of recently popular diffusion models have sparked research efforts to utilize them for enhanced reconstruction of anomalous images. Nonetheless, these methods mig…

2024

AnomalyDiffusion: Few-Shot Anomaly Image Generation with Diffusion Model

AAAI 2024technical

Anomaly inspection plays an important role in industrial manufacture. Existing anomaly inspection methods are limited in their performance due to insufficient anomaly data. Although anomaly generation methods have been proposed to augment the anomaly data, they either suffer from poor generation aut…

2024

Density Matters: Improved Core-Set for Active Domain Adaptive Segmentation

AAAI 2024technical

Active domain adaptation has emerged as a solution to balance the expensive annotation cost and the performance of trained models in semantic segmentation. However, existing works usually ignore the correlation between selected samples and its local context in feature space, which leads to inferior…

Cited by 2SourcePDFScholar
2024

Fetch and Forge: Efficient Dataset Condensation for Object Detection

NeurIPS 2024poster

Dataset condensation (DC) is an emerging technique capable of creating compact synthetic datasets from large originals while maintaining considerable performance. It is crucial for accelerating network training and reducing data storage requirements. However, current research on DC mainly focuses o…

Cited by 1SourcePDFScholar
2024

FreeMotion: A Unified Framework for Number-free Text-to-Motion Synthesis

ECCV 2024poster

"Text-to-motion synthesis is a crucial task in computer vision. Existing methods are limited in their universality, as they are tailored for single-person or two-person scenarios and can not be applied to generate motions for more individuals. To achieve the number-free motion synthesis, this paper…

Cited by 19SourcePDFScholar
2024

Learning Hybrid Negative Probability Model for Weakly-Supervised Whole Slide Image Recognition

ICASSP 2024accepted

Classifying an entire Whole Slide Image (WSI) in a single forward pass is challenging due to its vast resolution. Consequently, current effort on WSI classification resorts to multiple instance learning (MIL), using patch-wise instances to predict categories under image-wise supervision. However, re…

Cited by 0SourceScholar
2024

Rethinking Reverse Distillation for Multi-Modal Anomaly Detection

AAAI 2024technical

In recent years, there has been significant progress in employing color images for anomaly detection in industrial scenarios, but it is insufficient for identifying anomalies that are invisible in RGB images alone. As a supplement, introducing extra modalities such as depth and surface normal maps c…

Cited by 16SourcePDFScholar
2024

Search for Gravitational Wave Probes - A Self-Supervised Learning for Pulsars Based on Signal Contexts

ICASSP 2024accepted

The recent successful detection of gravitational waves (GWs) at nanohertz based on pulsar timing arrays has underscored the growing significance of searching for new pulsars, which serve as valuable probes for GWs. However, one of the challenges in this endeavor is the lack of labeled data, which ca…

Cited by 0SourceScholar
2024

Self-Supervised Likelihood Estimation with Energy Guidance for Anomaly Segmentation in Urban Scenes

AAAI 2024technical

Robust autonomous driving requires agents to accurately identify unexpected areas (anomalies) in urban scenes. To this end, some critical issues remain open: how to design advisable metric to measure anomalies, and how to properly generate training samples of anomaly data? Classical effort in anomal…

2024

TransAVS: End-to-End Audio-Visual Segmentation with Transformer

ICASSP 2024accepted

Audio-Visual Segmentation (AVS) is a challenging task, which aims to segment sounding objects in video frames by exploring audio signals. Generally AVS faces two key challenges: (1) Audio signals inherently exhibit a high degree of information density, as sounds produced by multiple objects are enta…

Cited by 0SourceScholar
2024

UniM-OV3D: Uni-Modality Open-Vocabulary 3D Scene Understanding with Fine-Grained Feature Representation

IJCAI 2024poster

3D open-vocabulary scene understanding aims to recognize arbitrary novel categories beyond the base label space. However, existing works not only fail to fully utilize all the available modal information in the 3D domain but also lack sufficient granularity in representing the features of each modal…

2023

Align, Perturb and Decouple: Toward Better Leverage of Difference Information for RSI Change Detection

IJCAI 2023poster

Change detection is a widely adopted technique in remote sense imagery (RSI) analysis in the discovery of long-term geomorphic evolution. To highlight the areas of semantic changes, previous effort mostly pays attention to learning representative feature descriptors of a single image, while the diff…

2023

Calibrated Teacher for Sparsely Annotated Object Detection

AAAI 2023technical

Fully supervised object detection requires training images in which all instances are annotated. This is actually impractical due to the high labor and time costs and the unavoidable missing annotations. As a result, the incomplete annotation in each image could provide misleading supervision and ha…

2023

Learning From Noisy Labels With Decoupled Meta Label Purifier

CVPR 2023poster

Training deep neural networks (DNN) with noisy labels is challenging since DNN can easily memorize inaccurate labels, leading to poor generalization ability. Recently, the meta-learning based label correction strategy is widely adopted to tackle this problem via identifying and correcting potential…

2023

Learning Global-aware Kernel for Image Harmonization

ICCV 2023poster

Image harmonization aims to solve the visual inconsistency problem in composited images by adaptively adjusting the foreground pixels with the background as references. Existing methods employ local color transformation or region matching between foreground and background, which neglects powerful pr…

Cited by 9PDFScholar
2023

MixTeacher: Mining Promising Labels With Mixed Scale Teacher for Semi-Supervised Object Detection

CVPR 2023poster

Scale variation across object instances is one of the key challenges in object detection. Although modern detection models have achieved remarkable progress in dealing with the scale variation, it still brings trouble in the semi-supervised case. Most existing semi-supervised object detection method…

2023

Multimodal Industrial Anomaly Detection via Hybrid Fusion

CVPR 2023poster

2D-based Industrial Anomaly Detection has been widely discussed, however, multimodal industrial anomaly detection based on 3D point clouds and RGB images still has many untouched fields. Existing multimodal industrial anomaly detection methods directly concatenate the multimodal features, which lead…

2023

Phasic Content Fusing Diffusion Model with Directional Distribution Consistency for Few-Shot Model Adaption

ICCV 2023poster

Training a generative model with limited number of samples is a challenging task. Current methods primarily rely on few-shot model adaption to train the network. However, in scenarios where data is extremely limited (less than 10), the generative network tends to overfit and suffers from content deg…

Cited by 14PDFcodeScholar
2023

RFENet: Towards Reciprocal Feature Evolution for Glass Segmentation

IJCAI 2023poster

Glass-like objects are widespread in daily life but remain intractable to be segmented for most existing methods. The transparent property makes it difficult to be distinguished from background, while the tiny separation boundary further impedes the acquisition of their exact contour. In this paper,…

2023

Remembering Normality: Memory-guided Knowledge Distillation for Unsupervised Anomaly Detection

ICCV 2023poster

Knowledge distillation (KD) has been widely explored in unsupervised anomaly detection (AD). The student is assumed to constantly produce representations of typical patterns within trained data, named "normality", and the representation discrepancy between the teacher and student model is identified…

Cited by 48PDFScholar
2023

Rethinking Mobile Block for Efficient Attention-based Models

ICCV 2023poster

This paper focuses on developing modern, efficient, lightweight models for dense predictions while trading off parameters, FLOPs, and performance. Inverted Residual Block (IRB) serves as the infrastructure for lightweight CNNs, but no counterpart has been recognized by attention-based studies. This…

Cited by 180PDFcodeScholar
2022

Designing One Unified Framework for High-Fidelity Face Reenactment and Swapping

ECCV 2022poster

"Face reenactment and swapping share a similar identity and attribute manipulating pattern, but most methods treat them separately, which is redundant and practical-unfriendly. In this paper, we propose an effective end-to-end unified framework to achieve both tasks. Unlike existing methods that dir…

2022

ISDNet: Integrating Shallow and Deep Networks for Efficient Ultra-High Resolution Segmentation

CVPR 2022poster

The huge burden of computation and memory are two obstacles in ultra-high resolution image segmentation. To tackle these issues, most of the previous works follow the global-local refinement pipeline, which pays more attention to the memory consumption but neglects the inference speed. In comparison…

Cited by 60PDFcodeScholar
2022

Iterative Few-shot Semantic Segmentation from Image Label Text

IJCAI 2022poster

Few-shot semantic segmentation aims to learn to segment unseen class objects with the guidance of only a few support images. Most previous methods rely on the pixel-level label of support images. In this paper, we focus on a more challenging setting, in which only the image-level labels are availabl…

2022

LCTR: On Awakening the Local Continuity of Transformer for Weakly Supervised Object Localization

AAAI 2022technical

Weakly supervised object localization (WSOL) aims to learn object localizer solely by using image-level labels. The convolution neural network (CNN) based techniques often result in highlighting the most discriminative part of objects while ignoring the entire object extent. Recently, the transforme…

Cited by 57SourcePDFScholar
2022

Learning Distinctive Margin Toward Active Domain Adaptation

CVPR 2022oral

Despite plenty of efforts focusing on improving the domain adaptation ability (DA) under unsupervised or few-shot semi-supervised settings, recently the solution of active learning started to attract more attention due to its suitability in transferring model in a more practical way with limited ann…

Cited by 42PDFcodeScholar
2022

Prototypical Contrast Adaptation for Domain Adaptive Semantic Segmentation

ECCV 2022poster

"Unsupervised Domain Adaptation (UDA) aims to adapt the model trained on the labeled source domain to an unlabeled target domain. In this paper, we present Prototypical Contrast Adaptation (ProCA), a simple and efficient contrastive learning method for unsupervised domain adaptive semantic segmentat…

2022

SCSNet: An Efficient Paradigm for Learning Simultaneously Image Colorization and Super-resolution

AAAI 2022technical

In the practical application of restoring low-resolution gray-scale images, we generally need to run three separate processes of image colorization, super-resolution, and dows-sampling operation for the target device. However, this pipeline is redundant and inefficient for the independent processes,…

Cited by 15SourcePDFScholar
2021

Analogous to Evolutionary Algorithm: Designing a Unified Sequence Model

NeurIPS 2021poster

Inspired by biological evolution, we explain the rationality of Vision Transformer by analogy with the proven practical Evolutionary Algorithm (EA) and derive that both of them have consistent mathematical representation. Analogous to the dynamic local population in EA, we improve the existing trans…

Cited by 21SourcePDFScholar
2021

Learning Comprehensive Motion Representation for Action Recognition

AAAI 2021technical

For action recognition learning, 2D CNN-based methods are efficient but may yield redundant features due to applying the same 2D convolution kernel to each frame. Recent efforts attempt to capture motion information by establishing inter-frame connections while still suffering the limited temporal r…

Cited by 13SourcePDFScholar
2021

Learning Salient Boundary Feature for Anchor-free Temporal Action Localization

CVPR 2021poster

Temporal action localization is an important yet challenging task in video understanding. Typically, such a task aims at inferring both the action category and localization of the start and end frame for each action instance in a long, untrimmed video. While most current models achieve good results…

Cited by 341PDFcodeScholar
2021

Rethinking Counting and Localization in Crowds: A Purely Point-Based Framework

ICCV 2021poster

Localizing individuals in crowds is more in accordance with the practical demands of subsequent high-level crowd analysis tasks than simply counting. However, existing localization based methods relying on intermediate representations (i.e., density maps or pseudo boxes) serving as learning targets…

Cited by 365PDFcodeScholar
2021

Rethinking Semantic Segmentation From a Sequence-to-Sequence Perspective With Transformers

CVPR 2021poster

Most recent semantic segmentation methods adopt a fully-convolutional network (FCN) with an encoder-decoder architecture. The encoder progressively reduces the spatial resolution and learns more abstract/semantic visual concepts with larger receptive fields. Since context modeling is critical for se…

Cited by 4009PDFcodeScholar
2021

SiamRCR: Reciprocal Classification and Regression for Visual Object Tracking

IJCAI 2021poster

Recently, most siamese network based trackers locate targets via object classification and bounding-box regression. Generally, they select the bounding-box with maximum classification confidence as the final prediction. This strategy may miss the right result due to the accuracy misalignment between…

Cited by 52SourcePDFScholar
2021

To Choose or to Fuse? Scale Selection for Crowd Counting

AAAI 2021technical

In this paper, we address the large scale variation problem in crowd counting by taking full advantage of the multi-scale feature representations in a multi-level network. We implement such an idea by keeping the counting error of a patch as small as possible with a proper feature level selection st…

2021

Uniformity in Heterogeneity: Diving Deep Into Count Interval Partition for Crowd Counting

ICCV 2021poster

Recently, the problem of inaccurate learning targets in crowd counting draws increasing attention. Inspired by a few pioneering work, we solve this problem by trying to predict the indices of pre-defined interval bins of counts instead of the count values themselves. However, an inappropriate interv…

Cited by 49PDFcodeScholar
2020

Chained-Tracker: Chaining Paired Attentive Regression Results for End-to-End Joint Multiple-Object Detection and Tracking

ECCV 2020poster

Existing Multiple-Object Tracking (MOT) methods either follow the tracking-by-detection paradigm to conduct object detection, feature extraction and data association separately, or have two of the three subtasks integrated to form a partially end-to-end solution. Going beyond these sub-optimal frame…

2020

Learning by Analogy: Reliable Supervision From Transformations for Unsupervised Optical Flow Estimation

CVPR 2020poster

Unsupervised learning of optical flow, which leverages the supervision from view synthesis, has emerged as a promising alternative to supervised methods. However, the objective of unsupervised learning is likely to be unreliable in challenging scenes. In this work, we present a framework to use more…

Cited by 213PDFcodeScholar
2020

Temporal Distinct Representation Learning for Action Recognition

ECCV 2020poster

Motivated by the previous success of Two-Dimensional Convolutional Neural Network (2D CNN) on image recognition, researchers endeavor to leverage it to characterize videos. However, one limitation of applying 2D CNN to analyze videos is that different frames of a video share the same 2D CNN kernels,…

Cited by 38SourcePDFScholar
2015

Probabilistic graph based spatial assembly relation inference for programming of assembly task by demonstration

IROS 2015poster

In robot programming by demonstration (PBD) for assembly tasks, one of the important topics is to inference the poses and spatial relations of parts during the demonstration. In this paper, we propose a world model called assembly graph (AG) to achieve this task. The model is able to represent the p…

Cited by 9SourceScholar