← Search

Chuanguang Yang

34 accepted papers

2026

A Comprehensive Survey of Interaction Techniques in 3D Scene Generation

IJCAI 2026

The rapid evolution of 3D scene generation has revolutionized content creation across domains such as gaming, film production, and architectural visualization. Within this landscape, interaction techniques serve as the pivotal bridge connecting user intent with generative models, enabling precise co

Cited by 0Scholar
2026

Asymmetric Cross-Modal Knowledge Distillation: Bridging Modalities with Weak Semantic Consistency

AAAI 2026technical

Cross-modal Knowledge Distillation has demonstrated promising performance on paired modalities with strong semantic connections, referred to as Symmetric Cross-modal Knowledge Distillation (SCKD). However, implementing SCKD becomes exceedingly constrained in real-world scenarios due to the limited a

Cited by 0SourcePDFScholar
2026

EfficientVPR: Toward Efficient Visual Place Recognition via Scene-Aware Prompt Tuning and Adaptive Feature Enhancement

CVPR 2026

Visual place recognition (VPR) faces critical challenges in handling extreme environmental variations while meeting the computational constraints of practical applications. Current methods predominantly address these challenges by either scaling up model capacity or employing computationally intensi

Cited by 0SourcecodeScholar
2026

Fast-SAM3D: 3Dfy Anything in Images but Faster

ICML 2026poster

SAM3D enables scalable, open-world 3D reconstruction from complex scenes, yet its deployment is hindered by prohibitive inference latency. In this work, we conduct the **first systematic investigation** into its inference dynamics, revealing that generic acceleration strategies are brittle in this c…

Cited by 0SourceScholar
2026

MultiAnimate: Pose-Guided Image Animation Made Extensible

CVPR 2026

Pose-guided human image animation aims to synthesize realistic videos of a reference character driven by a sequence of poses. While diffusion-based methods have achieved remarkable success, most existing approaches are limited to single-character animation. We observe that naively extending these me

Cited by 0SourcecodeScholar
2026

Parameterized Prompt for Incremental Object Detection

CVPR 2026

Recent studies have demonstrated that incorporating trainable prompts into pretrained models enables effective incremental learning. However, the application of prompts in incremental object detection (IOD) remains underexplored. Our study reveals that existing prompts-pool-based approaches assume d

Cited by 0SourcecodeScholar
2026

Point-UQ: An Uncertainty-Quantification Paradigm for Point Cloud Few-Shot Class Incremental Learning

ICLR 2026poster

3D few-shot class-incremental learning (3D FSCIL) requires effectively integrating novel classes from limited samples while preserving base-class knowledge, without succumbing to catastrophic forgetting the learned knowledge or overfitting the novel ones. Current 3D FSCIL approaches predominantly f…

Cited by 0SourceScholar
2026

QuantSparse: Comprehensively Compressing Video Diffusion Transformer with Model Quantization and Attention Sparsification

ICLR 2026poster

Diffusion transformers exhibit remarkable video generation capability, yet their prohibitive computational and memory costs hinder practical deployment. Model quantization and attention sparsification are two promising directions for compression, but each alone suffers severe performance degradation…

Cited by 0SourcecodeScholar
2026

Quantized Visual Geometry Grounded Transformer

ICLR 2026poster

Learning-based 3D reconstruction models, represented by Visual Geometry Grounded Transformers (VGGTs), have achieved remarkable progress with large-scale transformers. Their prohibitive computational and memory costs severely hinder real-world deployment. Post-Training Quantization (PTQ) has emerged…

Cited by 0SourcecodeScholar
2026

Representation-Steered Incremental Adapter-Tuning for Class-Incremental Learning with Pre-Trained Models

CVPR 2026

Class-Incremental Learning (CIL) aims to develop models to continuously learn new classes without forgetting learned old ones. Recent advances combine pre-trained models with parameter-efficient fine-tuning, achieving promising results. However, these approaches typically allocate new trainable para

Cited by 0SourcecodeScholar
2026

WorldCache: Accelerating World Models for Free via Heterogeneous Token Caching

ICML 2026poster

Diffusion-based world models have shown strong potential for unified world simulation, but the iterative denoising remains too costly for interactive use and long-horizon rollouts. While feature caching can accelerate inference without training, we find that policies designed for single-modal diffus…

Cited by 0SourceScholar
2025

$\text{S}^2$Q-VDiT: Accurate Quantized Video Diffusion Transformer with Salient Data and Sparse Token Distillation

NeurIPS 2025poster

Diffusion transformers have emerged as the mainstream paradigm for video generation models. However, the use of up to billions of parameters incurs significant computational costs. Quantization offers a promising solution by reducing memory usage and accelerating inference. Nonetheless, we observe t…

Cited by 0SourcecodeScholar
2025

Cross-Layer Graph Knowledge Distillation for Image Recognition

ICASSP 2025accepted

Knowledge Distillation (KD) aims to improve a light-weight student network supervised by a large teacher network. The core idea of KD is to explore valuable knowledge from the teacher. Previous works often extract information from a single sample, but ignore relation modeling among multiple samples…

Cited by 0SourceScholar
2025

Enhancing Image Generation Fidelity via Progressive Prompts

ICASSP 2025accepted

Diffusion transformer (DiT) architecture catches much attention in image generation, which achieves better fidelity, performance, and diversity. However, most existing DiT-based image generation methods are global-aware synthesis and regional prompt control is less explored. In this paper, we propos…

Cited by 0SourceScholar
2025

Frequency-Aligned Knowledge Distillation for Lightweight Spatiotemporal Forecasting

ICCV 2025poster

Spatiotemporal forecasting tasks, such as traffic flow, combustion dynamics, and weather forecasting, often require complex models that suffer from low training efficiency and high memory consumption. This paper proposes a lightweight framework, Spectral Decoupled Knowledge Distillation, which trans…

2025

Geometric Feature Embedding for Effective 3D Few-Shot Class Incremental Learning

ICML 2025poster

3D few-shot class incremental learning (FSCIL) aims to learn new point cloud categories from limited samples while preventing the forgetting of previously learned categories. This research area significantly enhances the capabilities of self-driving vehicles and computer vision systems. Existing 3D…

2025

HSRDiff: A Hierarchical Self-Regulation Diffusion Model for Stochastic Semantic Segmentation

AAAI 2025technical

In safety-critical domains such as medical diagnostics and autonomous driving, single-image evidence is sometimes insufficient to reflect the inherent ambiguity of vision problems. Therefore, multiple plausible assumptions that match the image semantics may be needed to reflect the actual distributi…

2025

MPQ-DM: Mixed Precision Quantization for Extremely Low Bit Diffusion Models

AAAI 2025technical

Diffusion models have received wide attention in generation tasks. However, the expensive computation cost prevents the application of diffusion models in resource-constrained scenarios. Quantization emerges as a practical solution that significantly saves storage and computation by reducing the bit…

2025

Multi-Teacher Knowledge Distillation with Reinforcement Learning for Visual Recognition

AAAI 2025technical

Multi-teacher Knowledge Distillation (KD) transfers diverse knowledge from a teacher pool to a student network. The core problem of multi-teacher KD is how to balance distillation strengths among various teachers. Most existing methods often develop weighting strategies from an individual perspectiv…

2025

Multi-party Collaborative Attention Control for Image Customization

CVPR 2025poster

The rapid development of diffusion models has fueled a growing demand for customized image generation. However, current customization methods face several limitations: 1) typically accept either image or text conditions alone; 2) customization in complex visual scenarios often leads to subject leaka…

2025

OLN++: Improved Object Localization Network for Open-world Object Detection

ICASSP 2025accepted

Open-world object detection (OWOD) is vital for identifying the new objects not encountered during training. Among the various methods for OWOD, Object Proposals without Learning Classification (OPwLC) stands out, with its Object Localization Network (OLN) stressing the localization features. Howeve…

Cited by 0SourceScholar
2025

Prototype-Driven Multi-Feature Generation for Visible-Infrared Person Re-identification

ICASSP 2025accepted

The primary challenges in visible-infrared person re-identification arise from the differences between visible (vis) and infrared (ir) images, including inter-modal and intra-modal variations. These challenges are further complicated by varying viewpoints and irregular movements. Existing methods of…

Cited by 0SourceScholar
2025

Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion Transformers

ICML 2025poster

Diffusion transformers (DiT) have demonstrated exceptional performance in video generation. However, their large number of parameters and high computational complexity limit their deployment on edge devices. Quantization can reduce storage requirements and accelerate inference by lowering the bit-wi…

Cited by 0SourcePDFScholar
2024

CLIP-KD: An Empirical Study of CLIP Model Distillation

CVPR 2024poster

Contrastive Language-Image Pre-training (CLIP) has become a promising language-supervised visual pre-training framework. This paper aims to distill small CLIP models supervised by a large teacher CLIP model. We propose several distillation strategies including relation feature gradient and contrasti…

2024

DetKDS: Knowledge Distillation Search for Object Detectors

ICML 2024poster

In this paper, we present DetKDS, the first framework that searches for optimal detection distillation policies. Manual design of detection distillers becomes challenging and time-consuming due to significant disparities in distillation behaviors between detectors with different backbones, paradigms…

2024

eTag: Class-Incremental Learning via Embedding Distillation and Task-Oriented Generation

AAAI 2024technical

Class incremental learning (CIL) aims to solve the notorious forgetting problem, which refers to the fact that once the network is updated on a new task, its performance on previously-learned tasks degenerates catastrophically. Most successful CIL methods store exemplars (samples of learned tasks) t…

2023

VL-Match: Enhancing Vision-Language Pretraining with Token-Level and Instance-Level Matching

ICCV 2023poster

Vision-Language Pretraining (VLP) has significantly improved the performance of various vision-language tasks with the matching of images and texts. In this paper, we propose VL-Match, a Vision-Language framework with Enhanced Token-level and Instance-level Matching. At the token level, a Vision-Lan…

Cited by 5PDFScholar
2022

Cross-Image Relational Knowledge Distillation for Semantic Segmentation

CVPR 2022poster

Current Knowledge Distillation (KD) methods for semantic segmentation often guide the student to mimic the teacher's structured information generated from individual data samples. However, they ignore the global semantic relations among pixels across various images that are valuable for KD. This pap…

Cited by 247PDFcodeScholar
2022

MixSKD: Self-Knowledge Distillation from Mixup for Image Recognition

ECCV 2022poster

"Unlike the conventional Knowledge Distillation (KD), Self-KD allows a network to learn knowledge from itself without any guidance from extra networks. This paper proposes to perform Self-KD from image Mixture (MixSKD), which integrates these two techniques into a unified framework. MixSKD mutually…

2022

Mutual Contrastive Learning for Visual Representation Learning

AAAI 2022technical

We present a collaborative learning method called Mutual Contrastive Learning (MCL) for general visual representation learning. The core idea of MCL is to perform mutual interaction and transfer of contrastive distributions among a cohort of networks. A crucial component of MCL is Interactive Contra…

2022

Prior Gradient Mask Guided Pruning-Aware Fine-Tuning

AAAI 2022technical

We proposed a Prior Gradient Mask Guided Pruning-aware Fine-Tuning (PGMPF) framework to accelerate deep Convolutional Neural Networks (CNNs). In detail, the proposed PGMPF selectively suppresses the gradient of those ”unimportant” parameters via a prior gradient mask generated by the pruning criteri…

Cited by 43SourcePDFScholar
2021

Hierarchical Self-supervised Augmented Knowledge Distillation

IJCAI 2021poster

Knowledge distillation often involves how to define and transfer knowledge from teacher to student effectively. Although recent self-supervised contrastive knowledge achieves the best performance, forcing the network to learn such knowledge may damage the representation learning of the original clas…