← Search

Yanyun Qu

62 accepted papers

2026

BeyondSparse: Facilitating Mamba to Enhance Cross-Domain 3D Semantic Segmentation in Adverse Weather

AAAI 2026technical

Domain generalization (DG) and domain adaptation (DA) for 3D semantic segmentation enable the model to maintain high performance while avoiding labor-intensive and time-consuming annotation of target-domain data. However, under adverse weather conditions, the injection of spatial noise will affect t

Cited by 0SourcePDFScholar
2026

Degradation-Consistent Test-Time Adaptation for All-in-One Image Restoration

CVPR 2026

All-in-one image restoration (AiOIR) methods have made remarkable progress in handling diverse degradations. However, their performance often deteriorates when the test distribution deviates from the training distribution. Exploring test-time adaptation for AiOIR is therefore crucial. To adapt a pre

Cited by 0SourcecodeScholar
2026

Diffusion Once and Done: Degradation-Aware LoRA for All-in-One Image Restoration

AAAI 2026technical

Diffusion models have revealed powerful potential in all-in-one image restoration (AiOIR), which is talented in generating abundant texture details. The existing AiOIR methods either retrain a diffusion model or fine-tune the pretrained diffusion model with extra conditional guidance. However, they

Cited by 0SourcePDFScholar
2026

Direct Segmentation without Logits Optimization for Training-Free Open-Vocabulary Semantic Segmentation

CVPR 2026

Open-vocabulary semantic segmentation (OVSS) aims to segment arbitrary category regions in images using open-vocabulary prompts, necessitating that existing methods possess pixel-level vision-language alignment capability. Typically, this capability involves computing the cosine similarity, ie, logi

Cited by 0SourcecodeScholar
2026

PC-CrossDiff: Point-Cluster Dual-Level Cross-Modal Differential Attention for Unified 3D Referring and Segmentation

AAAI 2026technical

3D Visual Grounding (3DVG) aims to localize the referent of natural language referring expressions through two core tasks: Referring Expression Comprehension (3DREC) and Segmentation (3DRES). While existing methods achieve high accuracy in simple, single-object scenes, they suffer from severe perfor

Cited by 0SourcePDFScholar
2026

SpikingIR: A Novel Converted Spiking Neural Network for Efficient Image Restoration

AAAI 2026technical

Image restoration has made great progress with the rise of deep learning, but its energy consumption limits its real-world applications. Spiking Neural Networks (SNNs) are seen as energy-efficient alternatives to Artificial Neural Networks (ANNs). Applying SNNs to image restoration (IR) remains chal

Cited by 0SourcePDFScholar
2026

Target Refocusing via Attention Redistribution for Open-Vocabulary Semantic Segmentation: An Explainability Perspective

AAAI 2026technical

Open-vocabulary semantic segmentation (OVSS) employs pixel-level vision-language alignment to associate category-related prompts with corresponding pixels. A key challenge is enhancing the multimodal dense prediction capability, specifically this pixel-level multimodal alignment. Although existing m

Cited by 0SourcePDFScholar
2026

UZ3DVG: Unaided Zero-Shot 3D Visual Grounding with Generated Language Conditions

CVPR 2026

Zero-Shot 3D Visual Grounding (Zero-Shot 3DVG) aims to localize target objects in 3D scenes from natural language descriptions without relying on instance-wise description annotations. Existing methods rely on extra 2D images during inference and/or require multi-turn interactions with large languag

Cited by 0SourcecodeScholar
2026

UniLDiff: Unlocking the Power of Diffusion Priors for All-in-One Image Restoration

CVPR 2026

All-in-One Image Restoration (AiOIR) has emerged as a promising yet challenging research direction. To address the core challenges of diverse degradation modeling and detail preservation, we propose UniLDiff, a unified framework enhanced with degradation- and detail-aware mechanisms, unlocking the p

Cited by 0SourceScholar
2026

xMHashSeg: Cross-modal Hash Learning for Training-free Unsupervised LiDAR Semantic Segmentation

AAAI 2026technical

3D semantic segmentation serves as a fundamental component in many applications, such as autonomous driving and medical image analysis. Although recent methods have advanced the field, adapting these methods to new environments or object categories without extensive retraining remains a significant

Cited by 0SourcePDFScholar
2025

Large Continual Instruction Assistant

ICML 2025poster

Continual Instruction Tuning (CIT) is adopted to continually instruct Large Models to follow human intent data by data. It is observed that existing gradient update would heavily destroy the performance on previous datasets during CIT process. Instead, Exponential Moving Average (EMA), owns the abil…

2025

MaskViM: Domain Generalized Semantic Segmentation with State Space Models

AAAI 2025technical

Domain Generalized Semantic Segmentation (DGSS) aims to utilize segmentation model training on known source domains to make predictions on unknown target domains. Currently, there are two network architectures: one based on Convolutional Neural Networks (CNNs) and the other based on Visual Transform…

Cited by 0SourcePDFScholar
2025

Multi-Schema Proximity Network for Composed Image Retrieval

ICCV 2025poster

Composed Image Retrieval (CIR) aims to retrieve a target image using a query that combines a reference image and a textual description, benefiting users to express their intent more effectively. Despite significant advances in CIR methods, two unresolved problems remain: 1) existing methods overlook…

Cited by 0SourcePDFScholar
2025

Omni-Query Active Learning for Source-Free Domain Adaptive Cross-Modality 3D Semantic Segmentation

AAAI 2025technical

Source-Free Domain Adaptation (SFDA) aims to transfer a pre-trained source model to the unlabeled target domain without accessing the source data, thereby effectively solving labeled data dependency and domain shift problems. However, the SFDA setting faces a bottleneck due to the absence of supervi…

2025

One-for-More: Continual Diffusion Model for Anomaly Detection

CVPR 2025poster

With the rise of generative models, there is a growing interest in unifying all tasks within a generative framework. Anomaly detection methods also fall into this scope and utilize diffusion models to generate or reconstruct normal samples when given arbitrary anomaly images. However, our study foun…

2025

Task-Aware Prompt Gradient Projection for Parameter-Efficient Tuning Federated Class-Incremental Learning

ICCV 2025poster

Federated Continual Learning (FCL) has recently garnered significant attention due to its ability to continuously learn new tasks while protecting user privacy. However, existing Data-Free Knowledge Transfer (DFKT) methods require training the entire model, leading to high training and communication…

Cited by 0SourcePDFScholar
2024

AdaFormer: Efficient Transformer with Adaptive Token Sparsification for Image Super-resolution

AAAI 2024technical

Efficient transformer-based models have made remarkable progress in image super-resolution (SR). Most of these works mainly design elaborate structures to accelerate the inference of the transformer, where all feature tokens are propagated equally. However, they ignore the underlying characteristic…

Cited by 7SourcePDFScholar
2024

Beyond the Label Itself: Latent Labels Enhance Semi-supervised Point Cloud Panoptic Segmentation

AAAI 2024technical

As the exorbitant expense of labeling autopilot datasets and the growing trend of utilizing unlabeled data, semi-supervised segmentation on point clouds becomes increasingly imperative. Intuitively, finding out more ``unspoken words'' (i.e., latent instance information) beyond the label itself shoul…

Cited by 5SourcePDFScholar
2024

Building a Strong Pre-Training Baseline for Universal 3D Large-Scale Perception

CVPR 2024poster

An effective pre-training framework with universal 3D representations is extremely desired in perceiving large-scale dynamic scenes. However establishing such an ideal framework that is both task-generic and label-efficient poses a challenge in unifying the representation of the same primitive acros…

2024

CLIP-FSAC: Boosting CLIP for Few-Shot Anomaly Classification with Synthetic Anomalies

IJCAI 2024poster

Few-shot anomaly classification (FSAC) is a vital task in manufacturing industry. Recent methods focus on utilizing CLIP in zero/few normal shot anomaly detection instead of custom models. However, there is a lack of specific text prompts in anomaly classification and most of them ignore the modalit…

Cited by 6SourcePDFScholar
2024

CLIP-Guided Federated Learning on Heterogeneity and Long-Tailed Data

AAAI 2024technical

Federated learning (FL) provides a decentralized machine learning paradigm where a server collaborates with a group of clients to learn a global model without accessing the clients' data. User heterogeneity is a significant challenge for FL, which together with the class-distribution imbalance furth…

2024

COTR: Compact Occupancy TRansformer for Vision-based 3D Occupancy Prediction

CVPR 2024poster

The autonomous driving community has shown significant interest in 3D occupancy prediction driven by its exceptional geometric perception and general object recognition capabilities. To achieve this current works try to construct a Tri-Perspective View (TPV) or Occupancy (OCC) representation extendi…

2024

Cross-Modal Match for Language Conditioned 3D Object Grounding

AAAI 2024technical

Language conditioned 3D object grounding aims to find the object within the 3D scene mentioned by natural language descriptions, which mainly depends on the matching between visual and natural language. Considerable improvement in grounding performance is achieved by improving the multimodal fusion…

Cited by 9SourcePDFScholar
2024

Efficient Lightweight Image Denoising with Triple Attention Transformer

AAAI 2024technical

Transformer has shown outstanding performance on image denoising, but the existing Transformer methods for image denoising are with large model sizes and high computational complexity, which is unfriendly to resource-constrained devices. In this paper, we propose a Lightweight Image Denoising Transf…

Cited by 6SourcePDFScholar
2024

Learning Commonality, Divergence and Variety for Unsupervised Visible-Infrared Person Re-identification

NeurIPS 2024poster

Unsupervised visible-infrared person re-identification (USVI-ReID) aims to match specified persons in infrared images to visible images without annotations, and vice versa. USVI-ReID is a challenging yet underexplored task. Most existing methods address the USVI-ReID through cluster-based contrastiv…

2024

Learning Task-Aware Language-Image Representation for Class-Incremental Object Detection

AAAI 2024technical

Class-incremental object detection (CIOD) is a real-world desired capability, requiring an object detector to continuously adapt to new tasks without forgetting learned ones, with the main challenge being catastrophic forgetting. Many methods based on distillation and replay have been proposed to al…

Cited by 5SourcePDFScholar
2024

One-Stage Training Generative Paradigm for Generalized Zero-Shot Learning

ICASSP 2024accepted

Zero-shot learning image classification aims to identify unseen classes not present during training. Generalized zero-shot learning (GZSL) is more in line with realistic scenarios due to its ability of recognizing both seen and unseen classes. Current GZSL methods mostly utilize generative adversari…

Cited by 0SourceScholar
2024

Prompt Gradient Projection for Continual Learning

ICLR 2024spotlight

Prompt-tuning has demonstrated impressive performance in continual learning by querying relevant prompts for each input instance, which can avoid the introduction of task identifier. Its forgetting is therefore reduced as this instance-wise query mechanism enables us to select and update only releva…

2024

PromptAD: Learning Prompts with only Normal Samples for Few-Shot Anomaly Detection

CVPR 2024poster

The vision-language model has brought great improvement to few-shot industrial anomaly detection which usually needs to design of hundreds of prompts through prompt engineering. For automated scenarios we first use conventional prompt learning with many-class paradigm as the baseline to automaticall…

2024

Relationship Prompt Learning is Enough for Open-Vocabulary Semantic Segmentation

NeurIPS 2024poster

Open-vocabulary semantic segmentation (OVSS) aims to segment unseen classes without corresponding labels. Existing Vision-Language Model (VLM)-based methods leverage VLM's rich knowledge to enhance additional explicit segmentation-specific networks, yielding competitive results, but at the cost of e…

Cited by 0SourcePDFScholar
2024

SkipDiff: Adaptive Skip Diffusion Model for High-Fidelity Perceptual Image Super-resolution

AAAI 2024technical

It is well-known that image quality assessment usually meets with the problem of perception-distortion (p-d) tradeoff. The existing deep image super-resolution (SR) methods either focus on high fidelity with pixel-level objectives or high perception with generative models. The emergence of diffusion…

Cited by 7SourcePDFScholar
2024

UniDSeg: Unified Cross-Domain 3D Semantic Segmentation via Visual Foundation Models Prior

NeurIPS 2024poster

3D semantic segmentation using an adapting model trained from a source domain with or without accessing unlabeled target-domain data is the fundamental task in computer vision, containing domain adaptation and domain generalization. The essence of simultaneously solving cross-domain tasks is to enha…

2023

BEV-DG: Cross-Modal Learning under Bird's-Eye View for Domain Generalization of 3D Semantic Segmentation

ICCV 2023poster

Cross-modal Unsupervised Domain Adaptation (UDA) aims to exploit the complementarity of 2D-3D data to overcome the lack of annotation in a new domain. However, UDA methods rely on access to the target domain during training, meaning the trained model only works in a specific target domain. In light…

Cited by 17PDFScholar
2023

Dual Pseudo-Labels Interactive Self-Training for Semi-Supervised Visible-Infrared Person Re-Identification

ICCV 2023poster

Visible-infrared person re-identification (VI-ReID) aims to match a specific person from a gallery of images captured from non-overlapping visible and infrared cameras. Most works focus on fully supervised VI-ReID, which requires substantial cross-modality annotation that is more expensive than the…

Cited by 42PDFcodeScholar
2023

Efficient Converted Spiking Neural Network for 3D and 2D Classification

ICCV 2023poster

Spiking Neural Networks (SNNs) have attracted enormous research interest due to their low-power and biologically plausible nature. Existing ANN-SNN conversion methods can achieve lossless conversion by converting a well-trained Artificial Neural Network (ANN) into an SNN. However, converted SNN requ…

Cited by 16PDFScholar
2023

Instance and Category Supervision are Alternate Learners for Continual Learning

ICCV 2023poster

Continual Learning (CL) is the constant development of complex behaviors by building upon previously acquired skills. Yet, current CL algorithms tend to incur class-level forgetting as the label information is often quickly overwritten by new knowledge. This motivates attempts to mine instance-level…

Cited by 2PDFScholar
2023

Learning Re-sampling Methods with Parameter Attribution for Image Super-resolution

NeurIPS 2023poster

Single image super-resolution (SISR) has made a significant breakthrough benefiting from the prevalent rise of deep neural networks and large-scale training samples. The mainstream deep SR models primarily focus on network architecture design as well as optimization schemes, while few pay attention…

Cited by 3SourcePDFScholar
2023

Memory-Friendly Scalable Super-Resolution via Rewinding Lottery Ticket Hypothesis

CVPR 2023poster

Scalable deep Super-Resolution (SR) models are increasingly in demand, whose memory can be customized and tuned to the computational recourse of the platform. The existing dynamic scalable SR methods are not memory-friendly enough because multi-scale models have to be saved with a fixed size for eac…

Cited by 9SourcePDFScholar
2023

Multi-Centroid Task Descriptor for Dynamic Class Incremental Inference

CVPR 2023poster

Incremental learning could be roughly divided into two categories, i.e., class- and task-incremental learning. The main difference is whether the task ID is given during evaluation. In this paper, we show this task information is indeed a strong prior knowledge, which will bring significant improvem…

Cited by 5SourcePDFScholar
2023

Rethinking Gradient Projection Continual Learning: Stability / Plasticity Feature Space Decoupling

CVPR 2023poster

Continual learning aims to incrementally learn novel classes over time, while not forgetting the learned knowledge. Recent studies have found that learning would not forget if the updated gradient is orthogonal to the feature space. However, previous approaches require the gradient to be fully ortho…

Cited by 29SourcePDFScholar
2023

VS-Boost: Boosting Visual-Semantic Association for Generalized Zero-Shot Learning

IJCAI 2023poster

Unlike conventional zero-shot learning (CZSL) which only focuses on the recognition of unseen classes by using the classifier trained on seen classes and semantic embeddings, generalized zero-shot learning (GZSL) aims at recognizing both the seen and unseen classes, so it is more challenging due to…

Cited by 17SourcePDFScholar
2023

Weakly Supervised 3D Segmentation via Receptive-Driven Pseudo Label Consistency and Structural Consistency

AAAI 2023technical

As manual point-wise label is time and labor-intensive for fully supervised large-scale point cloud semantic segmentation, weakly supervised method is increasingly active. However, existing methods fail to generate high-quality pseudo labels effectively, leading to unsatisfactory results. In this p…

Cited by 11SourcePDFScholar
2022

Comprehensive Regularization in a Bi-directional Predictive Network for Video Anomaly Detection

AAAI 2022technical

Video anomaly detection aims to automatically identify unusual objects or behaviours by learning from normal videos. Previous methods tend to use simplistic reconstruction or prediction constraints, which leads to the insufficiency of learned representations for normal data. As such, we propose a no…

Cited by 79SourcePDFScholar
2022

En-Compactness: Self-Distillation Embedding & Contrastive Generation for Generalized Zero-Shot Learning

CVPR 2022poster

Generalized zero-shot learning (GZSL) requires a classifier trained on seen classes that can recognize objects from both seen and unseen classes. Due to the absence of unseen training samples, the classifier tends to bias towards seen classes. To mitigate this problem, feature generation based model…

Cited by 89PDFScholar
2022

Optimal Transport for Label-Efficient Visible-Infrared Person Re-identification

ECCV 2022poster

"Visible-infrared person re-identification (VI-ReID) has been a key enabler for night intelligent monitoring system. However, the extensive laboring efforts significantly limit its applications. In this paper, we raise a new label-efficient training pipeline for VI-ReID. Our observation is: RGB ReID…

2021

Boundary-Aware Geometric Encoding for Semantic Segmentation of Point Clouds

AAAI 2021technical

Boundary information plays a significant role in 2D image segmentation, while usually being ignored in 3D point cloud segmentation where ambiguous features might be generated in feature extraction, leading to misclassification in the transition area between two objects. In this paper, firstly, we pr…

2021

Contrastive Learning for Compact Single Image Dehazing

CVPR 2021poster

Single image dehazing is a challenging ill-posed problem due to the severe information degeneration. However, existing deep learning based dehazing methods only adopt clear images as positive samples to guide the training of dehazing network while negative information is unexploited. Moreover, most…

Cited by 876PDFcodeScholar
2021

Farewell to Mutual Information: Variational Distillation for Cross-Modal Person Re-Identification

CVPR 2021poster

The Information Bottleneck (IB) provides an information theoretic principle for representation learning, by retaining all information relevant for predicting label while minimizing the redundancy. Though IB principle has been applied to a wide range of applications, its optimization remains a challe…

Cited by 165PDFcodeScholar
2021

Learn from Concepts: Towards the Purified Memory for Few-shot Learning

IJCAI 2021poster

Human beings have a great generalization ability to recognize a novel category by only seeing a few number of samples. This is because humans possess the ability to learn from the concepts that already exist in our minds. However, many existing few-shot approaches fail in addressing such a fundament…

Cited by 12SourcePDFScholar
2021

Omni-Supervised Point Cloud Segmentation via Gradual Receptive Field Component Reasoning

CVPR 2021poster

Hidden features in neural network usually fail to learn informative representation for 3D segmentation as supervisions are only given on output prediction, while this can be solved by omni-scale supervision on intermediate layers. In this paper, we bring the first omni-scale supervision method to po…

Cited by 61PDFcodeScholar
2021

Perturbed Self-Distillation: Weakly Supervised Large-Scale Point Cloud Semantic Segmentation

ICCV 2021poster

Large-scale point cloud semantic segmentation has wide applications. Current popular researches mainly focus on fully supervised learning which demands expensive and tedious manual point-wise annotation. Weakly supervised learning is an alternative way to avoid this exhausting annotation. However, f…

Cited by 162PDFScholar
2021

Towards Compact Single Image Super-Resolution via Contrastive Self-distillation

IJCAI 2021poster

Convolutional neural networks (CNNs) are highly successful for super-resolution (SR) but often require sophisticated architectures with heavy memory cost and computational overhead significantly restricts their practical deployments on resource-limited devices. In this paper, we proposed a novel con…

2021

Weakly Supervised Semantic Segmentation for Large-Scale Point Cloud

AAAI 2021technical

Existing methods for large-scale point cloud semantic segmentation require expensive, tedious and error-prone manual point-wise annotation. Intuitively, weakly supervised training is a direct solution to reduce the labeling costs. However, for weakly supervised large-scale point cloud semantic segme…

2020

LatticeNet: Towards Lightweight Image Super-resolution with Lattice Block

ECCV 2020poster

Deep neural networks with a massive number of layers have made a remarkable breakthrough on single image super-resolution (SR), but sacrifice computation complexity and memory storage. To address this problem, we focus on the lightweight models for fast and accurate image SR. Due to the frequent use…

2020

Meta Segmentation Network for Ultra-Resolution Medical Images

IJCAI 2020poster

Despite recent great progress on semantic segmentation, there still exist huge challenges in medical ultra-resolution image segmentation. The methods based on multi-branch structure can make a good balance between computational burdens and segmentation accuracy. However, the fusion structure in thes…

Cited by 0SourcePDFScholar