← Search

Qin Zhou

10 accepted papers

2025

FreeCus: Free Lunch Subject-driven Customization in Diffusion Transformers

ICCV 2025poster

In light of recent breakthroughs in text-to-image (T2I) generation, particularly with diffusion transformers (DiT), subject-driven technologies are increasingly being employed for high-fidelity customized production that preserves subject identity from reference inputs, enabling thrilling design wor…

2025

Learnable Retrieval Enhanced Visual-Text Alignment and Fusion for Radiology Report Generation

ICCV 2025poster

Automated radiology report generation is essential for improving diagnostic efficiency and reducing the workload of medical professionals. However, existing methods face significant challenges, such as disease class imbalance and insufficient cross-modal fusion. To address these issues, we propose t…

2025

Semantic-guided Masked Mutual Learning for Multi-modal Brain Tumor Segmentation with Arbitrary Missing Modalities

AAAI 2025technical

Malignant brain tumors have become an aggressive and dangerous disease that leads to death worldwide. Multi-modal MRI data is crucial for accurate brain tumor segmentation, but missing modalities common in clinical practice can severely degrade the segmentation performance. While incomplete multi-mo…

Cited by 0SourcePDFScholar
2024

Advancing Medical Image Segmentation via Self-supervised Instance-adaptive Prototype Learning

IJCAI 2024poster

Medical Image Segmentation (MIS) plays a crucial role in medical therapy planning and robot navigation. Prototype learning methods in MIS focus on generating segmentation masks through pixel-to-prototype comparison. However, current approaches often overlook sample diversity by using a fixed prototy…

Cited by 0SourcePDFScholar
2024

Attention Calibration for Disentangled Text-to-Image Personalization

CVPR 2024poster

Recent thrilling progress in large-scale text-to-image (T2I) models has unlocked unprecedented synthesis quality of AI-generated content (AIGC) including image generation 3D and video composition. Further personalized techniques enable appealing customized production of a novel concept given only se…

2024

The Orthogonality of Weight Vectors: The Key Characteristics of Normalization and Residual Connections

IJCAI 2024poster

Normalization and residual connections find extensive application within the intricate architecture of deep neural networks, contributing significantly to their heightened performance. Nevertheless, the precise factors responsible for this elevated performance have remained elusive. Our theoretical…

2021

Asynchronous Teacher Guided Bit-wise Hard Mining for Online Hashing

AAAI 2021technical

Online hashing for streaming data has attracted increasing attention recently. However, most existing algorithms focus on batch inputs and instance-balanced optimization, which is limited in the single datum input case and does not match the dynamic training in online hashing. Furthermore, constantl…

Cited by 9SourcePDFScholar
2019

Learning to Self-Train for Semi-Supervised Few-Shot Classification

NeurIPS 2019poster

Few-shot classification (FSC) is challenging due to the scarcity of labeled training data (e.g. only one labeled data point per class). Meta-learning has shown to achieve promising results by learning to initialize a classification model for FSC. In this paper we propose a novel semi-supervised meta…

2018

Recognizing Minimal Facial Sketch by Generating Photorealistic Faces With the Guidance of Descriptive Attributes

ICASSP 2018accepted

Cross-modal sketch-photo recognition is of vital importance in law enforcement and public security. Most existing methods are dedicated to bridging the gap between the low-level visual features of sketches and photo images, which is limited due to intrinsic differences in pixel values. In this paper…

Cited by 0SourceScholar
2016

Joint instance and feature importance re-weighting for person reidentification

ICASSP 2016accepted

Person reidentification refers to the task of recognizing the same person under different non-overlapping camera views. Presently, person reidentification based on metric learning is proved to be effective among various techniques, which exploits the labeled data to learn a subspace that maximizes t…

Cited by 0SourceScholar