← Search

Huaidong Zhang

21 accepted papers

2026

Healthcare Insurance Fraud Detection via Continual Fiedler Vector Graph Model

ICLR 2026poster

Healthcare insurance fraud detection presents unique machine learning challenges: labeled data are scarce due to delayed verification processes, and fraudulent behaviors evolve rapidly, often manifesting in complex, graph-structured interactions. Existing methods struggle in such settings. Pretraini…

Cited by 0SourceScholar
2026

M$^3$E: Continual Vision-and-Language Navigation via Mixture of Macro and Micro Experts

ICLR 2026poster

Vision-and-Language Navigation (VLN) agents have shown strong capabilities in following natural language instructions. However, they often struggle to generalize across environments due to catastrophic forgetting, which limits their practical use in real-world settings where agents must continually…

Cited by 0SourceScholar
2025

FR²Seg: Continual Segmentation Across Multiple Sites via Fourier Style Replay and Adaptive Consistency Regularization

AAAI 2025technical

In clinical imaging, medical segmentation networks typically require continually adapting to new data from multiple sites over time, as aggregating all data for learning at once can be impractical due to storage limitations and privacy concerns. However, existing methods basically overlook domain-s…

2025

Let's Chorus: Partner-aware Hybrid Song-Driven 3D Head Animation

CVPR 2025poster

Singing is a vital form of human emotional expression and social interaction, distinguished from speech by its richer emotional nuances and freer expressive style. Thus, investigating 3D facial animation driven by singing holds significant research value. Our work focuses on 3D singing facial animat…

Cited by 0SourcePDFScholar
2025

MODfinity: Unsupervised Domain Adaptation with Multimodal Information Flow Intertwining

CVPR 2025poster

Multimodal unsupervised domain adaptation leverages unlabeled data in the target domain to enhance multimodal systems continuously. While current state-of-the-art methods encourage interaction between sub-models of different modalities through pseudo-labeling and feature-level exchange, varying samp…

Cited by 2SourcePDFScholar
2025

MSSDA: Multi-Sub-Source Domain Adaptation for Diabetic Foot Neuropathy Recognition

AAAI 2025technical

Diabetic foot neuropathy (DFN) is a critical factor leading to diabetic foot ulcers, which is one of the most common and severe complications of diabetes mellitus (DM) and is associated with high risks of amputation and mortality. Despite its significance, existing datasets do not directly derive fr…

2025

NexusGS: Sparse View Synthesis with Epipolar Depth Priors in 3D Gaussian Splatting

CVPR 2025highlight

Neural Radiance Field (NeRF) and 3D Gaussian Splatting (3DGS) have noticeably advanced photo-realistic novel view synthesis using images from densely spaced camera viewpoints. However, these methods struggle in few-shot scenarios due to limited supervision. In this paper, we present NexusGS, a 3DGS-…

2025

PersonaMagic: Stage-Regulated High-Fidelity Face Customization with Tandem Equilibrium

AAAI 2025technical

Personalized image generation has made significant strides in adapting content to novel concepts. However, a persistent challenge remains: balancing the accurate reconstruction of unseen concepts with the need for editability according to the prompt, especially when dealing with the complex nuances…

2025

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition

ICCV 2025poster

3D visual grounding aims to identify and localize objects in a 3D space based on textual descriptions. However, existing methods struggle with disentangling targets from anchors in complex multi-anchor queries and resolving inconsistencies in spatial descriptions caused by perspective variations.To…

2024

Beat-It: Beat-Synchronized Multi-Condition 3D Dance Generation

ECCV 2024oral

"Dance, as an art form, fundamentally hinges on the precise synchronization with musical beats. However, achieving aesthetically pleasing dance sequences from music is challenging, with existing methods often falling short in controllability and beat alignment. To address these shortcomings, this pa…

2024

Beyond Textual Constraints: Learning Novel Diffusion Conditions with Fewer Examples

CVPR 2024poster

In this paper we delve into a novel aspect of learning novel diffusion conditions with datasets an order of magnitude smaller. The rationale behind our approach is the elimination of textual constraints during the few-shot learning process. To that end we implement two optimization strategies. The f…

2024

D3still: Decoupled Differential Distillation for Asymmetric Image Retrieval

CVPR 2024poster

Existing methods for asymmetric image retrieval employ a rigid pairwise similarity constraint between the query network and the larger gallery network. However these one-to-one constraint approaches often fail to maintain retrieval order consistency especially when the query network has limited repr…

2024

D4-VTON: Dynamic Semantics Disentangling for Differential Diffusion based Virtual Try-On

ECCV 2024poster

"In this paper, we introduce D4 -VTON, an innovative solution for image-based virtual try-on. We address challenges from previous studies, such as semantic inconsistencies before and after garment warping, and reliance on static, annotation-driven clothing parsers. Additionally, we tackle the comple…

2024

Mask4Align: Aligned Entity Prompting with Color Masks for Multi-Entity Localization Problems

CVPR 2024poster

In Visual Question Answering (VQA) recognizing and localizing entities pose significant challenges. Pretrained vision-and-language models have addressed this problem by providing a text description as the answer. However in visual scenes with multiple entities textual descriptions struggle to distin…

Cited by 0SourcePDFScholar
2023

Towards a Smaller Student: Capacity Dynamic Distillation for Efficient Image Retrieval

CVPR 2023poster

Previous Knowledge Distillation based efficient image retrieval methods employ a lightweight network as the student model for fast inference. However, the lightweight student model lacks adequate representation capacity for effective knowledge imitation during the most critical early training period…

Cited by 22SourcePDFScholar
2023

Where Is My Spot? Few-Shot Image Generation via Latent Subspace Optimization

CVPR 2023poster

Image generation relies on massive training data that can hardly produce diverse images of an unseen category according to a few examples. In this paper, we address this dilemma by projecting sparse few-shot samples into a continuous latent space that can potentially generate infinite unseen samples…

2021

Learning Semantic Context from Normal Samples for Unsupervised Anomaly Detection

AAAI 2021technical

Unsupervised anomaly detection aims to identify data samples that have low probability density from a set of input samples, and only the normal samples are provided for model training. The inference of abnormal regions on the input image requires an understanding of the surrounding semantic context.…

Cited by 179SourcePDFScholar
2021

Object Detection in Densely Packed Scenes via Semi-Supervised Learning with Dual Consistency

IJCAI 2021poster

Deep neural networks have been shown to be very powerful tools for object detection in various scenes. Their remarkable performance, however, heavily depends on the availability of a large number of high quality labeled data, which are time-consuming and costly to acquire for scenes with densely pac…

2020

Context-Aware and Scale-Insensitive Temporal Repetition Counting

CVPR 2020poster

Temporal repetition counting aims to estimate the number of cycles of a given repetitive action. Existing deep learning methods assume repetitive actions are performed in a fixed time-scale, which is invalid for the complex repetitive actions in real life. In this paper, we tailor a context-aware an…

Cited by 71PDFcodeScholar