← Search

Haoran Duan

10 accepted papers

2026

Debiasing Diffusion Priors via 3D Attention for Consistent Gaussian Splatting

AAAI 2026technical

Versatile 3D tasks (e.g., generation or editing) distilling Text-to-Image (T2I) diffusion models have attracted significant research interest for not relying on extensive 3D training data. However, T2I models exhibit limitations resulting from prior view bias, which produces conflicting appearances

Cited by 0SourcePDFScholar
2026

ReaSon: Reinforced Causal Search with Information Bottleneck for Video Understanding

AAAI 2026technical

Keyframe selection has become essential for video understanding with vision-language models (VLMs) due to limited input tokens and the temporal sparsity of relevant information across video frames. Video understanding often relies on effective keyframes that are not only informative but also causall

Cited by 0SourcePDFScholar
2026

Time Shuffle: A Transferability-Booster for Multiple Audio Adversarial Tasks

AAAI 2026technical

Existing audio adversarial attack methods suffer from poor transferability, primarily due to insufficient exploration of model decision mechanisms and overreliance on heuristic-driven algorithm design. This paper aims to alleviate this gap. Specifically, through observations across three mainstream

Cited by 0SourcePDFScholar
2026

vMFCoOp: Towards Equilibrium on a Unified Hyperspherical Manifold for Prompting Biomedical VLMs

AAAI 2026technical

Recent advances in context optimization (CoOp) guided by large language model (LLM)–distilled medical semantic priors offer a scalable alternative to manual prompt engineering and full fine-tuning for adapting biomedical CLIP-based vision-language models (VLMs). However, prompt learning in this cont

Cited by 1SourcePDFScholar
2025

Highlight What You Want: Weakly-Supervised Instance-Level Controllable Infrared-Visible Image Fusion

ICCV 2025poster

Infrared and visible image fusion (VIS-IR) aims to integrate complementary information from both source images to produce a fused image with enriched details. However, most existing fusion models lack controllability, making it difficult to customize the fused output according to user preferences. T…

2025

LAGD: Local Topological-Alignment and Global Semantic-Deconstruction for Incremental 3D Semantic Segmentation

AAAI 2025technical

Numerous deep learning-based works focusing on 3D semantic segmentation have been proposed and have achieved impressive performance. However, due to the catastrophic forgetting, existing methods will degrade dramatically in a real-world scenario where new 3D semantic categories are arriving continua…

Cited by 0SourcePDFScholar
2025

Multi-Modal Medical Image Fusion via 3D Manifold Fitting and Dual-Domain Cross-Attention

ICASSP 2025accepted

Medical image fusion (MIF) aims to extract complementary features from multi-modal source images and fuse them into a single image to assist in clinical diagnostics. Despite its importance, MIF faces two primary challenges: the lack of tailored paradigms for CMSF extraction and insufficient dual exp…

Cited by 0SourceScholar
2025

NoiseHGNN: Synthesized Similarity Graph-Based Neural Network for Noised Heterogeneous Graph Representation Learning

AAAI 2025technical

Real-world graph data environments intrinsically exist noise (e.g., link and structure errors) that inevitably disturb the effectiveness of graph representation and downstream learning tasks. For homogeneous graphs, the latest works use original node features to synthesize a similarity graph that ca…

2025

Rethinking Score Distilling Sampling for 3D Editing and Generation

ICML 2025poster

Score Distillation Sampling (SDS) has emerged as a prominent method for text-to-3D generation by leveraging the strengths of 2D diffusion models. However, SDS is limited to generation tasks and lacks the capability to edit existing 3D assets. Conversely, variants of SDS that introduce editing capabi…

Cited by 0SourcePDFScholar
2025

Towards Scalable Spatial Intelligence via 2D-to-3D Data Lifting

ICCV 2025poster

Spatial intelligence is emerging as a transformative frontier in AI, yet it remains constrained by the scarcity of large-scale 3D datasets. Unlike the abundant 2D imagery, acquiring 3D data typically requires specialized sensors and laborious annotation. In this work, we present a scalable pipeline…