← Search

Haoran Xie

25 accepted papers

2026

FreqEdit: Preserving High-Frequency Features for Robust Multi-Turn Image Editing

CVPR 2026

Instruction-based image editing through natural language has emerged as a powerful paradigm for intuitive visual manipulation. While recent models achieve impressive results on single edits, they suffer from severe quality degradation under multi-turn editing. Through systematic analysis, we identif

Cited by 0SourcecodeScholar
2026

Neural-Inspired Modeling of Auditory Selection and Compensation for Audio-Visual Speech Separation

ICML 2026poster

Current audio-visual speech separation (AVSS) models typically rely on implicit multimodal fusion, but the absence of explicit modality alignment and reliability modeling often causes semantic misalignment and contaminates speech representations. The brain addresses this with a hierarchy: top-down a…

Cited by 0SourceScholar
2026

RGGT: A Generative-Prior-Guided Transformer for Unified Rigid and Non-Rigid Point Cloud Registration

ICML 2026poster

Point cloud registration can be categorized into rigid and non-rigid settings depending on the motion characteristics of the underlying objects. Rigid alignment assumes a single global transformation under which corresponding points remain geometrically consistent across scales, whereas non-rigid al…

Cited by 0SourceScholar
2026

Simulation-Ready Tree: High-Quality Dynamic Tree Reconstruction from a Single RGB-D Sensor

ICRA 2026poster

Realistic animation of real trees is challenging due to the difficulty in accurately capturing and simulating their movements under varying environmental conditions. Most of real tree reconstruction methods focus on the static modeling of trees from RGB images or LiDAR point clouds. Rather than RGB …

Cited by 0Scholar
2026

Sketch-Guided Anime Hair Editing Using Multimodal Diffusion Transformer (Student Abstract)

AAAI 2026technical

Anime hair design is crucial but challenging, as it conveys personality and emotion through stylized geometry and layered structure. In this work, we propose a sketch-guided approach for intuitive control of multimodal diffusion transformers (MMDiT) to generate semantically consistent anime hairstyl

Cited by 0SourcePDFScholar
2025

A Survey on Multi-View Knowledge Graph: Generation, Fusion, Applications and Future Directions

IJCAI 2025

Knowledge Graphs (KGs) have revolutionized structured knowledge representation, yet their capacity to model real-world complexity and heterogeneity remains fundamentally constrained. The emerging paradigm of Multi-View Knowledge Graphs (MVKGs) addresses this gap through multi-view learning, but exis

Cited by 0SourcePDFScholar
2025

CondAmbigQA: A Benchmark and Dataset for Conditional Ambiguous Question Answering

EMNLP 2025

Users often assume that large language models (LLMs) share their cognitive alignment of context and intent, leading them to omit critical information in question-answering (QA) and produce ambiguous queries. Responses based on misaligned assumptions may be perceived as hallucinations. Therefore, ide

2025

LineArt: A Knowledge-guided Training-free High-quality Appearance Transfer for Design Drawing with Diffusion Model

CVPR 2025poster

Image rendering from line drawings is vital in design and image generation technologies reduce costs, yet professional line drawings demand preserving complex details. Text prompts struggle with accuracy, and image translation struggles with consistency and fine-grained control. We present LineArt,…

Cited by 1SourcePDFScholar
2024

Cross Initialization for Face Personalization of Text-to-Image Models

CVPR 2024poster

Recently there has been a surge in face personalization techniques benefiting from the advanced capabilities of pretrained text-to-image diffusion models. Among these a notable method is Textual Inversion which generates personalized images by inverting given images into textual embeddings. However…

2023

A Semi-Automatic Oriental Ink Painting Framework for Robotic Drawing From 3D Models

RA-L 2023

Creating visually pleasing stylized ink paintings from 3D models is a challenge in robotic manipulation. We propose a semi-automatic framework that can extract expressive strokes from 3D models and draw them in oriental ink painting styles by using a robotic arm. The framework consists of a simulati

Cited by 2SourceScholar
2023

Geogcn: Geometric Dual-Domain Graph Convolution Network For Point Cloud Denoising

ICASSP 2023accepted

We propose GeoGCN, a novel geometric dual-domain graph convolution network for point cloud denoising (PCD). Beyond the traditional wisdom of PCD, to fully exploit the geometric information of point clouds, we define two kinds of surface normals, one is called Real Normal (RN), and the other is Virtu…

Cited by 0SourceScholar
2023

ISmallNet: Densely Nested Network with Label Decoupling for Infrared Small Target Detection

ICASSP 2023accepted

Small targets are often submerged in cluttered backgrounds of infrared images. Conventional detectors tend to generate false alarms, while CNN-based detectors lose small targets in deep layers. To this end, we propose iSmallNet, a multi-stream densely nested network with label decoupling for infrare…

Cited by 0SourceScholar
2023

Recurrent Attention Networks for Long-text Modeling

ACL 2023findings

Self-attention-based models have achieved remarkable progress in short-text mining. However, the quadratic computational complexities restrict their application in long text processing. Prior works have adopted the chunking strategy to divide long documents into chunks and stack a self-attention bac…

2023

TransFace: Calibrating Transformer Training for Face Recognition from a Data-Centric Perspective

ICCV 2023poster

Vision Transformers (ViTs) have demonstrated powerful representation ability in various visual tasks thanks to their intrinsic data-hungry nature. However, we unexpectedly find that ViTs perform vulnerably when applied to face recognition (FR) scenarios with extremely large datasets. We investigate…

Cited by 34PDFcodeScholar
2023

ifUNet++: Iterative Feedback UNet++ for Infrared Small Target Detection

ICASSP 2023accepted

Small targets are often submerged in the cluttered backgrounds of infrared images. In this paper, we propose an iterative feedback UNet++ for infrared small target detection, dubbed ifUNet++. Unlike most of existing methods, ifU-Net++ enables to concentrate on small targets while weakening the inter…

Cited by 0SourceScholar
2022

Fine-tuning Deep Neural Networks by Interactively Refining the 2D Latent Space of Ambiguous Images

IJCAI 2022poster

Deep neural networks (DNNs) have achieved excellent results currently in classification, while they may still suffer from ambiguous images which are similar across classes. By contrast, humans have a relatively good ability to distinguish these categories of images. Therefore, we propose a human-in-…

Cited by 5SourcePDFScholar
2022

I Can Find You! Boundary-Guided Separated Attention Network for Camouflaged Object Detection

AAAI 2022technical

Can you find me? By simulating how humans to discover the so-called 'perfectly'-camouflaged object, we present a novel boundary-guided separated attention network (call BSA-Net). Beyond the existing camouflaged object detection (COD) wisdom, BSA-Net utilizes two-stream separated attention modules to…

2022

MBA-RainGAN: A Multi-Branch Attention Generative Adversarial Network for Mixture of Rain Removal

ICASSP 2022accepted

Rain severely degrades the visibility of scene objects, especially when images are captured through the glass under rainy weather. We observe three intriguing phenomena: 1) rain is a mixture of raindrops, rain streaks and rainy haze; 2) the depth from the camera determines the degree of object visib…

Cited by 0SourceScholar
2021

Direction-aware Feature-level Frequency Decomposition for Single Image Deraining

IJCAI 2021poster

We present a novel direction-aware feature-level frequency decomposition network for single image deraining. Compared with existing solutions, the proposed network has three compelling characteristics. First, unlike previous algorithms, we propose to perform frequency decomposition at feature-level…

Cited by 3SourcePDFScholar
2021

Merging Statistical Feature via Adaptive Gate for Improved Text Classification

AAAI 2021technical

Currently, text classification studies mainly focus on training classifiers by using textual input only, or enhancing semantic features by introducing external knowledge (e.g., hand-craft lexicons and domain knowledge). In contrast, some intrinsic statistical features of the corpus, like word freque…

2021

Tree-Structured Topic Modeling with Nonparametric Neural Variational Inference

ACL 2021long

Topic modeling has been widely used for discovering the latent semantic structure of documents, but most existing methods learn topics with a flat structure. Although probabilistic models can generate topic hierarchies by introducing nonparametric priors like Chinese restaurant process, such methods…

2020

Detail-recovery Image Deraining via Context Aggregation Networks

CVPR 2020poster

This paper looks at this intriguing question: are single images with their details lost during deraining, reversible to their artifact-free status? We propose an end-to-end detail-recovery image deraining network (termed a DRDNet) to solve the problem. Unlike existing image deraining approaches that…

Cited by 217PDFcodeScholar
2020

Geometry and Learning Co-Supported Normal Estimation for Unstructured Point Cloud

CVPR 2020poster

In this paper, we propose a normal estimation method for unstructured point cloud. We observe that geometric estimators commonly focus more on feature preservation but are hard to tune parameters and sensitive to noise, while learning-based approaches pursue an overall normal estimation accuracy but…

Cited by 42PDFScholar
2017

Least Squares Generative Adversarial Networks

ICCV 2017poster

Unsupervised learning with generative adversarial networks (GANs) has proven hugely successful. Regular GANs hypothesize the discriminator as a classifier with the sigmoid cross entropy loss function. However, we found that this loss function may lead to the vanishing gradients problem during the le…

Cited by 6527PDFcodeScholar