← Search

Haojie Li

29 accepted papers

2026

Breaking Alignment Barriers: TPS-Driven Semantic Correlation Learning for Alignment-Free RGB-T Salient Object Detection

AAAI 2026technical

Existing RGB-T salient object detection methods predominantly rely on manually aligned and annotated datasets, struggling to handle real-world scenarios with raw, unaligned RGB-T image pairs. In practical applications, due to significant cross-modal disparities such as spatial misalignment, scale va

Cited by 0SourcePDFScholar
2025

2.5D Top-K Ranked Multiple Instance Learning to Classify NSCLC PD-L1 Status on CT Images

ICASSP 2025accepted

Classifying the status of NSCLC PD-L1 on chest CT is a cost-effective and non-invasive method. The existing multiple instance learning (MIL) methods are not effective for this task, due to the lack of an efficient feature encoder for 3D instances and ignoring the importance of representative instanc…

Cited by 0SourceScholar
2025

Application of soft constraints on mirror position to improve robustness of optical target positioning in shallow water

IROS 2025

The unique optical characteristics of the underwater environment, such as light refraction and loss of salient features, pose a significant challenge to traditional vision sensors, especially in the swarm operation scenario where multiple autonomous underwater vehicles (AUVs) cooperate with the moth

Cited by 0SourceScholar
2025

Delving into Transformer-based Network Architecture for Guided Depth Super-Resolution

ICASSP 2025accepted

Guided Depth Super-Resolution (GDSR) enhances low-resolution (LR) depth maps by leveraging high-resolution (HR) color images. The primary challenges involve achieving effective cross-modal data alignment and fusion, as well as incorporating multi-scale information within the Transformer architecture…

Cited by 0SourceScholar
2025

Few-shot Image Classification based on Attribute Prediction and Selection

ICASSP 2025accepted

Few-shot learning addresses the challenges of image classification with limited samples, but current methods often fail to fully utilize sample correlations and external semantic information, leading to low accuracy. To overcome these limitations, we propose a few-shot image classification method ba…

Cited by 0SourceScholar
2025

Is this Generated Person Existed in Real-world? Fine-grained Detecting and Calibrating Abnormal Human-body

CVPR 2025highlight

Recent improvements in visual synthesis have significantly enhanced the depiction of generated human photos, which are pivotal due to their wide applicability and demand. Nonetheless, the existing text-to-image or text-to-video models often generate low-quality human photos that might differ conside…

Cited by 1SourcePDFScholar
2025

Mining Scene Structural Guidance for Thermal Images in Self-Supervised Monocular Depth Estimation

ICASSP 2025accepted

Self-supervised monocular depth estimation from RGB images has seen significant advancements recently, primarily because it eliminates the need for ground truth data during training. However, applying this technique to thermal images remains challenging due to their inherent characteristics, such as…

Cited by 0SourceScholar
2025

Rethinking Joint Maximum Mean Discrepancy for Visual Domain Adaptation

NeurIPS 2025oral

In domain adaption (DA), joint maximum mean discrepancy (JMMD), as a famous distribution-distance metric, aims to measure joint probability distribution difference between the source domain and target domain, while it is still not fully explored and especially hard to be applied into a subspace-lear…

Cited by 0SourceScholar
2025

Self-Supervised Monocular Depth Estimation from Videos via Pose-Adaptive Reconstruction

ICASSP 2025accepted

Self-supervised depth estimation from videos involves predicting the depth map of a target frame and the pose changes between source and target frames. The reconstructed source frame is aligned with the target view using the predicted pose and depth information. Precise pose estimation significantly…

Cited by 0SourceScholar
2024

Bridging the Gap: Sketch to Color Diffusion Model with Semantic Prompt Learning

ICASSP 2024accepted

Automatic anime sketch colorization aims to generate a color image from a sketch image, which is challenging due to limited structure and semantic understanding, leading to constrained style, and semantic color inconsistency. In this paper, we introduce a sketch to color diffusion model with semanti…

Cited by 0SourceScholar
2024

KMTalk: Speech-Driven 3D Facial Animation with Key Motion Embedding

ECCV 2024poster

"We present a novel approach for synthesizing 3D facial motions from audio sequences using key motion embeddings. Despite recent advancements in data-driven techniques, accurately mapping between audio signals and 3D facial meshes remains challenging. Direct regression of the entire sequence often l…

2024

Mimic: Speaking Style Disentanglement for Speech-Driven 3D Facial Animation

AAAI 2024technical

Speech-driven 3D facial animation aims to synthesize vivid facial animations that accurately synchronize with speech and match the unique speaking style. However, existing works primarily focus on achieving precise lip synchronization while neglecting to model the subject-specific speaking style, of…

2023

Fine-Grained Retrieval Prompt Tuning

AAAI 2023technical

Fine-grained object retrieval aims to learn discriminative representation to retrieve visually similar objects. However, existing top-performing works usually impose pairwise similarities on the semantic embedding spaces or design a localization sub-network to continually fine-tune the entire model…

Cited by 21SourcePDFScholar
2023

Learning to Parameterize Visual Attributes for Open-set Fine-grained Retrieval

NeurIPS 2023poster

Open-set fine-grained retrieval is an emerging challenging task that allows to retrieve unknown categories beyond the training set. The best solution for handling unknown categories is to represent them using a set of visual attributes learnt from known categories, as widely used in zero-shot learn…

Cited by 11SourcePDFScholar
2023

Open-Set Fine-Grained Retrieval via Prompting Vision-Language Evaluator

CVPR 2023poster

Open-set fine-grained retrieval is an emerging challenge that requires an extra capability to retrieve unknown subcategories during evaluation. However, current works are rooted in the close-set scenarios, where all the subcategories are pre-defined, and make it hard to capture discriminative knowle…

Cited by 22SourcePDFScholar
2023

Towards Fair and Comprehensive Comparisons for Image-Based 3D Object Detection

ICCV 2023poster

In this work, we build a modular-designed codebase, formulate strong training recipes, design an error diagnosis toolbox, and discuss current methods for image-based 3D object detection. Specifically, different from other highly mature tasks, e.g., 2D object detection, the community of image-based 3…

Cited by 3PDFcodeScholar
2022

Category-Specific Nuance Exploration Network for Fine-Grained Object Retrieval

AAAI 2022technical

Employing additional prior knowledge to model local features as a final fine-grained object representation has become a trend for fine-grained object retrieval (FGOR). A potential limitation of these methods is that they only focus on common parts across the dataset (e.g. head, body or even leg) by…

Cited by 14SourcePDFScholar
2022

MonoDistill: Learning Spatial Features for Monocular 3D Object Detection

ICLR 2022poster

3D object detection is a fundamental and challenging task for 3D scene understanding, and the monocular-based methods can serve as an economical alternative to the stereo-based or LiDAR-based methods. However, accurately locating objects in the 3D space from a single image is extremely difficult due…

2021

Cost Affinity Learning Network for Stereo Matching

ICASSP 2021accepted

Existing stereo matching methods mainly tend to directly aggregate features output from Convolutional Neural Network to obtain more discriminative cost features, but ignore the affinity of each element in the cost feature which also plays a key role in enhancing the cost feature. In this work, we pr…

Cited by 0SourceScholar
2021

Delving Into Localization Errors for Monocular 3D Object Detection

CVPR 2021poster

Estimating 3D bounding boxes from monocular images is an essential component in autonomous driving, while accurate 3D object detection from this kind of data is very challenging. In this work, by intensive diagnosis experiments, we quantify the impact introduced by each sub-task and found the `local…

Cited by 272PDFcodeScholar
2021

Dynamic Position-aware Network for Fine-grained Image Recognition

AAAI 2021technical

Most weakly supervised fine-grained image recognition (WFGIR) approaches predominantly focus on learning the discriminative details which contain the visual variances and position clues. The position clues can be indirectly learnt by utilizing context information of discriminative visual content. Ho…

Cited by 35SourcePDFScholar
2021

Learning Scene Structure Guidance via Cross-Task Knowledge Transfer for Single Depth Super-Resolution

CVPR 2021poster

Existing color-guided depth super-resolution (DSR) approaches require paired RGB-D data as training examples where the RGB image is used as structural guidance to recover the degraded depth map due to their geometrical similarity. However, the paired data may be limited or expensive to be collected…

Cited by 55PDFScholar
2020

SIRI: Spatial Relation Induced Network For Spatial Description Resolution

NeurIPS 2020poster

Spatial Description Resolution, as a language-guided localization task, is proposed for target location in a panoramic street view, given corresponding language descriptions. Explicitly characterizing an object-level relationship while distilling spatial relationships are currently absent but crucia…

2020

Weakly Supervised Fine-Grained Image Classification via Guassian Mixture Model Oriented Discriminative Learning

CVPR 2020oral

Existing weakly supervised fine-grained image recognition (WFGIR) methods usually pick out the discriminative regions from the high-level feature maps directly. We discover that due to the operation of stacking local receptive filed, Convolutional Neural Network causes the discriminative region diff…

Cited by 104PDFScholar
2019

Accurate Monocular 3D Object Detection via Color-Embedded 3D Reconstruction for Autonomous Driving

ICCV 2019poster

In this paper, we propose a monocular 3D object detection framework in the domain of autonomous driving. Unlike previous image-based methods which focus on RGB feature extracted from 2D images, our method solves this problem in the reconstructed 3D space in order to exploit 3D contexts explicitly. T…

Cited by 396PDFScholar
2018

Deep Layer Prior Optimization for Single Image Rain Streaks Removal

ICASSP 2018accepted

Visible distortions caused by rain streaks have significant negative effects on the performance of many vision and learning algorithms. Most of the existing deraining approaches propose to build complex prior models to formulate the appearance of rain streaks. Unfortunately, these human-designed pri…

Cited by 0SourceScholar
2018

Depth Super-Resolution with Deep Edge-Inference Network and Edge-Guided Depth Filling

ICASSP 2018accepted

In this paper, we propose a novel depth super-resolution framework with deep edge-inference network and edge-guided depth filling. We first construct a convolutional neural network (CNN) architecture to learn a binary map of depth edge location from low resolution depth map and corresponding color i…

Cited by 0SourceScholar
2018

Robust Haze Removal Via Joint Deep Transmission and Scene Propagation

ICASSP 2018accepted

Haze is one of the most important factors which reduce the outdoor image quality. Existing approaches often aim to design their models based on principles of hazes. However, even with exactly modeled haze distribution, it is still a challenging task due to factors in real scenario, such as noises, h…

Cited by 0SourceScholar