← Search

Licheng Jiao

34 accepted papers

2026

DSGCR: Decomposed Spectral Geometry-Aware Cross-Modal Semantic Representation for 3D Visual Grounding

ICML 2026poster

3D visual grounding encompassing 3D referring expression comprehension (3DREC) and segmentation (3DRES) requires robust cross-modal representation to achieve fine-grained semantic alignment and precise geometric reasoning. However, most methods employ unimodal pre-trained encoders that transfer visu…

Cited by 0SourceScholar
2026

Delving Aleatoric Uncertainty in Medical Image Segmentation via Vision Foundation Models

CVPR 2026

Medical image segmentation supports clinical workflows by precisely delineating anatomical structures and lesions. However, medical image datasets medical image datasets suffer from acquisition noise and annotation ambiguity, causing pervasive data uncertainty that substantially undermines model rob

Cited by 0SourceScholar
2026

Evolving Semantic Propagation for Aerial Semantic 3D Gaussian Splatting

AAAI 2026technical

Semantic understanding of large-scale aerial scenes represents a critical challenge in 3D computer vision, hindered by the prohibitive cost of dense annotation. This paper introduces EvoPropGS, a novel approach for the semantic segmentation of 3D Gaussian Splatting models that requires only minimal

Cited by 0SourcePDFScholar
2026

Foreground-Aware Token Routing Vision Transformer for Real-Time Satellite Video Tracking

ICML 2026poster

Real-time satellite video tracking poses distinct challenges, including accommodating high spatial-temporal resolution, dynamic backgrounds, and constrained onboard computational resources. While Discriminative Correlation Filter (DCF)-based methods offer high-speed inference, they suffer from limit…

Cited by 0SourceScholar
2026

HTTrack: Learning to Perceive Targets via Historical Trajectories in Satellite Video Tracking

AAAI 2026technical

In recent years, the rapid progress of deep learning has driven notable advancements in satellite video tracking, a critical task for applications such as environmental monitoring, disaster management, and defense. Despite these strides, existing approaches remain constrained by their inability to h

Cited by 0SourcePDFScholar
2026

RECS4R: Bridging Semantics and Geometry for Referring Remote Sensing Interpretation

CVPR 2026

Referring expression comprehension and segmentation (RECS) task plays a vital role in remote sensing due to its high efficiency in multi-tasking. However, RECS has reached a performance bottleneck rooted in representational insufficiency, primarily due to cross-task representational fragmentation in

Cited by 0SourcecodeScholar
2026

SOAR: Semi-Supervised Open-Vocabulary Aerial Object Detection via Dual-Aware Enhanced Prior Denoising

AAAI 2026technical

Open-Vocabulary Object Detection (OVOD) shows promise in remote sensing (RS), but due to its unique value, there are challenges such as the predominance of background regions, sparse labels, limited semantic information, and difficulties in semi-supervised training. To tackle these challenges, we pr

Cited by 0SourcePDFScholar
2026

STVG-R1: Incentivizing Instance-Level Reasoning and Grounding in Videos via Reinforcement Learning

ICLR 2026poster

In vision–language models (VLMs), misalignment between textual descriptions and visual coordinates often induces hallucinations. This issue becomes particularly severe in dense prediction tasks such as spatial–temporal video grounding (STVG). Prior approaches typically focus on enhancing visual–text…

Cited by 0SourceScholar
2026

Task-free Adaptive Meta Black-box Optimization

ICLR 2026oral

Handcrafted optimizers become prohibitively inefficient for complex black-box optimization (BBO) tasks. MetaBBO addresses this challenge by meta-learning to automatically configure optimizers for low-level BBO tasks, thereby eliminating heuristic dependencies. However, existing methods typically req…

Cited by 0SourceScholar
2025

Cross-Rejective Open-Set SAR Image Registration

CVPR 2025poster

Synthetic Aperture Radar (SAR) image registration is an essential upstream task in geoscience applications, in which pre-detected keypoints from two images are employed as observed objects to seek matched-point pairs. In general, the registration is regarded as a typical closed-set classification, w…

2025

Domain-aware Category-level Geometry Learning Segmentation for 3D Point Clouds

ICCV 2025poster

Domain generalization in 3D segmentation is a critical challenge in deploying models to unseen environments. Current methods mitigate the domain shift by augmenting the data distribution of point clouds. However, the model learns global geometric patterns in point clouds while ignoring the category-…

2025

Hierarchical Variational Test-Time Prompt Generation for Zero-Shot Generalization

ICCV 2025poster

Vision-language models like CLIP have demonstrated strong zero-shot generalization, making them valuable for various downstream tasks through prompt learning. However, existing test-time prompt tuning methods, such as entropy minimization, treat both text and visual prompts as fixed learnable parame…

Cited by 0SourcePDFScholar
2025

Language-Guided Hybrid Representation Learning for Visual Grounding on Remote Sensing Images

IJCAI 2025

Visual grounding (VG) refers to detecting the specific objects in images based on linguistic expressions, and it has profound significance in the advanced interpretation of natural images. In remote sensing image interpretation, visual grounding is limited by characteristics such as the complex scen

Cited by 0SourcePDFScholar
2025

RegionMatch: Pixel-Region Collaboration for Semi-Supervised Semantic Segmentation in Remote Sensing Images

IJCAI 2025

Semi-supervised semantic segmentation (S4) has shown significant promise in reducing the burden of labor-intensive data annotation. However, existing methods mainly rely on pixel-level information, neglecting the strong region consistency inherent in remote sensing images (RSIs), which limits their

Cited by 0SourcePDFScholar
2025

Rethinking Multiple-Instance Learning From Feature Space to Probability Space

ICLR 2025poster

Multiple-instance learning (MIL) was initially proposed to identify key instances within a set (bag) of instances when only one bag-level label is provided. Current deep MIL models mostly solve multi-instance problem in feature space. Nevertheless, with the increasing complexity of data, we found th…

2024

Masked Angle-Aware Autoencoder for Remote Sensing Images

ECCV 2024poster

"To overcome the inherent domain gap between remote sensing (RS) images and natural images, some self-supervised representation learning methods have made promising progress. However, they have overlooked the diverse angles present in RS objects. This paper proposes the Masked Angle-Aware Autoencode…

2024

Multiplane Prior Guided Few-Shot Aerial Scene Rendering

CVPR 2024poster

Neural Radiance Fields (NeRF) have been successfully applied in various aerial scenes yet they face challenges with sparse views due to limited supervision. The acquisition of dense aerial views is often prohibitive as unmanned aerial vehicles (UAVs) may encounter constraints in perspective range an…

Cited by 3SourcePDFScholar
2024

Stable Neighbor Denoising for Source-free Domain Adaptive Segmentation

CVPR 2024poster

We study source-free unsupervised domain adaptation (SFUDA) for semantic segmentation which aims to adapt a source-trained model to the target domain without accessing the source data. Many works have been proposed to address this challenging problem among which uncertainty based self-training is a…

2024

ViLT-CLIP: Video and Language Tuning CLIP with Multimodal Prompt Learning and Scenario-Guided Optimization

AAAI 2024technical

Pre-trained vision-language(V-L) models such as CLIP have demonstrated impressive Zero-Shot performance in many downstream tasks. Since adopting contrastive video-text pairs methods like CLIP to video tasks is limited by its high cost and scale, recent approaches focus on efficiently transferring th…

Cited by 16SourcePDFScholar
2023

Curvature-Balanced Feature Manifold Learning for Long-Tailed Classification

CVPR 2023poster

To address the challenges of long-tailed classification, researchers have proposed several approaches to reduce model bias, most of which assume that classes with few samples are weak classes. However, recent studies have shown that tail classes are not always hard to learn, and model bias has been…

Cited by 57SourcePDFScholar
2023

Learning Pseudo-Relations for Cross-domain Semantic Segmentation

ICCV 2023poster

Domain adaptive semantic segmentation aims to adapt a model trained on labeled source domain to the unlabeled target domain. Self-training shows competitive potential in this field. Existing methods along this stream mainly focus on selecting reliable predictions on target data as pseudo-labels for…

Cited by 23PDFcodeScholar
2023

RGMIL: Guide Your Multiple-Instance Learning Model with Regressor

NeurIPS 2023poster

In video analysis, an important challenge is insufficient annotated data due to the rare occurrence of the critical patterns, and we need to provide discriminative frame-level representation with limited annotation in some applications. Multiple Instance Learning (MIL) is suitable for this scenario.…

2023

Towards Better Stability and Adaptability: Improve Online Self-Training for Model Adaptation in Semantic Segmentation

CVPR 2023highlight

Unsupervised domain adaptation (UDA) in semantic segmentation transfers the knowledge of the source domain to the target one to improve the adaptability of the segmentation model in the target domain. The need to access labeled source data makes UDA unable to handle adaptation scenarios involving pr…

2022

Absolute Wrong Makes Better: Boosting Weakly Supervised Object Detection via Negative Deterministic Information

IJCAI 2022poster

Weakly supervised object detection (WSOD) is a challenging task, in which image-level labels (e.g., categories of the instances in the whole image) are used to train an object detector. Many existing methods follow the standard multiple instance learning (MIL) paradigm and have achieved promising pe…

Cited by 16SourcePDFScholar
2022

Self-Training Multi-Sequence Learning with Transformer for Weakly Supervised Video Anomaly Detection

AAAI 2022technical

Weakly supervised Video Anomaly Detection (VAD) using Multi-Instance Learning (MIL) is usually based on the fact that the anomaly score of an abnormal snippet is higher than that of a normal snippet. In the beginning of training, due to the limited accuracy of the model, it is easy to select the wro…

Cited by 219SourcePDFScholar
2022

Unsupervised Few-Shot Image Classification by Learning Features into Clustering Space

ECCV 2022poster

"Most few-shot image classification methods are trained based on tasks. Usually, tasks are built on base classes with a large number of labeled images, which consumes large effort. Unsupervised few-shot image classification methods do not need labeled images, because they require tasks to be built o…

2021

Cascade Attention Fusion for Fine-Grained Image Captioning Based on Multi-Layer LSTM

ICASSP 2021accepted

The conventional visual attention-based image captioning approaches typically use image information to guide caption generation. Results from these models tend to be coarse and ignore the details in the image, such as objects, attributes and the distinguishing aspects of each image. In this paper, w…

Cited by 0SourceScholar
2021

Multi-Scale Progressive Attention Network for Video Question Answering

ACL 2021short

Understanding the multi-scale visual information in a video is essential for Video Question Answering (VideoQA). Therefore, we propose a novel Multi-Scale Progressive Attention Network (MSPAN) to achieve relational reasoning between cross-scale video information. We construct clips of different leng…

Cited by 23SourcePDFScholar
2020

AttAN: Attention Adversarial Networks for 3D Point Cloud Semantic Segmentation

IJCAI 2020poster

3D point cloud semantic segmentation has attracted wide attention with its extensive applications in autonomous driving, AR/VR, and robot sensing fields. However, in existing methods, each point in the segmentation results is predicted independently from each other. This property causes the non-cont…

Cited by 0SourcePDFScholar
2020

Supervised Edge Attention Network for Accurate Image Instance Segmentation

ECCV 2020poster

Effectively keeping boundary of the mask complete is important in instance segmentation. In this task, many works segment instance based on a bounding box from the box head, which means the quality of the detection also affects the completeness of the mask. To circumvent this issue, we propose a ful…

2019

AFD-Net: Aggregated Feature Difference Learning for Cross-Spectral Image Patch Matching

ICCV 2019oral

Image patch matching across different spectral domains is more challenging than in a single spectral domain. We consider the reason is twofold: 1. the weaker discriminative feature learned by conventional methods; 2. the significant appearance difference between two images domains. To tackle these p…

Cited by 38PDFScholar
2019

Better and Faster: Exponential Loss for Image Patch Matching

ICCV 2019poster

Recent studies on image patch matching are paying more attention on hard sample learning, because easy samples do not contribute much to the network optimization. They have proposed various hard negative sample mining strategies, but very few addressed this problem from the perspective of loss funct…

Cited by 31PDFScholar
2017

Accelerated First-order Methods for Geodesically Convex Optimization on Riemannian Manifolds

NeurIPS 2017poster

In this paper, we propose an accelerated first-order method for geodesically convex optimization, which is the generalization of the standard Nesterov's accelerated method from Euclidean space to nonlinear Riemannian space. We first derive two equations and obtain two nonlinear operators for geodesi…

Cited by 101SourcePDFScholar