← Search

Lingling Li

21 accepted papers

2026

DSGCR: Decomposed Spectral Geometry-Aware Cross-Modal Semantic Representation for 3D Visual Grounding

ICML 2026poster

3D visual grounding encompassing 3D referring expression comprehension (3DREC) and segmentation (3DRES) requires robust cross-modal representation to achieve fine-grained semantic alignment and precise geometric reasoning. However, most methods employ unimodal pre-trained encoders that transfer visu…

Cited by 0SourceScholar
2026

Delving Aleatoric Uncertainty in Medical Image Segmentation via Vision Foundation Models

CVPR 2026

Medical image segmentation supports clinical workflows by precisely delineating anatomical structures and lesions. However, medical image datasets medical image datasets suffer from acquisition noise and annotation ambiguity, causing pervasive data uncertainty that substantially undermines model rob

Cited by 0SourceScholar
2026

Disentangled Hypergraph-Guided Mamba Scanning for Fine-Grained Visual Recognition

AAAI 2026technical

Fine-grained Visual Recognition (FGVR) aims to distinguish between categories with subtle inter-class differences and large intra-class variations. While Vision Transformers with attention mechanisms have been widely adopted for FGVR, they usually suffer from high computational complexity and entang

Cited by 0SourcePDFScholar
2026

Evolving Semantic Propagation for Aerial Semantic 3D Gaussian Splatting

AAAI 2026technical

Semantic understanding of large-scale aerial scenes represents a critical challenge in 3D computer vision, hindered by the prohibitive cost of dense annotation. This paper introduces EvoPropGS, a novel approach for the semantic segmentation of 3D Gaussian Splatting models that requires only minimal

Cited by 0SourcePDFScholar
2026

Foreground-Aware Token Routing Vision Transformer for Real-Time Satellite Video Tracking

ICML 2026poster

Real-time satellite video tracking poses distinct challenges, including accommodating high spatial-temporal resolution, dynamic backgrounds, and constrained onboard computational resources. While Discriminative Correlation Filter (DCF)-based methods offer high-speed inference, they suffer from limit…

Cited by 0SourceScholar
2026

HTTrack: Learning to Perceive Targets via Historical Trajectories in Satellite Video Tracking

AAAI 2026technical

In recent years, the rapid progress of deep learning has driven notable advancements in satellite video tracking, a critical task for applications such as environmental monitoring, disaster management, and defense. Despite these strides, existing approaches remain constrained by their inability to h

Cited by 0SourcePDFScholar
2026

RECS4R: Bridging Semantics and Geometry for Referring Remote Sensing Interpretation

CVPR 2026

Referring expression comprehension and segmentation (RECS) task plays a vital role in remote sensing due to its high efficiency in multi-tasking. However, RECS has reached a performance bottleneck rooted in representational insufficiency, primarily due to cross-task representational fragmentation in

Cited by 0SourcecodeScholar
2026

SOAR: Semi-Supervised Open-Vocabulary Aerial Object Detection via Dual-Aware Enhanced Prior Denoising

AAAI 2026technical

Open-Vocabulary Object Detection (OVOD) shows promise in remote sensing (RS), but due to its unique value, there are challenges such as the predominance of background regions, sparse labels, limited semantic information, and difficulties in semi-supervised training. To tackle these challenges, we pr

Cited by 0SourcePDFScholar
2026

STVG-R1: Incentivizing Instance-Level Reasoning and Grounding in Videos via Reinforcement Learning

ICLR 2026poster

In vision–language models (VLMs), misalignment between textual descriptions and visual coordinates often induces hallucinations. This issue becomes particularly severe in dense prediction tasks such as spatial–temporal video grounding (STVG). Prior approaches typically focus on enhancing visual–text…

Cited by 0SourceScholar
2026

Task-free Adaptive Meta Black-box Optimization

ICLR 2026oral

Handcrafted optimizers become prohibitively inefficient for complex black-box optimization (BBO) tasks. MetaBBO addresses this challenge by meta-learning to automatically configure optimizers for low-level BBO tasks, thereby eliminating heuristic dependencies. However, existing methods typically req…

Cited by 0SourceScholar
2026

VDFE: Difference-Aware 3D Scene Editing with Non-Intrusive Video Diffusion Priors for Multi-View Consistency and Efficiency

CVPR 2026

Text-driven 3D editing, enabled by advancements in 3D reconstruction techniques such as NeRF and 3D Gaussian Splatting, aims to provide intuitive scene customization. However, existing methods frequently exhibit limitations in controllability and consistency. To address these shortcomings, we propos

Cited by 0SourceScholar
2025

Domain-aware Category-level Geometry Learning Segmentation for 3D Point Clouds

ICCV 2025poster

Domain generalization in 3D segmentation is a critical challenge in deploying models to unseen environments. Current methods mitigate the domain shift by augmenting the data distribution of point clouds. However, the model learns global geometric patterns in point clouds while ignoring the category-…

2025

Hierarchical Variational Test-Time Prompt Generation for Zero-Shot Generalization

ICCV 2025poster

Vision-language models like CLIP have demonstrated strong zero-shot generalization, making them valuable for various downstream tasks through prompt learning. However, existing test-time prompt tuning methods, such as entropy minimization, treat both text and visual prompts as fixed learnable parame…

Cited by 0SourcePDFScholar
2025

Language-Guided Hybrid Representation Learning for Visual Grounding on Remote Sensing Images

IJCAI 2025

Visual grounding (VG) refers to detecting the specific objects in images based on linguistic expressions, and it has profound significance in the advanced interpretation of natural images. In remote sensing image interpretation, visual grounding is limited by characteristics such as the complex scen

Cited by 0SourcePDFScholar
2025

Logits DeConfusion with CLIP for Few-Shot Learning

CVPR 2025poster

With its powerful visual-language alignment capability, CLIP performs well in zero-shot and few-shot learning tasks. However, we found in experiments that CLIP's logits suffer from serious inter-class confusion problems in downstream tasks, and the ambiguity between categories seriously affects the…

2024

Gesture Generation Via Diffusion Model with Attention Mechanism

ICASSP 2024accepted

Generating natural and semantically aligned gestures from speech remains a challenging task in human-computer interaction due to the intricate relationship between speech and gestures. While recent advances in learning-based methodologies have shown progress, they exhibit limitations like limited di…

Cited by 0SourceScholar
2024

Multiplane Prior Guided Few-Shot Aerial Scene Rendering

CVPR 2024poster

Neural Radiance Fields (NeRF) have been successfully applied in various aerial scenes yet they face challenges with sparse views due to limited supervision. The acquisition of dense aerial views is often prohibitive as unmanned aerial vehicles (UAVs) may encounter constraints in perspective range an…

Cited by 3SourcePDFScholar
2024

ViLT-CLIP: Video and Language Tuning CLIP with Multimodal Prompt Learning and Scenario-Guided Optimization

AAAI 2024technical

Pre-trained vision-language(V-L) models such as CLIP have demonstrated impressive Zero-Shot performance in many downstream tasks. Since adopting contrastive video-text pairs methods like CLIP to video tasks is limited by its high cost and scale, recent approaches focus on efficiently transferring th…

Cited by 16SourcePDFScholar
2023

Curvature-Balanced Feature Manifold Learning for Long-Tailed Classification

CVPR 2023poster

To address the challenges of long-tailed classification, researchers have proposed several approaches to reduce model bias, most of which assume that classes with few samples are weak classes. However, recent studies have shown that tail classes are not always hard to learn, and model bias has been…

Cited by 57SourcePDFScholar
2021

Multi-Scale Progressive Attention Network for Video Question Answering

ACL 2021short

Understanding the multi-scale visual information in a video is essential for Video Question Answering (VideoQA). Therefore, we propose a novel Multi-Scale Progressive Attention Network (MSPAN) to achieve relational reasoning between cross-scale video information. We construct clips of different leng…

Cited by 23SourcePDFScholar