← Search

Jianan Li

30 accepted papers

2026

HyperCOD: The First Challenging Benchmark and Baseline for Hyperspectral Camouflaged Object Detection

AAAI 2026technical

RGB-based camouflaged object detection struggles in real-world scenarios where color and texture cues are ambiguous. While hyperspectral image offers a powerful alternative by capturing fine-grained spectral signatures, progress in hyperspectral camouflaged object detection (HCOD) has been criticall

Cited by 0SourcePDFScholar
2026

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning

ICASSP 2026poster

Multimodal Large Language Models (MLLMs) perform well in single-image visual grounding but struggle with real-world tasks that demand cross-image reasoning and multi-modal instructions. To address this, we adopt a reinforcement learning (RL) based post-training strategy for MLLMs in multi-image grou…

Cited by 0SourcePDFScholar
2026

Learning to Control Physically-simulated 3D Characters via Generating and Mimicking 2D Motions

CVPR 2026

Video data is more cost-effective than motion capture data for learning 3D character controllers, yet using it to generate realistic and physically plausible motions remains challenging. Previous approaches typically rely on off-the-shelf motion reconstruction techniques to extract 3D kinematic traj

Cited by 0SourcecodeScholar
2026

MODA: The First Challenging Benchmark for Multispectral Object Detection in Aerial Images

AAAI 2026technical

Aerial object detection faces significant challenges in real-world scenarios, such as small objects and extensive background interference, which limit the performance of RGB-based detectors with insufficient discriminative information. Multispectral images (MSIs) capture additional spectral cues acr

Cited by 0SourcePDFScholar
2026

OpenFly: A COMPREHENSIVE PLATFORM FOR AERIAL VISION-LANGUAGE NAVIGATION

ICLR 2026poster

Aerial Vision-Language Navigation (VLN) seeks to guide UAVs by leveraging language instructions and visual cues, establishing a new paradigm for human-UAV interaction. However, the collection of VLN data demands extensive human effort to construct trajectories and corresponding instructions, hinderi…

Cited by 0SourcecodeScholar
2026

Reasoning as an Attack Surface: Adaptive Evolutionary CoT Jailbreaks for LLMs

ICML 2026poster

Large Reasoning Models (LRMs) have demonstrated remarkable capabilities in reasoning and generation tasks and are increasingly deployed in real-world applications. However, their explicit chain-of-thought (CoT) mechanism introduces new security risks, making them particularly vulnerable to jailbreak…

Cited by 0SourceScholar
2025

Cooperative Bearing-Only Target Pursuit via Multiagent Reinforcement Learning: Design and Experiment

IROS 2025

This paper addresses the multi-robot pursuit problem for an unknown target, encompassing both target state estimation and pursuit control. First, in state estimation, we focus on using only bearing information, as it is readily available from vision sensors and effective for small, distant targets.

Cited by 2SourceScholar
2025

Dual Multi-Scale GCN with Deformable Temporal Kernel for Skeleton-based Action Recognition

ICASSP 2025accepted

Skeleton sequences for action recognition are with complex temporal dynamics due to various factors such as speed variation and different activities. It is crucial and essential to model variation changes in the temporal dimension. In recent years, skeleton sequence is always modeled as a graph stru…

Cited by 0SourceScholar
2025

FBRT-YOLO: Faster and Better for Real-Time Aerial Image Detection

AAAI 2025technical

Embedded flight devices with visual capabilities have become essential for a wide range of applications. In aerial image detection, while many existing methods have partially addressed the issue of small target detection, challenges remain in optimizing small target detection and balancing detectio…

2025

HSOD-BIT-V2: A Challenging Benchmark for Hyperspectral Salient Object Detection

AAAI 2025technical

Salient Object Detection (SOD) is crucial in computer vision, yet RGB-based methods face limitations in challenging scenes, such as small objects and similar color features. Hyperspectral images provide a promising solution for more accurate Hyperspectral Salient Object Detection (HSOD) by abundant…

Cited by 0SourcePDFScholar
2025

Knowledge Starts with Practice: Knowledge-Aware Exercise Generative Recommendation with Adaptive Multi-Agent Cooperation

NeurIPS 2025poster

Adaptive learning, which requires the in-depth understanding of students' learning processes and rational planning of learning resources, plays a crucial role in intelligent education. However, how to effectively model these two processes and seamlessly integrate them poses significant implementatio…

Cited by 0SourcecodeScholar
2025

Learning Humanoid Standing-up Control across Diverse Postures

RSS 2025poster

Standing-up control is crucial for humanoid robots, with the potential for integration into current locomotion and loco-manipulation systems. Existing approaches are either limited to simulations that neglect hardware constraints or rely on predefined ground-specific motion trajectories, failing to…

Cited by 6PDFScholar
2025

MMOT: The First Challenging Benchmark for Drone-based Multispectral Multi-Object Tracking

NeurIPS 2025poster

Drone-based multi-object tracking is essential yet highly challenging due to small targets, severe occlusions, and cluttered backgrounds. Existing RGB-based multi-object tracking algorithms heavily depend on spatial appearance cues such as color and texture, which often degrade in aerial views, comp…

Cited by 0SourcecodeScholar
2025

MUST: The First Dataset and Unified Framework for Multispectral UAV Single Object Tracking

CVPR 2025poster

UAV tracking faces significant challenges in real-world scenarios, such as small-size targets and occlusions, which limit the performance of RGB-based trackers. Multispectral images (MSI), which capture additional spectral information, offer a promising solution to these challenges. However, progres…

2025

PvNeXt: Rethinking Network Design and Temporal Motion for Point Cloud Video Recognition

ICLR 2025poster

Point cloud video perception has become an essential task for the realm of 3D vision. Current 4D representation learning techniques typically engage in iterative processing coupled with dense query operations. Although effective in capturing temporal features, this approach leads to substantial comp…

Cited by 0SourcePDFScholar
2024

Target-Guided Adversarial Point Cloud Transformer Towards Recognition Against Real-world Corruptions

NeurIPS 2024poster

Achieving robust 3D perception in the face of corrupted data presents an challenging hurdle within 3D vision research. Contemporary transformer-based point cloud recognition models, albeit advanced, tend to overfit to specific patterns, consequently undermining their robustness against corruption. I…

2024

Unified Single-Stage Transformer Network for Efficient RGB-T Tracking

IJCAI 2024poster

Most existing RGB-T tracking networks extract modality features in a separate manner, which lacks interaction and mutual guidance between modalities. This limits the network's ability to adapt to the diverse dual-modality appearances of targets and the dynamic relationships between the modalities. A…

2023

Rethinking Few-Shot Medical Segmentation: A Vector Quantization View

CVPR 2023poster

The existing few-shot medical segmentation networks share the same practice that the more prototypes, the better performance. This phenomenon can be theoretically interpreted in Vector Quantization (VQ) view: the more prototypes, the more clusters are separated from pixel-wise feature points distrib…

Cited by 17SourcePDFScholar
2023

Sample-adaptive Augmentation for Point Cloud Recognition Against Real-world Corruptions

ICCV 2023poster

Robust 3D perception under corruption has become an essential task for the realm of 3D vision. While current data augmentation techniques usually perform random transformations on all point cloud objects in an offline way and ignore the structure of the samples, resulting in over-or-under enhancemen…

Cited by 8PDFcodeScholar
2023

Value-Informed Skill Chaining for Policy Learning of Long-Horizon Tasks with Surgical Robot

IROS 2023poster

Reinforcement learning is still struggling with solving long-horizon surgical robot tasks which involve multiple steps over an extended duration of time due to the policy exploration challenge. Recent methods try to tackle this problem by skill chaining, in which the long-horizon task is decomposed…

Cited by 7SourcecodeScholar
2022

CAGroup3D: Class-Aware Grouping for 3D Object Detection on Point Clouds

NeurIPS 2022accept

We present a novel two-stage fully sparse convolutional 3D object detection framework, named CAGroup3D. Our proposed method first generates some high-quality 3D proposals by leveraging the class-aware local group strategy on the object surface voxels with the same semantic predictions, which conside…

2022

Delving into Sample Loss Curve to Embrace Noisy and Imbalanced Data

AAAI 2022technical

Corrupted labels and class imbalance are commonly encountered in practically collected training data, which easily leads to over-fitting of deep neural networks (DNNs). Existing approaches alleviate these issues by adopting a sample re-weighting strategy, which is to re-weight sample by designing…

2022

FH-Net: A Fast Hierarchical Network for Scene Flow Estimation on Real-World Point Clouds

ECCV 2022poster

"Estimating scene flow from real-world point clouds is a fundamental task for practical 3D vision. Previous methods often rely on deep models to first extract expensive per-point features at full resolution, and then get the flow either from complex matching mechanism or feature decoding, suffering…

2022

MsSVT: Mixed-scale Sparse Voxel Transformer for 3D Object Detection on Point Clouds

NeurIPS 2022accept

3D object detection from the LiDAR point cloud is fundamental to autonomous driving. Large-scale outdoor scenes usually feature significant variance in instance scales, thus requiring features rich in long-range and fine-grained information to support accurate detection. Recent detectors leverage th…

2021

Image-Based Visual Servoing of Rotorcrafts to Planar Visual Targets of Arbitrary Orientation

RA-L 2021

This letter for the first time extends the virtual camera image-based visual servoing (IBVS) scheme to enable an underactuated rotorcraft UAV to regulate its translational motion and heading relative to a planar visual target of arbitrary orientation. The conversion from real camera images to virtua

Cited by 30SourceScholar
2019

LayoutGAN: Generating Graphic Layouts with Wireframe Discriminators

ICLR 2019poster

Layout is important for graphic design and scene generation. We propose a novel Generative Adversarial Network, called LayoutGAN, that synthesizes layouts by modeling geometric relations of different types of 2D elements. The generator of LayoutGAN takes as input a set of randomly-placed 2D graphic…

Cited by 262SourcePDFScholar
2017

Perceptual Generative Adversarial Networks for Small Object Detection

CVPR 2017poster

Detecting small objects is notoriously challenging due to their low resolution and noisy representation. Existing object detection pipelines usually detect small objects through learning representations of all the objects at multiple scales. However, the performance gain of such ad hoc architectures…

Cited by 1052PDFScholar