← Search

Ning Li

27 accepted papers

2026

Boosting Self-Supervised Tracking with Contextual Prompts and Noise Learning

CVPR 2026

Learning robust contextual knowledge from unlabeled videos is essential for advancing self-supervised tracking. However, conventional self-supervised trackers lack effective context modeling, while existing context association methods based on non-semantic queries struggle to adapt to unlabeled trac

Cited by 0SourceScholar
2026

Exposing and Defending the Achilles' Heel of Video Mixture-of-Experts

ICLR 2026poster

Mixture-of-Experts (MoE) has demonstrated strong performance in video understanding tasks, yet its adversarial robustness remains underexplored. Existing attack methods often treat MoE as a unified architecture, overlooking the independent and collaborative weaknesses of key components such as route…

Cited by 0SourcecodeScholar
2026

MUTrack: A Memory-Aware Unified Representation Framework for Visual Tracking

AAAI 2026technical

Building a unified target representation that simultaneously achieves short-term adaptability and long-term stability is crucial for robust visual tracking. However, existing trackers typically face an inherent trade-off. Methods primarily relying on short-term appearance and motion cues achieve ra

Cited by 0SourcePDFScholar
2025

ARIG-GCN: Anatomical Relationship and Isomorphic Graph Approximation Guided Graph Convolutional Network for Automated ASPECTS Scoring on Non-Contrast CT

ICASSP 2025accepted

The Alberta Stroke Program Early CT Score (AS-PECTS) is a systematic method for assessing the extent of early ischemic changes on non-contrast CT (NCCT) of patients with acute ischemic stroke (AIS). The ASPECTS regions are anatomically and physiologically interconnected, making them suitable for ana…

Cited by 0SourceScholar
2025

Artificial Muscle: A Sarcomere-inspired Magnetic Approach

IROS 2025

Soft artificial muscle actuators have gained attention in robotics for their remote control, fast response, and high compliance. However, replicating the intricate and efficient motions of natural muscles remains a challenge. Existing designs often lack the hierarchical and anisotropic properties of

Cited by 0SourceScholar
2025

Decoupled Spatio-Temporal Consistency Learning for Self-Supervised Tracking

AAAI 2025technical

The success of visual tracking has been largely driven by datasets with manual box annotations. However, these box annotations require tremendous human effort, limiting the scale and diversity of existing tracking datasets. In this work, we present a novel Self-Supervised Tracking framework, named S…

2025

MobileUse: A Hierarchical Reflection-Driven GUI Agent for Autonomous Mobile Operation

NeurIPS 2025poster

Recent advances in Multimodal Large Language Models (MLLMs) have enabled the development of mobile agents that can understand visual inputs and follow user instructions, unlocking new possibilities for automating complex tasks on mobile devices. However, applying these models to real-world mobile sc…

Cited by 0SourcecodeScholar
2025

Robust Tracking via Mamba-based Context-aware Token Learning

AAAI 2025technical

How to make a good trade-off between performance and computational cost is crucial for a tracker. However, current famous methods typically focus on complicated and time-consuming learning that combining temporal and appearance information by input more and more images (or features). Consequently, t…

2025

Similarity-Guided Layer-Adaptive Vision Transformer for UAV Tracking

CVPR 2025poster

Vision transformers (ViTs) have emerged as a popular backbone for visual tracking. However, complete ViT architectures are too cumbersome to deploy for unmanned aerial vehicle (UAV) tracking which extremely emphasizes efficiency. In this study, we discover that many layers within lightweight ViT-bas…

2024

4DBInfer: A 4D Benchmarking Toolbox for Graph-Centric Predictive Modeling on RDBs

NeurIPS 2024poster

Given a relational database (RDB), how can we predict missing column values in some target table of interest? Although RDBs store vast amounts of rich, informative data spread across interconnected tables, the progress of predictive machine learning models as applied to such tasks arguably falls we…

2024

AACP: Aesthetics Assessment of Children’s Paintings Based on Self-Supervised Learning

AAAI 2024technical

The Aesthetics Assessment of Children's Paintings (AACP) is an important branch of the image aesthetics assessment (IAA), playing a significant role in children's education. This task presents unique challenges, such as limited available data and the requirement for evaluation metrics from multiple…

Cited by 1SourcePDFScholar
2024

Arbitrary Motion Style Transfer with Multi-condition Motion Latent Diffusion Model

CVPR 2024poster

Computer animation's quest to bridge content and style has historically been a challenging venture with previous efforts often leaning toward one at the expense of the other. This paper tackles the inherent challenge of content-style duality ensuring a harmonious fusion where the core narrative of t…

2024

Explicit Visual Prompts for Visual Object Tracking

AAAI 2024technical

How to effectively exploit spatio-temporal information is crucial to capture target appearance changes in visual tracking. However, most deep learning-based trackers mainly focus on designing a complicated appearance model or template updating strategy, while lacking the exploitation of context betw…

2024

HOIAnimator: Generating Text-prompt Human-object Animations using Novel Perceptive Diffusion Models

CVPR 2024poster

To date the quest to rapidly and effectively produce human-object interaction (HOI) animations directly from textual descriptions stands at the forefront of computer vision research. The underlying challenge demands both a discriminating interpretation of language and a comprehensive physics-centric…

Cited by 11SourcePDFScholar
2024

Towards a Novel Soft Magnetic Laparoscope for Single Incision Laparoscopic Surgery

ICRA 2024poster

In single-incision laparoscopic surgery (SILS), magnetic anchoring and guidance system (MAGS) is a promising technique to prevent clutter in the surgical workspace and provide a larger vision field. Existing camera designs mainly rely on rigid structure design, resulting in risks of losing magnetic…

Cited by 0SourceScholar
2023

Autonomous Exploration and Mapping for Mobile Robots via Cumulative Curriculum Reinforcement Learning

IROS 2023poster

Deep reinforcement learning (DRL) has been widely applied in autonomous exploration and mapping tasks, but often struggles with the challenges of sampling efficiency, poor adaptability to unknown map sizes, and slow simulation speed. To speed up convergence, we combine curriculum learning (CL) with…

Cited by 10SourcecodeScholar
2023

UMIFormer: Mining the Correlations between Similar Tokens for Multi-View 3D Reconstruction

ICCV 2023poster

In recent years, many video tasks have achieved breakthroughs by utilizing the vision transformer and establishing spatial-temporal decoupling for feature extraction. Although multi-view 3D reconstruction also faces multiple images as input, it cannot immediately inherit their success due to complet…

Cited by 14PDFcodeScholar
2021

Real-time 3D-Lidar, MMW Radar and GPS/IMU fusion based vehicle detection and tracking in unstructured environment

ICRA 2021poster

To solve the problem of unmanned ground vehicle leader-follower formation transportation in unstructured environment, we propose a novel target detection and tracking method based on multi-sensor fusion perception. Combined with 3D-Lidar, millimeter wave Radar and GPS/IMU, the proposed method can ac…

Cited by 15SourceScholar
2021

Recovering Stress Distribution on Deformable Tissue for a Magnetic Actuated Insertable Laparoscopic Surgical Camera

ICRA 2021poster

Fully insertable laparoscopic cameras represent a promising future of minimally invasive surgery. The most characteristic technology adopted on these devices is transabdominal anchoring and actuation based on magnetic coupling. However, few have paid adequate attention to the safety concerns. As the…

Cited by 5SourceScholar
2020

Full-Time Monocular Road Detection Using Zero-Distribution Prior of Angle of Polarization

ECCV 2020poster

This paper presents a road detection technique based on long-wave infrared (LWIR) polarization imaging for autonomous navigation regardless of illumination conditions, day and night. Division of Focal Plane (DoFP) imaging technology enables acquisition of infrared polarization images in real time us…

2019

A Noninvasive Approach to Recovering the Lost Force Feedback for a Robotic-Assisted Insertable Laparoscopic Surgical Camera

ICRA 2019poster

Fully insertable laparoscopic cameras feature more locomotive flexibility in a larger workspace compared to conventional trocar-based laparoscopes and thus represent a promising future of minimally invasive surgery. These cameras are principally anchored and actuated by transabdominal magnetic coupl…

Cited by 6SourceScholar
2018

Self-Supervised Adversarial Hashing Networks for Cross-Modal Retrieval

CVPR 2018poster

Thanks to the success of deep learning, cross-modal retrieval has made significant progress recently. However, there still remains a crucial bottleneck: how to bridge the modality gap to further enhance the retrieval accuracy. In this paper, we propose a self-supervised adversarial hashing (SSAH) ap…

2017

A novel laparoscopic camera robot with in-vivo lens cleaning and debris prevention modules

IROS 2017poster

Robotic systems have recently drawn attention in minimally invasive surgeries due to their increased dexterity feature. A major drawback of these systems is image blurring due to lens contamination which cause imaging impairment during up to 40% of surgery time. This paper demonstrates a novel lapar…

Cited by 15SourceScholar
2016

A robust Gaussian approximate filter for nonlinear systems with heavy tailed measurement noises

ICASSP 2016accepted

The scale matrix and degrees of freedom (dof) parameter of a Student's t distribution are important for nonlinear robust inference, and it is difficult to determine exact values in practical application due to complex environments. To solve this problem, an improved robust Gaussian approximate (GA)…

Cited by 0SourceScholar