← Search

Nan Li

29 accepted papers

2026

EgoRoC: Towards Egocentric Robotic Control via Task-Agnostic Visual Alignment

CVPR 2026

Recent Vision-Language-Action (VLA) models map visual-textual inputs to robotic actions via end-to-end architectures, yet this approach entangles visual understanding with task-specific actions. This leads to an exhaustive collection of full operational sequences and parameter redundancy across task

Cited by 0SourceScholar
2026

No Retraining at Edge: Efficient Resource-Aware Mixed-Precision Quantization via Federated Supernet Learning

ICML 2026poster

Federated learning (FL) enables collaborative training across distributed edge devices, but deploying lightweight models in dynamic edge environments remains challenging. Existing methods typically require retraining whenever device resource constraints change, resulting in excessive computational o…

Cited by 0SourceScholar
2026

UV-RGS: Relightable 3D Gaussian Splatting from Unposed Views Under Varied Illuminations

AAAI 2026technical

The latest advancements in scene relighting have been predominantly driven by inverse rendering with 3D Gaussian Splatting (3DGS). However, existing methods remain overly reliant on precise camera parameters under static illumination conditions, which is prohibitively expensive and even impractical

Cited by 0SourcePDFScholar
2026

Write Where It Matters: Policy-Guided Watermarks for 3D Gaussian Splatting

CVPR 2026

Recent advances in 3D Gaussian Splatting (3DGS) enable photorealistic real-time rendering but also increase the risks of unauthorized copying and redistribution. Existing 3DGS watermarking methods typically rely on handcrafted thresholds or globally fixed hyperparameters to balance invisibility and

Cited by 0SourceScholar
2025

A Visual Servo System for Robotic on-Orbit Servicing Based on 3D Perception of Non-Cooperative Satellite

ICRA 2025

The 3D perception of satellites, including both their shape and pose, is a key foundation for robotic on-orbit servicing. However, the demanding space environment-such as intense and dim illumination-presents significant challenges. Previous non-cooperative methods focus on specific geometric featur

Cited by 0SourceScholar
2025

DASA-Trans-STM: Adaptive Efficient Transformer for Short Text Matching using Data Augmentation and Semantic Awareness

EMNLP 2025

Rencent advancements in large language models (LLM) have shown impressive versatility across various tasks. Short text matching is one of the fundamental technologies in natural language processing. In previous studies, the common approach to applying them to Chinese is segmenting each sentence into

Cited by 0SourcePDFScholar
2025

Estimation of Slip Ratio and Side Slip Angle of Wheeled Planetary Rovers Based on Trace Imprint

RA-L 2025

This paper proposes a method to estimate the wheel slip ratio and side slip angle of wheeled rovers by processing images of wheel trace imprints. The proposed method extracts structural features from trace imprint images, such as the trace unit, trace contour, and angle between the centerline of the

Cited by 2SourceScholar
2025

SU-RGS: Relightable 3D Gaussian Splatting from Sparse Views under Unconstrained Illuminations

ICCV 2025poster

The latest advancements in scene relighting have been predominantly driven by inverse rendering with 3D Gaussian Splatting (3DGS). However, existing methods remain overly reliant on densely sampled images under static illumination conditions, which is prohibitively expensive and even impractical in…

Cited by 0SourcePDFScholar
2025

Transferable Relativistic Predictor: Mitigating Cross-Task Cold-Start Issue in NAS

IJCAI 2025

In neural architecture search (NAS), the relativistic predictor has recently emerged as an attractive technique to solve ranking issue for performance evaluation by predicting the relativistic ranking of architecture pair rather than the absolute performance of an architecture. However, it suffers f

Cited by 0SourcePDFScholar
2024

BAE-Net: a Low Complexity and High Fidelity Bandwidth-Adaptive Neural Network for Speech Super-Resolution

ICASSP 2024accepted

Speech bandwidth extension (BWE) has demonstrated promising performance in enhancing the perceptual speech quality in real communication systems. Most existing BWE researches primarily focus on fixed upsampling ratios, disregarding the fact that the effective bandwidth of captured audio may fluctuat…

Cited by 0SourceScholar
2024

Implicit Coarse-to-Fine 3D Perception for Category-level Object Pose Estimation from Monocular RGB Image

ICRA 2024poster

Category-level object pose estimation demonstrates robust generalization capabilities that benefit robotics applications. However, exclusive reliance on RGB images without leveraging any 3D information introduces ambiguity in the translation and size of objects, leading to suboptimal performance. In…

Cited by 0SourceScholar
2024

SVAD: A Robust, Low-Power, and Light-Weight Voice Activity Detection with Spiking Neural Networks

ICASSP 2024accepted

Speech applications are expected to be low-power and robust under noisy conditions. An effective Voice Activity Detection (VAD) front-end lowers the computational need. Spiking Neural Networks (SNNs) are known to be biologically plausible and power-efficient. However, SNN-based VADs have yet to achi…

Cited by 0SourceScholar
2023

A Low-Latency Deep Hierarchical Fusion Network for Fullband Acoustic Echo Cancellation

ICASSP 2023accepted

This paper describes our submission to the fourth Acoustic Echo Cancellation (AEC) Challenge, which is part of ICASSP 2023 Signal Processing Grand Challenge. The proposed system is developed based on our earlier system submitted to the ICASSP 2022 AEC challenge with significant latency and network s…

Cited by 3SourceScholar
2023

A Model-Based Hearing Compensation Method Using a Self-Supervised Framework

ICASSP 2023accepted

Hearing aids can improve auditory perception for hearing-impaired (HI) listeners, but even state-of-art devices provide only limited benefits if not configured correctly for the listeners. The prescriptive fittings of hearing aids ignore the individual difference among HI listeners with identical he…

Cited by 0SourceScholar
2023

LADA-Trans-NER: Adaptive Efficient Transformer for Chinese Named Entity Recognition Using Lexicon-Attention and Data-Augmentation

AAAI 2023technical

Recently, word enhancement has become very popular for Chinese Named Entity Recognition (NER), reducing segmentation errors and increasing the semantic and boundary information of Chinese words. However, these methods tend to ignore the semantic relationship before and after the sentence after integ…

Cited by 10SourcePDFScholar
2023

Speech and Noise Dual-Stream Spectrogram Refine Network With Speech Distortion Loss For Robust Speech Recognition

ICASSP 2023accepted

In recent years, the joint training of speech enhancement front-end and automatic speech recognition (ASR) back-end has been widely used to improve the robustness of ASR systems. Traditional joint training methods only use enhanced speech as input for the backend. However, it is difficult for speech…

Cited by 0SourceScholar
2022

A Deep Hierarchical Fusion Network for Fullband Acoustic Echo Cancellation

ICASSP 2022accepted

Deep learning based wideband (16kHz) acoustic echo cancellation (AEC) approaches have surpassed traditional methods. This work proposes a deep hierarchical fusion (DHF) network with intra-network and inter-network fusion to further improve the wideband AEC performance. Meanwhile, this work extends t…

Cited by 0SourceScholar
2022

Cost-Effective Sensing for Goal Inference: A Model Predictive Approach

ICRA 2022poster

Goal inference is of great importance for a variety of applications that involve interaction, coordination, and/or competition with goal-oriented agents. Typical goal inference approaches use as many pointwise measurements of the agent's trajectory as possible to pursue a most accurate a-posteriori…

Cited by 0SourceScholar
2021

IIAS: An Intelligent Insurance Assessment System through Online Real-time Conversation Analysis

IJCAI 2021poster

With the development of Chinese medical insurance industry, the amount of claim cases is growing rapidly. Ultimately, more claims necessarily indicate that the insurance company has to spend much time assessing claims and decides how much compensation the claimant should receive, which is a highly p…

2021

Robust Voice Activity Detection Using a Masked Auditory Encoder Based Convolutional Neural Network

ICASSP 2021accepted

Voice activity detection (VAD) based on deep learning has achieved remarkable success. However, when the traditional features (e.g., raw waveforms and MFCCs) are directly fed to the deep neural network model, the performance decreases because of noise interference. Here, we propose a robust VAD appr…

Cited by 0SourceScholar
2020

End-to-End Learnable Geometric Vision by Backpropagating PnP Optimization

CVPR 2020poster

Deep networks excel in learning patterns from large amounts of data. On the other hand, many geometric vision tasks are specified as optimization problems. To seamlessly combine deep learning and geometric vision, it is vital to perform learning and geometric optimization end-to-end. Towards this ai…

Cited by 128PDFcodeScholar
2019

Mapping for Planetary Rovers from Terramechanics Perspective

IROS 2019poster

In an autonomous scientific exploration system, the terrain map generated from mapping process integrates sensing information from multiple aspects and lays the base for decision making processes. With the increasing challenges in planetary exploration, equipping planetary rovers with the principles…

Cited by 13SourceScholar