← Search

Yue Li

56 accepted papers

2026

Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding

CVPR 2026

While 3DGS has emerged as a high-fidelity scene representation, encoding rich, general-purpose features directly from its primitives remains under-explored. We address this gap by introducing Chorus, a multi-teacher pretraining framework that learns a holistic feed-forward 3D Gaussian Splatting (3DG

Cited by 0SourcecodeScholar
2026

Drive-R1: Bridging Reasoning and Planning in VLMs for Autonomous Driving with Reinforcement Learning

AAAI 2026technical

Large vision-language models (VLMs) for autonomous driving (AD) are evolving beyond perception and cognition tasks toward motion planning. However, we identify two critical challenges in this direction: (1) VLMs tend to learn shortcuts by relying heavily on history input information, achieving seemi

Cited by 0SourcePDFScholar
2026

ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency Scenarios

ICML 2026oral

Speculative Decodin promises to accelerate Large Language Model inference, yet its efficacy often degrades in production-grade scenarios. Existing evaluations typically overlook the compute-bound nature of high-concurrency regimes, where verification compute becomes the dominant bottleneck. Conseque…

Cited by 0SourceScholar
2026

FedCARE: Federated Unlearning with Conflict-Aware Projection and Relearning-Resistant Recovery

IJCAI 2026

Federated learning (FL) enables collaborative model training without centralizing raw data, but privacy regulations such as the right to be forgotten require FL systems to remove the influence of previously used training data upon request. Retraining a federated model from scratch is prohibitively e

Cited by 0Scholar
2026

GS-Playground: A High-Throughput Photorealistic Simulator for Vision-Informed Robot Learning

RSS 2026poster

Embodied AI research is undergoing a shift toward vision-centric perceptual paradigms. While massively parallel simulators have catalyzed breakthroughs in proprioception-based locomotion, their potential remains largely untapped for vision-centric tasks due to the prohibitive computational overhead …

Cited by 0SourceScholar
2026

Modeling Rapid Contextual Learning in the Visual Cortex with Fast-Weight Deep Autoencoder Networks

AAAI 2026technical

Recent neurophysiological studies have revealed that the early visual cortex can rapidly learn global image context, as evidenced by a sparsification of population responses and a reduction in mean activity when exposed to familiar versus novel image contexts. This phenomenon has been attributed pri

Cited by 0SourcePDFScholar
2026

Synthetic Forgetting Without Access: A Few-Shot Zero-Glance Framework for Machine Unlearning

AAAI 2026technical

Machine unlearning aims to eliminate the influence of specific data from trained models to ensure privacy compliance. However, most existing methods assume full access to the original training dataset, which is often impractical. We address a more realistic yet challenging setting: few-shot zero-gla

Cited by 0SourcePDFScholar
2026

Taming the Long Tail: Rebalancing Adversarial Training via Adaptive Perturbation

CVPR 2026

Deep neural networks are highly vulnerable to adversarial examples, i.e.,small perturbations that can significantly degrade model performance. While adversarial training has become the primary defense strategy, most studies focus on balanced datasets, overlooking the challenges posed by real-world l

Cited by 0SourcecodeScholar
2025

A Quality-Aware Sampling Framework for Efficient 3D Point Cloud Transmission

ICASSP 2025accepted

The large volume of data from the point cloud brings significant demands on network bandwidth. However, the current transmission framework only considers using lossy compression to control the size of data, while ignoring visually redundant information due to the setting of rendering devices. Based…

Cited by 0SourceScholar
2025

Can Real-Time Lipreading Improve Speech Recognition? A Systematic Exploration Using Human-Robot Interaction Data

IROS 2025

Speech recognition in Human-Robot Interaction (HRI) fully relies on audio-based Automatic Speech Recognition. However, speech recognition that relies solely on audio faces significant challenges in noisy environments and may lead to poor performance in such environments. One approach to address this

Cited by 0SourceScholar
2025

DISCOVERSE: Efficient Robot Simulation in Complex High-Fidelity Environments

IROS 2025

We present Discoverse, the first unified, modular, open-source 3DGS-based simulation framework for Real2Sim2Real robot learning. It features a holistic Real2Sim pipeline that synthesizes hyper-realistic geometry and appearance of complex real-world scenarios, paving the way for analyzing and bridgin

Cited by 14SourcecodeScholar
2025

Dynamic Dictionary Learning for Remote Sensing Image Segmentation

ICCV 2025poster

Remote sensing image segmentation faces persistent challenges in distinguishing morphologically similar categories and adapting to diverse scene variations. While existing methods rely on implicit representation learning paradigms, they often fail to dynamically adjust semantic embeddings according…

2025

Exploiting Robust Model Watermarking Against the Model Fine-Tuning Attack via Flat Minima Aware Optimizers

ICASSP 2025accepted

With the rapid advancement of deep neural networks (DNNs), model watermarking has emerged as a widely adopted technique for safeguarding model copyrights. A prevalent method involves utilizing a watermark decoder to retrieve watermark bits from generated outputs, but such methods are often vulnerabl…

Cited by 0SourceScholar
2025

Fine-Grained Evaluation of Large Vision-Language Models in Autonomous Driving

ICCV 2025poster

Existing benchmarks for Vision-Language Model (VLM) in autonomous driving (AD) primarily assess interpretability through open-form visual question answering (QA) within coarse-grained tasks, which remain insufficient to assess capabilities in complex driving scenarios. To this end, we introduce VLAD…

2025

Generalizable Non-Line-of-Sight Imaging with Learnable Physical Priors

ICCV 2025poster

Non-line-of-sight (NLOS) imaging, recovering the hidden volume from indirect reflections, has attracted increasing attention due to its potential applications. Despite promising results, existing NLOS reconstruction approaches are constrained by the reliance on empirical physical priors, e.g., singl…

2025

Hierarchical Safety Realignment: Lightweight Restoration of Safety in Pruned Large Vision-Language Models

ACL 2025finding

With the increasing size of Large Vision-Language Models (LVLMs), network pruning techniques aimed at compressing models for deployment in resource-constrained environments have garnered significant attention. However, we observe that pruning often leads to a degradation in safety performance. To ad…

2025

It’s All About In-Context Learning! Teaching Extremely Low-Resource Languages to LLMs

EMNLP 2025

Extremely low-resource languages, especially those written in rare scripts, remain largely unsupported by large language models (LLMs). This is due in part to compounding factors such as the lack of training data. This paper delivers the first comprehensive analysis of whether LLMs can acquire such

2025

Label Set Optimization via Activation Distribution Kurtosis for Zero-Shot Classification with Generative Models

EMNLP 2025

In-context learning (ICL) performance is highly sensitive to prompt design, yet the impact of class label options (e.g. lexicon or order) in zero-shot classification remains underexplored. This study proposes LOADS (Label set Optimization via Activation Distribution kurtosiS), a post-hoc method for

Cited by 0SourcePDFScholar
2025

Rethinking Cancer Gene Identification Through Graph Anomaly Analysis

AAAI 2025technical

Graph neural networks (GNNs) have shown promise in integrating protein-protein interaction (PPI) networks for identifying cancer genes in recent studies. However, due to the insufficient modeling of the biological information in PPI networks, more faithfully depiction of complex protein interaction…

2025

SceneSplat++: A Large Dataset and Comprehensive Benchmark for Language Gaussian Splatting

NeurIPS 2025poster

3D Gaussian Splatting (3DGS) serves as a highly performant and efficient encoding of scene geometry, appearance, and semantics. Moreover, grounding language in 3D scenes has proven to be an effective strategy for 3D scene understanding. Current Language Gaussian Splatting line of work fall into thre…

Cited by 0SourceScholar
2025

SceneSplat: Gaussian Splatting-based Scene Understanding with Vision-Language Pretraining

ICCV 2025poster

Recognizing arbitrary or previously unseen categories is essential for comprehensive real-world 3D scene understanding. Currently, all existing methods rely on 2D or textual modalities during training, or together at inference. This highlights a clear absence of a model capable of processing 3D data…

2025

ShortcutsBench: A Large-Scale Real-world Benchmark for API-based Agents

ICLR 2025poster

Recent advancements in integrating large language models (LLMs) with application programming interfaces (APIs) have gained significant interest in both academia and industry. Recent work demonstrates that these API-based agents exhibit relatively strong autonomy and planning capabilities. However, t…

2024

Can We Identify Stance without Target Arguments? A Study for Rumour Stance Classification

COLING 2024main

Considering a conversation thread, rumour stance classification aims to identify the opinion (e.g. agree or disagree) of replies towards a target (rumour story). Although the target is expected to be an essential component in traditional stance classification, we show that rumour stance classificati…

2024

Cell ontology guided transcriptome foundation model

NeurIPS 2024spotlight

Transcriptome foundation models (TFMs) hold great promises of deciphering the transcriptomic language that dictate diverse cell functions by self-supervised learning on large-scale single-cell gene expression data, and ultimately unraveling the complex mechanisms of human diseases. However, current…

2024

Event-assisted Low-Light Video Object Segmentation

CVPR 2024poster

In the realm of video object segmentation (VOS) the challenge of operating under low-light conditions persists resulting in notably degraded image quality and compromised accuracy when comparing query and memory frames for similarity computation. Event cameras characterized by their high dynamic ran…

2024

FewViewGS: Gaussian Splatting with Few View Matching and Multi-stage Training

NeurIPS 2024poster

The field of novel view synthesis from images has seen rapid advancements with the introduction of Neural Radiance Fields (NeRF) and more recently with 3D Gaussian Splatting. Gaussian Splatting became widely adopted due to its efficiency and ability to render novel views accurately. While Gaussian S…

Cited by 2SourcePDFScholar
2024

High-Resolution and Few-shot View Synthesis from Asymmetric Dual-lens Inputs

ECCV 2024poster

"Novel view synthesis has achieved remarkable quality and efficiency by the paradigm of 3D Gaussian Splatting (3D-GS), but still faces two challenges: 1) significant performance degradation when trained with only few-shot samples due to a lack of geometry constraint, and 2) incapability of rendering…

2024

Optical-Waveguide Based 3-Axial Tactile Sensor for Minimally Invasive Surgical Instruments

RA-L 2024

Force feedback is of importance in Minimally Invasive Surgery (MIS) as it reduces surgical risks and enhances surgical safety. However, equipping force sensing to the tip of surgical instruments presents challenges due to their diminutive dimensions and often curved shapes. To address this issue, a

Cited by 4SourceScholar
2024

Revealing Hierarchical Structure of Leaf Venations in Plant Science via Label-Efficient Segmentation: Dataset and Method

IJCAI 2024poster

Hierarchical leaf vein segmentation is a crucial but under-explored task in agricultural sciences, where analysis of the hierarchical structure of plant leaf venation can contribute to plant breeding. While current segmentation techniques rely on data-driven models, there is no publicly available da…

2024

Toward Dynamic Non-Line-of-Sight Imaging with Mamba Enforced Temporal Consistency

NeurIPS 2024poster

Dynamic reconstruction in confocal non-line-of-sight imaging encounters great challenges since the dense raster-scanning manner limits the practical frame rate. A fewer pioneer works reconstruct high-resolution volumes from the under-scanning transient measurements but overlook temporal consistency…

2023

Characterisation of Antagonistically Actuated, Stiffness-Controllable Joint-Link Units for Cobots

ICRA 2023poster

Soft robotic structures may play a major role in the 4th industrial revolution. Researchers have successfully demonstrated the advantages of soft robotics over traditional robots made of rigid links and joints in many application areas. Variable stiffness links (VSL) and joints (VSJ) have been inves…

Cited by 1SourceScholar
2023

Deep Non-line-of-sight Imaging from Under-scanning Measurements

NeurIPS 2023poster

Active confocal non-line-of-sight (NLOS) imaging has successfully enabled seeing around corners relying on high-quality transient measurements. However, acquiring spatial-dense transient measurement is time-consuming, raising the question of how to reconstruct satisfactory results from under-scannin…

2023

Distance-Based Weight Transfer for Fine-Tuning From Near-Field to Far-Field Speaker Verification

ICASSP 2023accepted

The scarcity of labeled far-field speech is a constraint for training superior far-field speaker verification systems. In general, fine-tuning the model pre-trained on large-scale near- field speech through a small amount of far-field speech substantially outperforms training from scratch. However,…

Cited by 0SourceScholar
2023

DocTrack: A Visually-Rich Document Dataset Really Aligned with Human Eye Movement for Machine Reading

EMNLP 2023long findings

The use of visually-rich documents in various fields has created a demand for Document AI models that can read and comprehend documents like humans, which requires the overcoming of technical, linguistic, and cognitive barriers. Unfortunately, the lack of appropriate datasets has significantly hinde…

Cited by 0SourcecodeScholar
2023

Don't waste a single annotation: improving single-label classifiers through soft labels

EMNLP 2023short findings

In this paper, we address the limitations of the common data annotation and training methods for objective single-label classification tasks. Typically, when annotating such tasks annotators are only asked to provide a single label for each sample and annotator disagreement is discarded when a final…

Cited by 0SourceScholar
2023

NLOST: Non-Line-of-Sight Imaging With Transformer

CVPR 2023poster

Time-resolved non-line-of-sight (NLOS) imaging is based on the multi-bounce indirect reflections from the hidden objects for 3D sensing. Reconstruction from NLOS measurements remains challenging especially for complicated scenes. To boost the performance, we present NLOST, the first transformer-base…

Cited by 29SourcePDFScholar
2023

Swarm Robotics Search and Rescue: A Bee-Inspired Swarm Cooperation Approach without Information Exchange

ICRA 2023poster

Swarm robotics plays a non-negligible role in actual practice because of its scalability and robustness. Besides some specific studies, there is still a lack of overall approaches to solving the search and rescue problem in a communication-denied environment. This paper presents a bee-inspired swarm…

Cited by 3SourceScholar
2022

Is Discourse Role Important for Emotion Recognition in Conversation?

AAAI 2022technical

A conversation is a sequence of utterances, where each utterance plays a specific discourse role while expressing a particular emotion. This paper proposes a novel method to exploit latent discourse role information of an utterance to determine the emotion it conveys in a conversation. Specifically,…

Cited by 30SourcePDFScholar
2022

Polymer-Based Optical Waveguide Triaxial Tactile Sensing for 3-Dimensional Curved Shell

RA-L 2022

To realize dexterous robotic manipulation and enhance the human-machine interaction, nowadays increasing efforts have been made towards multi-dimensional force sensing. However, there are still bottlenecks in integrating these sensors into robots because of the limitation on conformability to comple

Cited by 17SourceScholar
2021

NTopo: Mesh-free Topology Optimization using Implicit Neural Representations

NeurIPS 2021poster

Recent advances in implicit neural representations show great promise when it comes to generating numerical solutions to partial differential equations. Compared to conventional alternatives, such representations employ parameterized neural networks to define, in a mesh-free manner, signals that are…

Cited by 86SourcePDFScholar
2021

TRQ: Ternary Neural Networks With Residual Quantization

AAAI 2021technical

Ternary neural networks (TNNs) are potential for network acceleration by reducing the full-precision weights in network to ternary ones, e.g., {-1,0,1}. However, existing TNNs are mostly calculated based on rule-of-thumb quantization methods by simply thresholding operations, which causes a signifi…

Cited by 35SourcePDFScholar
2020

Randomized tests for high-dimensional regression: A more efficient and powerful solution

NeurIPS 2020poster

We investigate the problem of testing the global null in the high-dimensional regression models when the feature dimension $p$ grows proportionally to the number of observations $n$. Despite a number of prior work studying this problem, whether there exists a test that is model-agnostic, efficient t…

Cited by 1SourcePDFScholar
2019

Gradient Image Super-resolution for Low-resolution Image Recognition

ICASSP 2019accepted

In visual object recognition problems essential to surveillance and navigation problems in a variety of military and civilian use cases, low-resolution and low-quality images present great challenges to this problem. Recent advancements in deep learning based methods like EDSR/VDSR have boosted pixe…

Cited by 0SourceScholar
2018

A Fluid-Filled Tubular Dielectric Elastomer Variable Stiffness Structure Inspired by the Hydrostatic Skeleton Principle

ICRA 2018poster

This work presents a novel variable stiffness structure consisting of a fiber-constrained dielectric elastomer tube filled with insulating oil. The tensile stiffness of the structure can be adjusted by voltages and its initial value can be customized according to the initial pre-stretch of the mater…

Cited by 3SourceScholar
2018

A Fluid-Filled Tubular Dielectric Elastomer Variable Stiffness Structure Inspired by the Hydrostatic Skeleton Principle *Research supported by the National Natural Science Foundation of China (No.51675413)

ICRA 2018

This work presents a novel variable stiffness structure consisting of a fiber-constrained dielectric elastomer tube filled with insulating oil. The tensile stiffness of the structure can be adjusted by voltages and its initial value can be customized according to the initial pre-stretch of the mater

Cited by 15SourceScholar
2017

Design and control of an inchworm-inspired soft robot with omega-arching locomotion

ICRA 2017poster

This paper presents an inchworm inspired soft robot composed of the soft body, the front foot as well as the back foot. Compared to the traditional inchworm-type robot consisting of rigid components, the driven mode for the soft robot is more simple. The soft robot inspired by the inchworm has highe…

Cited by 46SourceScholar