← Search

Fei Wen

20 accepted papers

2026

Boosting Vision-Language-Action Finetuning with Feasible Action Neighborhood Prior

CVPR 2026

In real-world robotic manipulation, states typically admit a neighborhood of near-equivalent actions. That is, for each state, there exist a feasible action neighborhood (FAN) rather than a single correct action, within which motions yield indistinguishable progress. However, prevalent VLA training

Cited by 0SourcecodeScholar
2026

Curvature-Aware Zeroth-Order Optimization for Memory-Efficient Test-Time Adaptation

CVPR 2026

Test-time adaptation (TTA) aims to enhance the cross-domain performance of pre-trained models by adapting to unlabeled test data.While most existing TTA methods rely on backpropagation (BP) for finetuning, BP-free methods such as zeroth-order (ZO) methods are more desired in practical on-device scen

Cited by 0SourcecodeScholar
2026

Efficient Plane Segmentation in Depth Image Based on Adaptive Patch-Wise Region Growing

ICRA 2026poster

Plane segmentation algorithms are widely used in robotics, serving key roles in scenarios such as indoor localization, scene understanding, and robotic manipulation. These applications typically require real-time, precise, and robust plane segmentation processing, which presents a significant challe…

Cited by 0SourceScholar
2026

OPTIMAL TRANSPORT BASED UNSUPERVISED RESTORATION LEARNING EXPLOITING DEGRADATION SPARSITY

ICASSP 2026poster

Optimal transport (OT) has recently been shown as a promising criterion for unsupervised restoration when no explicit prior model is available. Despite its theoretical appeal, OT still significantly falls short of supervised methods on challenging tasks such as super-resolution, deraining, and dehaz…

Cited by 0SourcePDFScholar
2026

SSMG-Nav: Enhancing Lifelong Object Navigation with Semantic Skeleton Memory Graph

ICRA 2026poster

Navigating to out-of-sight targets from human instructions in unfamiliar environments is a core capability for service robots. Despite substantial progress, most approaches underutilize reusable, persistent memory, constraining performance in lifelong settings. Many are additionally limited to singl…

2026

SWITCHCODEC: ADAPTIVE RESIDUAL-EXPERT SPARSE QUANTIZATION FOR HIGH-FIDELITY NEURAL AUDIO CODING

ICASSP 2026oral

Recent neural audio compression models often rely on residual vector quantization for high-fidelity coding, but using a fixed number of per-frame codebooks is suboptimal for the wide variability of audio content-especially for signals that are either very simple or highly complex. To address this li…

Cited by 0SourcePDFScholar
2026

Zeroth-Order Forward-Only SNN Training Inspiring Neuromorphic On-Chip Learning

ICML 2026poster

The human brain is a biologically instantiated on-device neural system that integrates both learning and inference in a unified architecture, which enables rapid and flexible learning on-the-fly. This extraordinary capability is achieved through non-BP learning mechanisms, whereas BP is computationa…

Cited by 0SourceScholar
2025

A Skeleton-Based Topological Planner for Exploration in Complex Unknown Environments

ICRA 2025

The capability of autonomous exploration in complex, unknown environments is important in many robotic applications. While recent research on autonomous exploration have achieved much progress, there are still limitations, e.g., existing methods relying on greedy heuristics or optimal path planning

Cited by 5SourcecodeScholar
2025

Efficient Plane Segmentation in Depth Image Based on Adaptive Patch-Wise Region Growing

RA-L 2025

Plane segmentation algorithms are widely used in robotics, serving key roles in scenarios such as indoor localization, scene understanding, and robotic manipulation. These applications typically require real-time, precise, and robust plane segmentation processing, which presents a significant challe

Cited by 2SourceScholar
2025

Lifelong Test-Time Adaptation via Online Learning in Tracked Low-Dimensional Subspace

NeurIPS 2025poster

Test-time adaptation (TTA) aims to adapt a source model to a target domain using only test data. Existing methods predominantly rely on unsupervised entropy minimization or its variants, which suffer from degeneration, leading to trivial solutions with low-entropy but inaccurate predictions. In this…

Cited by 0SourceScholar
2025

MUZO: Leveraging Multiple Queries and Momentum for Zeroth-Order Fine-Tuning of Large Language Models

EMNLP 2025

Fine-tuning pre-trained large language models (LLMs) on downstream tasks has achieved significant success across various domains. However, as model sizes grow, traditional first-order fine-tuning algorithms incur substantial memory overhead due to the need for activation storage for back-propagation

Cited by 0SourcePDFScholar
2024

BEVGM: A Visual Place Recognition Method With Bird's Eye View Graph Matching

RA-L 2024

Visual place recognition (VPR) is an essential tool in robotics perception and navigation. Though much progress has been made recently, the performance of VPR is far from satisfactory in challenging scenarios, such as large appearance variations, reverse viewpoints, and heterogeneous data. This work

Cited by 1SourceScholar
2022

Active SLAM With Prior Topo-Metric Graph Starting At Uncertain Position

RA-L 2022

Active simultaneous localization and mapping (SLAM) is an important technique for mobile robots to autonomously explore and map an environment. This letter considers the problem of active SLAM with a prior topo-metric graph. Unlike existing works, we consider a more challenging scenario that there e

Cited by 7SourceScholar
2021

On Perceptual Lossy Compression: The Cost of Perceptual Reconstruction and An Optimal Training Framework

ICML 2021spotlight

Lossy compression algorithms are typically designed to achieve the lowest possible distortion at a given bit rate. However, recent studies show that pursuing high perceptual quality would lead to increase of the lowest achievable distortion (e.g., MSE). This paper provides nontrivial results theoret…

2019

Action Recognition Based on 3D Skeleton and RGB Frame Fusion

IROS 2019poster

Action recognition has wide applications in assisted living, health monitoring, surveillance, and human-computer interaction. In traditional action recognition methods, RGB video-based ones are effective but computationally inefficient, while skeleton-based ones are computationally efficient but do…

Cited by 40SourceScholar
2016

Robust sparse recovery for compressive sensing in impulsive noise using ℓp-norm model fitting

ICASSP 2016accepted

This work considers the robust sparse recovery problem in compressive sensing (CS) in the presence of impulsive measurement noise. We propose a robust formulation for sparse recovery using the generalized lp-norm with 0 < p < 2 as the metric for the residual error under l1-norm regularization. An al…

Cited by 0SourceScholar