← Search

Yuxin Zheng

8 accepted papers

2026

FROM KNOWING TO DOING PRECISELY: A GENERAL SELF-CORRECTION AND TERMINATION FRAMEWORK FOR VLA MODELS

ICASSP 2026poster

While vision-language-action (VLA) models for embodied agents integrate perception, reasoning, and control, they remain constrained by two critical weaknesses: first, during grasping tasks, the action tokens generated by the language model often exhibit subtle spatial deviations from the target obje…

Cited by 0SourcePDFScholar
2025

Detail-Preserving Latent Diffusion for Stable Shadow Removal

CVPR 2025poster

Achieving high-quality shadow removal with strong generalizability is challenging in scenes with complex global illumination. Due to the limited diversity in shadow removal datasets, current methods are prone to overfitting training data, often leading to reduced performance on unseen cases. To addr…

Cited by 0SourcePDFScholar
2025

Multi-Modal Synergistic Implicit Image Enhancement for Efficient Optical Flow Estimation

CVPR 2025poster

As a fundamental visual task, optical flow estimation has widespread applications in computer vision. However, it faces significant challenges under adverse lighting conditions, where low texture and noise make accurate optical flow estimation particularly difficult.In this paper, we propose an opti…

Cited by 0SourcePDFScholar
2025

OmniSR: Shadow Removal Under Direct and Indirect Lighting

AAAI 2025technical

Shadows can originate from occlusions in both direct and indirect illumination. Although most current shadow removal research focuses on shadows caused by direct illumination, shadows from indirect illumination are often just as pervasive, particularly in indoor scenes. A significant challenge in re…

2024

ECPNet: An Enhanced Curve Perception Network for Lane Detection

ICASSP 2024accepted

Lane detection methods based on anchors have received increasing attention, but fixed-shape anchors make it difficult to model complex lane line shapes. To solve this problem, we propose an Enhanced Curve Perception Network (ECPNet). Specifically, we propose a Layer-by-layer Context Fusion (LCF) mod…

Cited by 0SourceScholar
2023

Binarized Spectral Compressive Imaging

NeurIPS 2023poster

Existing deep learning models for hyperspectral image (HSI) reconstruction achieve good performance but require powerful hardwares with enormous memory and computational resources. Consequently, these methods can hardly be deployed on resource-limited mobile devices. In this paper, we propose a nove…

2023

Hierarchical Spatiotemporal Feature Fusion Network For Video Saliency Prediction

ICASSP 2023accepted

Current video saliency prediction methods have made great progress relying on the feature extraction capability of CNN, but there are still many defects in hierarchical feature fusion, limiting the further improvement of accuracy. To address this issue, we propose a 3D convolutional Hierarchical Spa…

Cited by 0SourceScholar
2021

A Light-Weight Semantic Map for Visual Localization towards Autonomous Driving

ICRA 2021poster

Accurate localization is of crucial importance for autonomous driving tasks. Nowadays, we have seen a lot of sensor-rich vehicles (e.g. Robo-taxi) driving on the street autonomously, which rely on high-accurate sensors (e.g. Lidar and RTK GPS) and high-resolution map. However, low-cost production ca…

Cited by 128SourceScholar