← Search

Sheng Yang

30 accepted papers

2026

Don't Overthink with Pixels: Efficient Reasoning for Segmentation

ICML 2026poster

Existing reasoning segmentation approaches typically fine-tune multimodal large language models (MLLMs) using image-text pairs and corresponding mask labels. While recent efforts leverage reinforcement fine-tuning to further enhance reasoning ability, they often suffer from overthinking and produce …

Cited by 0SourceScholar
2026

GUIDE: Gaussian Unified Instance Detection for Enhanced Obstacle Perception in Autonomous Driving

AAAI 2026technical

In the realm of autonomous driving, accurately detecting surrounding obstacles is crucial for effective decision-making. Traditional methods primarily rely on 3D bounding boxes to represent these obstacles, which often fail to capture the complexity of irregularly shaped, real-world objects. To over

Cited by 0SourcePDFScholar
2026

LiDAR Prompted Spatio-Temporal Multi-View Stereo for Autonomous Driving

CVPR 2026

Accurate metric depth is critical for autonomous driving perception and simulation, yet current approaches struggle to achieve high metric accuracy, multi-view and temporal consistency, and cross-domain generalization. To address these challenges, we present DriveMVS, a novel multi-view stereo frame

Cited by 0SourcecodeScholar
2026

LiDAR-GS++: Improving LiDAR Gaussian Reconstruction via Diffusion Priors

AAAI 2026technical

Recent GS-based rendering has made significant progress for LiDAR, surpassing Neural Radiance Fields (NeRF) in both quality and speed. However, these methods exhibit artifacts in extrapolated novel view synthesis due to the incomplete reconstruction from single traversal scans. To address this limit

Cited by 0SourcePDFScholar
2026

Second-Order Bilevel Optimization with Accelerated Convergence Rates

ICML 2026poster

This paper studies second-order methods for nonconvex-strongly-convex bilevel optimization. We propose a novel fully second-order bilevel approximation method (FSBA) that achieves an iteration complexity of $\tilde{\mathcal{O}}(\epsilon^{-1.5})$ for finding the $(\epsilon, \mathcal{O}(\sqrt{\epsilon…

Cited by 0SourceScholar
2026

TIGaussian: Disentangle Gaussians for Spatial-Awared Text-Image-3D Alignment

ICLR 2026poster

While visual-language models have profoundly linked features between texts and images, the incorporation of 3D modality data, such as point clouds and 3D Gaussians, further enables pretraining for 3D-related tasks, e.g., cross-modal retrieval, zero-shot classification, and scene recognition. As chal…

Cited by 0SourcecodeScholar
2025

BIAWDiff: Enhancing Low-Light Images with Bio-Inspired Attention and Wavelet Diffusion

ICASSP 2025accepted

Low-light image enhancement aims to improve visual quality under challenging lighting conditions while preserving details and color fidelity. Existing traditional algorithms and deep learning approaches, often struggle with balancing brightness enhancement and detail preservation, leading to issues…

Cited by 0SourceScholar
2025

Industrial-Grade Sensor Simulation via Gaussian Splatting: A Modular Framework for Scalable Editing and Full-Stack Validation

IROS 2025

Sensor simulation is pivotal for scalable validation of autonomous driving systems, yet existing Neural Radiance Fields (NeRF) based methods face applicability and efficiency challenges in industrial workflows. This paper introduces a Gaussian Splatting (GS) based system to address these challenges:

Cited by 3SourceScholar
2025

Language Driven Occupancy Prediction

ICCV 2025poster

We introduce LOcc, an effective and generalizable framework for open-vocabulary occupancy (OVO) prediction. Previous approaches typically supervise the networks through coarse voxel-to-text correspondences via image features as intermediates or noisy and sparse correspondences from voxel-based model…

2025

MSANet: Mixed Spectral and Attention Network for Robust 3D Human Pose Estimation

ICASSP 2025accepted

Despite significant advances in 3D human pose estimation from a single-view video, existing methods often struggle to produce reasonable human poses when the human is heavily occluded or blurred. To address this issue, we propose a Mixed Spectral and Attention Network (MSANet) that stacks spectral a…

Cited by 0SourceScholar
2025

RGE-GS: Reward-Guided Expansive Driving Scene Reconstruction via Diffusion Priors

ICCV 2025poster

A single-pass driving clip frequently results in incomplete scanning of the road structure, making reconstructed scene expanding a critical requirement for sensor simulators to effectively regress driving actions. Although contemporary 3D Gaussian Splatting (3DGS) techniques achieve remarkable recon…

2025

RTMap: Real-Time Recursive Mapping with Change Detection and Localization

ICCV 2025poster

While recent online HD mapping methods relieve burdened offline pipelines and solve map freshness, they remain limited by perceptual inaccuracies, occlusion in dense traffic, and an inability to fuse multi-agent observations. We propose RTMap to enhance these single-traversal methods by persistently…

2025

SAM4D: Segment Anything in Camera and LiDAR Streams

ICCV 2025poster

We present SAM4D, a multi-modal and temporal foundation model designed for promptable segmentation across camera and LiDAR streams. Unified Multi-modal Positional Encoding (UMPE) is introduced to align camera and LiDAR features in a shared 3D space, enabling seamless cross-modal prompting and intera…

Cited by 0SourcePDFScholar
2024

AMPO: Automatic Multi-Branched Prompt Optimization

EMNLP 2024main

Prompt engineering is very important to enhance the performance of large language models (LLMs). When dealing with complex issues, prompt engineers tend to distill multiple patterns from examples and inject relevant solutions to optimize the prompts, achieving satisfying results. However, existing a…

Cited by 3SourcePDFScholar
2024

Not All Prompts Are Secure: A Switchable Backdoor Attack Against Pre-trained Vision Transfomers

CVPR 2024poster

Given the power of vision transformers a new learning paradigm pre-training and then prompting makes it more efficient and effective to address downstream visual recognition tasks. In this paper we identify a novel security threat towards such a paradigm from the perspective of backdoor attacks. Spe…

2024

StraGo: Harnessing Strategic Guidance for Prompt Optimization

EMNLP 2024finding

Prompt engineering is pivotal for harnessing the capabilities of large language models (LLMs) across diverse applications. While existing prompt optimization methods improve prompt effectiveness, they often lead to prompt drifting, wherein newly generated prompts canadversely impact previously succe…

2023

The Numerical Stability of Hyperbolic Representation Learning

ICML 2023poster

The hyperbolic space is widely used for representing hierarchical datasets due to its ability to embed trees with small distortion. However, this property comes at a price of numerical instability such that training hyperbolic learning models will sometimes lead to catastrophic NaN problems, encount…

2022

DIDO: Deep Inertial Quadrotor Dynamical Odometry

RA-L 2022

In this work, we propose an interoceptive-only state estimation system for a quadrotor with deep neural network processing, where the quadrotor dynamics is considered as a perceptive supplement of the inertial kinematics. To improve the precision of multi-sensor fusion, we train cascaded networks on

Cited by 25SourcecodeScholar
2022

SuperLine3D: Self-Supervised Line Segmentation and Description for LiDAR Point Cloud

ECCV 2022poster

"Poles and building edges are frequently observable objects on urban roads, conveying reliable hints for various computer vision tasks. To repetitively extract them as features and perform association between discrete LiDAR frames for registration, we propose the first learning-based feature segment…

2022

The Visual-Inertial- Dynamical Multirotor Dataset

ICRA 2022poster

Recently, the community has witnessed numerous datasets built for developing and testing state estimators. However, for some applications such as aerial transportation or search-and-rescue, the contact force or other disturbance must be perceived for robust planning and control, which is beyond the…

Cited by 8SourcecodeScholar
2021

Road Mapping and Localization Using Sparse Semantic Visual Features

RA-L 2021

We present a novel method for visual mapping and localization for autonomous vehicles, by extracting, modeling, and optimizing semantic road elements. Specifically, our method integrates cascaded deep models to detect standardized road elements instead of traditional point features, to seek for impr

Cited by 31SourceScholar
2020

ClusterVO: Clustering Moving Instances and Estimating Visual Odometry for Self and Surroundings

CVPR 2020poster

We present ClusterVO, a stereo Visual Odometry which simultaneously clusters and estimates the motion of both ego and surrounding rigid clusters/objects. Unlike previous solutions relying on batch input or imposing priors on scene structure or dynamic object models, ClusterVO is online, general and…

Cited by 123PDFScholar
2020

Learning Efficient Parameter Server Synchronization Policies for Distributed SGD

ICLR 2020poster

We apply a reinforcement learning (RL) based approach to learning optimal synchronization policies used for Parameter Server-based distributed training of machine learning models with Stochastic Gradient Descent (SGD). Utilizing a formal synchronization policy description in the PS-setting, we are a…

Cited by 10SourceScholar
2019

ClusterSLAM: A SLAM Backend for Simultaneous Rigid Body Clustering and Motion Estimation

ICCV 2019poster

We present a practical backend for stereo visual SLAM which can simultaneously discover individual rigid bodies and compute their motions in dynamic environments. While recent factor graph based state optimization algorithms have shown their ability to robustly solve SLAM problems by treating dynami…

Cited by 97PDFScholar
2019

Probabilistic Projective Association and Semantic Guided Relocalization for Dense Reconstruction

ICRA 2019poster

We present a real-time dense mapping system which uses the predicted 2D semantic labels for optimizing the geometric quality of reconstruction. With a combination of Convolutional Neural Networks (CNNs) for 2D labeling and a Simultaneous Localization and Mapping (SLAM) system for camera trajectory e…

Cited by 14SourceScholar
2019

Towards Robust Curve Text Detection With Conditional Spatial Expansion

CVPR 2019poster

It is challenging to detect curve texts due to their irregular shapes and varying sizes. In this paper, we first investigate the deficiency of the existing curve detection methods and then propose a novel Conditional Spatial Expansion (CSE) mechanism to improve the performance of curve detection. In…

Cited by 105PDFScholar
2019

Unsupervised Feature Selection Based on Reconstruction Error Minimization

ICASSP 2019accepted

In this paper, we propose a novel unsupervised feature selection method, which is to minimize the data reconstruction error between each sample and a linear combination of its neighbors. Different from the conventional reconstruction-based feature selection method, we impose a nonnegative orthogonal…

Cited by 0SourceScholar
2018

A robust pose graph approach for city scale LiDAR mapping

IROS 2018poster

This paper presents a method for reconstructing globally consistent 3D High-Definition (HD) maps at city scale. Current approaches for eliminating cumulative drift are mainly based on the pose graph optimization under the constraint of scan-matching factors. The misaligned edges in the graph may hav…

Cited by 64SourceScholar
2018

Learning Markov Clustering Networks for Scene Text Detection

CVPR 2018poster

A novel framework named Markov Clustering Network (MCN) is proposed for fast and robust scene text detection. MCN predicts instance-level bounding boxes by firstly converting an image into a Stochastic Flow Graph (SFG) and then performing Markov Clustering on this graph. Our method can detect text o…

Cited by 139SourcePDFScholar