← Search

Baiyong Ding

2 accepted papers

2026

SpikeVLA: Vision-Language-Action Models with Spiking Neural Networks

ICML 2026poster

Vision-Language-Action (VLA) models have become a central paradigm for embodied intelligence. However, most existing approaches are built on large-scale Transformers, resulting in substantial inference latency and energy consumption that limit their practical deployment in low-power, real-time scena…

Cited by 0SourceScholar
2026

UnsOcc: 3D Semantic Occupancy Prediction in Unstructured Scene Via Rendering Fusion

ICRA 2026poster

Unstructured scenes present unique challenges for autonomous driving, as irregular obstacles and sparse scene layouts undermine the effectiveness of traditional perception methods such as 3D object detection. 3D semantic occupancy prediction has emerged as a prominent focus due to its ability to pro…