← Search

Linna Song

5 accepted papers

2026

CoT-VLNBench: A Benchmark for Visual Chain-of-Thought Reasoning in Vision-Language-Navigation Robots

AAAI 2026technical

Recent advances in vision language models (VLMs) have demonstrated remarkable potential in embodied navigation tasks. However, existing robot-centric datasets primarily focus on traditional 3D tasks such as perception and prediction, lacking adequate support for vision-language tasks. Vision-languag

Cited by 0SourcePDFScholar
2026

GaussianFusion: Unified 3D Gaussian Representation for Multi-Modal Fusion Perception

ICLR 2026poster

The bird’s-eye view (BEV) representation enables multi-sensor features to be fused within a unified space, serving as the primary approach for achieving comprehensive multi-task perception. However, the discrete grid representation of BEV leads to significant detail loss and limits feature alignment…

Cited by 0SourceScholar
2025

TAD-E2E: A Large-scale End-to-end Autonomous Driving Dataset

ICCV 2025poster

End-to-end autonomous driving technology has recently become a focal point of research and application in autonomous driving. State-of-the-art (SOTA) methods are often trained and evaluated on the NuScenes dataset. However, the NuScenes dataset, introduced in 2019 for 3D perception tasks, faces seve…

Cited by 0SourcePDFScholar
2024

Spatial-Aware Dynamic Lightweight Self-Supervised Monocular Depth Estimation

RA-L 2024

Self-supervised monocular depth estimation has attracted extensive attention in recent years. Lightweight depth estimation methods are crucial for resource-constrained edge devices. However, existing lightweight methods often encounter the challenge of limited representation capacity and increased c

Cited by 10SourceScholar
2024

Unified Single-Stage Transformer Network for Efficient RGB-T Tracking

IJCAI 2024poster

Most existing RGB-T tracking networks extract modality features in a separate manner, which lacks interaction and mutual guidance between modalities. This limits the network's ability to adapt to the diverse dual-modality appearances of targets and the dynamic relationships between the modalities. A…