← Search

Dongyue Lu

12 accepted papers

2026

EventDrive: Event Cameras for Vision-Language Driving Intelligence

CVPR 2026

Event cameras sense the world through asynchronous brightness changes with microsecond latency and high dynamic range, offering motion fidelity far beyond frame-based sensors and capturing temporal structure that conventional exposures often miss. These properties make events a powerful complement t

Cited by 0SourceScholar
2026

LiDARCrafter: Dynamic 4D World Modeling from LiDAR Sequences

AAAI 2026technical

Generative world models have become essential data engines for autonomous driving, yet most focus on videos or occupancy grids and overlook the unique challenges of LiDAR. Extending LiDAR generation to dynamic 4D modeling requires addressing controllability, temporal coherence, and standardized eval

Cited by 0SourcePDFScholar
2026

WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World

CVPR 2026

Generative world models are reshaping embodied AI, enabling agents to synthesize realistic 4D driving environments that look convincing but often fail physically or behaviorally. Despite rapid progress, the field still lacks a unified way to assess whether generated worlds preserve geometry, obey ph

Cited by 0SourcecodeScholar
2025

EventFly: Event Camera Perception from Ground to the Sky

CVPR 2025poster

Cross-platform adaptation in event-based dense perception is crucial for deploying event cameras across diverse settings, such as vehicles, drones, and quadrupeds, each with unique motion dynamics, viewpoints, and class distributions. In this work, we introduce EventFly, a framework for robust cross…

Cited by 0SourcePDFScholar
2025

FlexEvent: Towards Flexible Event-Frame Object Detection at Varying Operational Frequencies

NeurIPS 2025poster

Event cameras offer unparalleled advantages for real-time perception in dynamic environments, thanks to the microsecond-level temporal resolution and asynchronous operation. Existing event detectors, however, are limited by fixed-frequency paradigms and fail to fully exploit the high-temporal resolu…

Cited by 0SourceScholar
2025

GEAL: Generalizable 3D Affordance Learning with Cross-Modal Consistency

CVPR 2025poster

Identifying affordance regions on 3D objects from semantic cues is essential for robotics and human-machine interaction. However, existing 3D affordance learning methods struggle with generalization and robustness due to limited annotated data and a reliance on 3D backbones focused on geometric enco…

2025

Perspective-Invariant 3D Object Detection

ICCV 2025poster

With the rise of robotics, LiDAR-based 3D object detection has garnered significant attention in both academia and industry. However, existing datasets and methods predominantly focus on vehicle-mounted platforms, leaving other autonomous platforms underexplored. To bridge this gap, we introduce Pi3…

2025

Reconstructing Close Human Interaction with Appearance and Proxemics Reasoning

CVPR 2025poster

Due to visual ambiguities and inter-person occlusions, existing human pose estimation methods cannot recover plausible close interactions from in-the-wild videos. Even state-of-the-art large foundation models (e.g., SAM) cannot accurately distinguish human semantics in such challenging scenarios. In…

Cited by 0SourcePDFScholar
2025

Spiral: Semantic-Aware Progressive LiDAR Scene Generation and Understanding

NeurIPS 2025poster

Leveraging diffusion models, 3D LiDAR scene generation has achieved great success in both range-view and voxel-based representations. While recent voxel-based approaches can generate both geometric structures and semantic labels, existing range-view methods are limited to producing unlabeled LiDAR s…

Cited by 0SourceScholar
2025

Talk2Event: Grounded Understanding of Dynamic Scenes from Event Cameras

NeurIPS 2025spotlight

Event cameras offer microsecond-level latency and robustness to motion blur, making them ideal for understanding dynamic environments. Yet, connecting these asynchronous streams to human language remains an open challenge. We introduce Talk2Event, the first large-scale benchmark for language-driven…

Cited by 0SourceScholar