← Search

Wenxuan Zhu

7 accepted papers

2026

LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion

RSS 2026poster

Recent robot foundation models largely rely on large-scale behavior cloning, which imitates expert actions but discards transferable dynamics knowledge embedded in heterogeneous embodied data. While the Unified World Model (UWM) formulation has the potential to leverage such diverse data, existing i…

Cited by 0SourceScholar
2025

4D-Bench: Benchmarking Multi-modal Large Language Models for 4D Object Understanding

ICCV 2025poster

Multimodal Large Language Models (MLLMs) have demonstrated impressive 2D image/video understanding capabilities.However, there are no publicly standardized benchmarks to assess the abilities of MLLMs in understanding the 4D objects.In this paper, we introduce 4D-Bench, the first benchmark to evaluat…

2024

Exploring Learngene via Stage-wise Weight Sharing for Initializing Variable-sized Models

IJCAI 2024poster

In practice, we usually need to build variable-sized models adapting for diverse resource constraints in different application scenarios, where weight initialization is an important step prior to training. The Learngene framework, introduced recently, firstly learns one compact part termed as learng…

2024

TrackNeRF: Bundle Adjusting NeRF from Sparse and Noisy Views via Feature Tracks

ECCV 2024poster

"Neural radiance fields (NeRFs) generally require many images with accurate poses for accurate novel view synthesis, which does not reflect realistic setups where views can be sparse and poses can be noisy. Previous solutions for learning NeRFs with sparse views and noisy poses only consider local g…

2024

Unleashing Channel Potential: Space-Frequency Selection Convolution for SAR Object Detection

CVPR 2024poster

Deep Convolutional Neural Networks (DCNNs) have achieved remarkable performance in synthetic aperture radar (SAR) object detection but this comes at the cost of tremendous computational resources partly due to extracting redundant features within a single convolutional layer. Recent works either del…

Cited by 14SourcePDFScholar
2024

Vivid-ZOO: Multi-View Video Generation with Diffusion Model

NeurIPS 2024poster

While diffusion models have shown impressive performance in 2D image/video generation, diffusion-based Text-to-Multi-view-Video (T2MVid) generation remains underexplored. The new challenges posed by T2MVid generation lie in the lack of massive captioned multi-view videos and the complexity of modeli…

Cited by 11SourcePDFScholar