← Search

Ziyu Zhu

21 accepted papers

2026

Mean Estimation from Coarse Data: Characterizations and Efficient Algorithms

ICLR 2026poster

Coarse data arise when learners observe only partial information about samples; namely, a set containing the sample rather than its exact value. This occurs naturally through measurement rounding, sensor limitations, and lag in economic systems. We study Gaussian mean estimation from coarse data, wh…

Cited by 0SourceScholar
2026

RefDiffMap: Diffusion-Guided Progressive Refinement for Vectorized HD Map Construction

RA-L 2026

High-definition (HD) map learning serves as an essential component of autonomous driving scene understanding, providing structured priors for planning and prediction. Recent transformer-based methods regress vectorized map elements via deformable attention over Bird's-Eye View (BEV) features. They t

Cited by 0SourceScholar
2026

RefDiffMap: Diffusion-Guided Progressive Refinement for Vectorized HD Map Construction

ICRA 2026poster

High-definition (HD) map learning serves as an essential component of autonomous driving scene understanding, providing structured priors for planning and prediction. Recent transformer-based methods regress vectorized map elements via deformable attention over Bird’s-Eye View (BEV) features. They t…

Cited by 0SourceScholar
2026

SceneCOT: Eliciting Chain-of-Thought Reasoning in 3D Scenes

ICLR 2026poster

Existing research of 3D LLMs still struggles to achieve efficient and explainable reasoning, primarily due to the under-exploration of the mechanism of human-like scene-object grounded reasoning. This paper bridges the gap by presenting a novel framework. We first introduce a Chain-of-Thought reason…

Cited by 0SourcecodeScholar
2026

Sparse Meets Dense: Correspondence Guided Robotic Manipulation with Rigid-Deformable Interactions

ICRA 2026poster

Manipulation involving rigid-deformable interactions, such as hanging clothes or dressing humans, is essential for household robots. Compared to single-object manipulation or interactions between rigid bodies, these tasks are particularly challenging due to the rich multi-point contacts and the comp…

Cited by 0Scholar
2025

DexGarmentLab: Dexterous Garment Manipulation Environment with Generalizable Policy

NeurIPS 2025spotlight

Garment manipulation is a critical challenge due to the diversity in garment categories, geometries, and deformations. Despite this, humans can effortlessly handle garments, thanks to the dexterity of our hands. However, existing research in the field has struggled to replicate this level of dexteri…

Cited by 0SourcecodeScholar
2025

From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes

NeurIPS 2025poster

3D visual grounding has made notable progress in localizing objects within complex 3D scenes. However, grounding referring expressions beyond objects in 3D scenes remains unexplored. In this paper, we introduce Anywhere3D-Bench, a holistic 3D visual grounding benchmark consisting of 2,886 referring…

Cited by 0SourceScholar
2025

GarmentPile: Point-Level Visual Affordance Guided Retrieval and Adaptation for Cluttered Garments Manipulation

CVPR 2025poster

Cluttered garments manipulation poses significant challenges in robotics due to the complex, deformable nature of garments and intricate garment relations. Unlike single-garment manipulation, cluttered scenarios require managing complex garment entanglements and interactions, while maintaining garme…

2025

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding

CVPR 2025poster

Open-vocabulary 3D scene understanding is pivotal for enhancing physical intelligence, as it enables embodied agents to interpret and interact dynamically within real-world environments. This paper introduces MPEC, a novel Masked Point-Entity Contrastive learning method for open-vocabulary 3D semant…

Cited by 2SourcePDFScholar
2025

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation

ICCV 2025poster

Embodied scene understanding requires not only comprehending visual-spatial information that has been observed but also determining where to explore next in the 3D physical world. Existing 3D Vision-Language (3D-VL) models primarily focus on grounding objects in static observations from 3D reconstru…

Cited by 0SourcePDFScholar
2025

On Domain-Adaptive Post-Training for Multimodal Large Language Models

EMNLP 2025

Adapting general multimodal large language models (MLLMs) to specific domains, such as scientific and industrial fields, is highly significant in promoting their practical applications. This paper systematically investigates domain adaptation of MLLMs via post-training, focusing on data synthesis, t

Cited by 0SourcePDFScholar
2025

SAMap: Semantic Alignment for HD Map Detection Domain Generalization Under Varying Weather and Lighting

IROS 2025

High-definition (HD) maps are crucial for autonomous driving systems. Despite recent advances in learning-based HD map prediction methods, these approaches experience significant performance degradation when encountering unseen weather or lighting conditions due to feature distribution discrepancies

Cited by 0SourceScholar
2025

Unveiling the Mist over 3D Vision-Language Understanding: Object-centric Evaluation with Chain-of-Analysis

CVPR 2025poster

Existing 3D vision-language (3D-VL) benchmarks fall short in evaluating 3D-VL models, creating a "mist" that obscures rigorous insights into model capabilities and 3D-VL tasks. This mist persists due to three key limitations. First, flawed test data, like ambiguous referential text in the grounding…

2024

ECM-OPCC: Efficient Context Model for Octree-Based Point Cloud Compression

ICASSP 2024accepted

Recently, deep learning methods have shown promising results in point cloud compression. However, previous octree-based approaches either lack sufficient context or have high decoding complexity (e.g. > 900s). To address this problem, we propose a sufficient yet efficient context model and design an…

Cited by 0SourceScholar
2024

GarmentLab: A Unified Simulation and Benchmark for Garment Manipulation

NeurIPS 2024poster

Manipulating garments and fabrics has long been a critical endeavor in the development of home-assistant robots. However, due to complex dynamics and topological structures, garment manipulations pose significant challenges. Recent successes in reinforcement learning and vision-based methods offer p…

2023

3D-VisTA: Pre-trained Transformer for 3D Vision and Text Alignment

ICCV 2023poster

3D vision-language grounding (3D-VL) is an emerging field that aims to connect the 3D physical world with natural language, which is crucial for achieving embodied intelligence. Current 3D-VL models rely heavily on sophisticated modules, auxiliary losses, and optimization tricks, which calls for a s…

Cited by 129PDFScholar
2023

Bit Allocation using Optimization

ICML 2023poster

In this paper, we consider the problem of bit allocation in Neural Video Compression (NVC). First, we reveal a fundamental relationship between bit allocation in NVC and Semi-Amortized Variational Inference (SAVI). Specifically, we show that SAVI with GoP (Group-of-Picture)-level likelihood is equiv…

2023

Learnable Flow Model Conditioned on Graph Representation Memory for Anomaly Detection

ICASSP 2023accepted

Anomaly detection could be applied in a wide range of fields from industrial scene to medical imaging analysis. Although invertible flow models are developed to accomplish unsupervised anomaly detection, they are usually hard to train and have limited capabilities of accurately modeling the distribu…

Cited by 0SourceScholar
2022

DuMLP-Pin: A Dual-MLP-Dot-Product Permutation-Invariant Network for Set Feature Extraction

AAAI 2022technical

Existing permutation-invariant methods can be divided into two categories according to the aggregation scope, i.e. global aggregation and local one. Although the global aggregation methods, e. g., PointNet and Deep Sets, get involved in simpler structures, their performance is poorer than the local…