← Search

Xiaoyang Wu

25 accepted papers

2026

Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds

ICML 2026poster

Existing Vision-Language-Action (VLA) models typically take 2D images as visual input, which limits their spatial understanding in complex scenes. How can we incorporate 3D information to enhance VLA capabilities? We conduct a pilot study across different observation spaces and visual representation…

Cited by 0SourceScholar
2026

Less Gaussians, Texture More: 4K Feed-Forward Textured Splatting

ICLR 2026poster

Existing feed-forward 3D Gaussian Splatting methods typically rely on pixel-aligned primitives, which makes scaling to higher resolutions (e.g., 4K) prohibitive as the number of Gaussians grows quadratically with image resolution. We introduce LGTM (Less Gaussians, Texture More), a feed-forward and…

Cited by 0SourcecodeScholar
2026

Utonia: Toward One Encoder for All Point Clouds

ICML 2026poster

We dream of a future where point clouds from all domains can come together to shape a single model that benefits them all. Toward this goal, we present Utonia, a first step toward training a single self-supervised point transformer encoder across heterogeneous domains, spanning remote sensing, outdo…

Cited by 0SourceScholar
2025

Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations

NeurIPS 2025poster

Humans learn abstract concepts through multisensory synergy, and once formed, such representations can often be recalled from a single modality. Inspired by this principle, we introduce Concerto, a minimalist simulation of human concept learning for spatial cognition, combining 3D intra-modal self-d…

Cited by 0SourceScholar
2025

LiteReality: Graphic-Ready 3D Scene Reconstruction from RGB-D Scans

NeurIPS 2025poster

We propose LiteReality, a novel pipeline that converts RGB-D scans of indoor environments into compact, realistic, and interactive 3D virtual replicas. LiteReality not only reconstructs scenes that visually resemble reality but also supports key features essential for graphics pipelines, such as obj…

Cited by 0SourceScholar
2025

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning

NeurIPS 2025poster

This work explores enabling Chain-of-Thought (CoT) reasoning to link visual cues across multiple images. A straightforward solution is to adapt rule-based reinforcement learning for Vision-Language Models (VLMs). However, such methods typically rely on manually curated question-answer pairs, which c…

Cited by 0SourceScholar
2025

Sonata: Self-Supervised Learning of Reliable Point Representations

CVPR 2025highlight

In this paper, we question whether we have a reliable self-supervised point cloud model that can be used for diverse 3D tasks via simple linear probing, even with limited data and minimal computation. We find that existing 3D self-supervised learning approaches fall short when evaluated on represent…

2025

VIP: Vision Instructed Pre-training for Robotic Manipulation

ICML 2025poster

The effectiveness of scaling up training data in robotic manipulation is still limited. A primary challenge in manipulation is the tasks are diverse, and the trained policy would be confused if the task targets are not specified clearly. Existing works primarily rely on text instruction to describe…

Cited by 0SourcePDFScholar
2024

DreamComposer: Controllable 3D Object Generation via Multi-View Conditions

CVPR 2024poster

Utilizing pre-trained 2D large-scale generative models recent works are capable of generating high-quality novel views from a single in-the-wild image. However due to the lack of information from multiple views these works encounter difficulties in generating controllable novel views. In this paper…

2024

GPT4Point: A Unified Framework for Point-Language Understanding and Generation

CVPR 2024highlight

Multimodal Large Language Models (MLLMs) have excelled in 2D image-text comprehension and image generation but their understanding of the 3D world is notably deficient limiting progress in 3D language understanding and generation. To solve this problem we introduce GPT4Point an innovative groundbrea…

Cited by 43SourcePDFScholar
2024

GroupContrast: Semantic-aware Self-supervised Representation Learning for 3D Understanding

CVPR 2024poster

Self-supervised 3D representation learning aims to learn effective representations from large-scale unlabeled point clouds. Most existing approaches adopt point discrimination as the pretext task which assigns matched points in two distinct views as positive pairs and unmatched points as negative pa…

2024

LiT: Unifying LiDAR "Languages" with LiDAR Translator

NeurIPS 2024poster

LiDAR data exhibits significant domain gaps due to variations in sensors, vehicles, and driving environments, creating “language barriers” that limit the effective use of data across domains and the scalability of LiDAR perception models. To address these challenges, we introduce the LiDAR Translato…

2024

Multi-Space Alignments Towards Universal LiDAR Segmentation

CVPR 2024poster

A unified and versatile LiDAR segmentation model with strong robustness and generalizability is desirable for safe autonomous driving perception. This work presents M3Net a one-of-a-kind framework for fulfilling multi-task multi-dataset multi-modality LiDAR segmentation in a universal manner using j…

2024

OA-CNNs: Omni-Adaptive Sparse CNNs for 3D Semantic Segmentation

CVPR 2024poster

The booming of 3D recognition in the 2020s began with the introduction of point cloud transformers. They quickly overwhelmed sparse CNNs and became state-of-the-art models especially in 3D semantic segmentation. However sparse CNNs are still valuable networks due to their efficiency treasure and eas…

2024

OpenIns3D: Snap and Lookup for 3D Open-vocabulary Instance Segmentation

ECCV 2024poster

"In this work, we introduce OpenIns3D, a new 3D-input-only framework for 3D open-vocabulary scene understanding. The OpenIns3D framework employs a “Mask-Snap-Lookup” scheme. The “Mask” module learns class-agnostic mask proposals in 3D point clouds, the “Snap” module generates synthetic scene-level i…

2024

Point Transformer V3: Simpler Faster Stronger

CVPR 2024poster

This paper is not motivated to seek innovation within the attention mechanism. Instead it focuses on overcoming the existing trade-offs between accuracy and efficiency within the context of point cloud processing leveraging the power of scale. Drawing inspiration from recent advances in 3D large-sca…

Cited by 981SourcePDFScholar
2024

Towards Large-scale 3D Representation Learning with Multi-dataset Point Prompt Training

CVPR 2024poster

The rapid advancement of deep learning models is often attributed to their ability to leverage massive training data. In contrast such privilege has not yet fully benefited 3D deep learning mainly due to the limited availability of large-scale 3D datasets. Merging multiple available data sources and…

2024

UniPAD: A Universal Pre-training Paradigm for Autonomous Driving

CVPR 2024poster

In the context of autonomous driving the significance of effective feature learning is widely acknowledged. While conventional 3D self-supervised pre-training methods have shown widespread success most methods follow the ideas originally designed for 2D images. In this paper we present UniPAD a nove…

2023

A Spatio-Temporal Decomposition Network for Compressed Video Quality Enhancement

ICASSP 2023accepted

Compressed video quality enhancement has always been a widely concerned research. However, existing methods rarely build models from the consideration of object motion diversity and feature frequency distribution. In this paper, we propose a Spatio-Temporal Decomposition Network (STDN) to reduce the…

Cited by 0SourceScholar
2023

Hierarchical Dense Correlation Distillation for Few-Shot Segmentation

CVPR 2023highlight

Few-shot semantic segmentation (FSS) aims to form class-agnostic models segmenting unseen classes with only a handful of annotations. Previous methods limited to the semantic feature and prototype representation suffer from coarse segmentation granularity and train-set overfitting. In this work, we…

2023

MarS3D: A Plug-and-Play Motion-Aware Model for Semantic Segmentation on Multi-Scan 3D Point Clouds

CVPR 2023poster

3D semantic segmentation on multi-scan large-scale point clouds plays an important role in autonomous systems. Unlike the single-scan-based semantic segmentation task, this task requires distinguishing the motion states of points in addition to their semantic categories. However, methods designed fo…

2023

Masked Scene Contrast: A Scalable Framework for Unsupervised 3D Representation Learning

CVPR 2023poster

As a pioneering work, PointContrast conducts unsupervised 3D representation learning via leveraging contrastive learning over raw RGB-D frames and proves its effectiveness on various downstream tasks. However, the trend of large-scale unsupervised learning in 3D has yet to emerge due to two stumblin…

2023

Understanding Imbalanced Semantic Segmentation Through Neural Collapse

CVPR 2023poster

A recent study has shown a phenomenon called neural collapse in that the within-class means of features and the classifier weight vectors converge to the vertices of a simplex equiangular tight frame at the terminal phase of training for classification. In this paper, we explore the corresponding st…

2022

Point Transformer V2: Grouped Vector Attention and Partition-based Pooling

NeurIPS 2022accept

As a pioneering work exploring transformer architecture for 3D point cloud understanding, Point Transformer achieves impressive results on multiple highly competitive benchmarks. In this work, we analyze the limitations of the Point Transformer and propose our powerful and efficient Point Transforme…