← Search

Yikai Wang

62 accepted papers

2026

A Comprehensive Survey of Interaction Techniques in 3D Scene Generation

IJCAI 2026

The rapid evolution of 3D scene generation has revolutionized content creation across domains such as gaming, film production, and architectural visualization. Within this landscape, interaction techniques serve as the pivotal bridge connecting user intent with generative models, enabling precise co

Cited by 0Scholar
2026

ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation

CVPR 2026

Recent advances in robot manipulation have leveraged pre-trained vision-language models (VLMs) and explored integrating 3D spatial signals into these models for effective action prediction, giving rise to the promising vision-language-action (VLA) paradigm. However, most existing approaches overlook

Cited by 0SourcecodeScholar
2026

BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models

ICLR 2026poster

Recent progress in aligning image and video generative models with Group Relative Policy Optimization (GRPO) has improved human preference alignment, but existing variants remain inefficient due to sequential rollouts and large numbers of sampling steps, unreliable credit assignment,as sparse termin…

Cited by 0SourcecodeScholar
2026

Conformal Reliability: A New Evaluation Metric for Conditional Generation

ICML 2026poster

Conditional generative models have recently achieved remarkable success in various applications. However, a suitable metric for evaluating the reliability of these models, which takes into account their inherent uncertainty, is still lacking. Existing metrics, which typically assess a single output,…

Cited by 0SourceScholar
2026

Direct 3D-Aware Object Insertion via Decomposed Visual Proxies

ICML 2026poster

Object insertion aims to seamlessly composite a reference object into a specified region of a background image. Recent diffusion-based methods achieve high visual quality but formulate insertion as a simple 2D inpainting task, providing no explicit control over the object’s 3D pose and limiting thei…

Cited by 0SourceScholar
2026

FlowDreamer: A RGB-D World Model With Flow-Based Motion Representations for Robot Manipulation

RA-L 2026

This paper investigates training better visual world models for robot manipulation, i.e., models that can predict future visual observations by conditioning on past frames and robot actions. Specifically, we consider world models that operate on RGB-D frames (RGB-D world models). As opposed to canon

Cited by 12SourcecodeScholar
2026

FlowDreamer: A RGB-D World Model with Flow-Based Motion Representations for Robot Manipulation

ICRA 2026poster

This paper investigates training better visual world models for robot manipulation, i.e., models that can predict future visual observations by conditioning on past frames and robot actions. Specifically, we consider world models that operate on RGB-D frames (RGB-D world models). As opposed to canon…

2026

GaussianPile: A Unified Sparse Gaussian Splatting Framework for Slice-based Volumetric Reconstruction

CVPR 2026

Slice-based volumetric imaging is widely applied and it demands representations that compress aggressively while preserving internal structure for analysis. This paper introduces GaussianPile, unifying 3D Gaussian splatting with an imaging system-aware focus model to address this challenge. Our new

Cited by 0SourceScholar
2026

Learning Human-Robot Collaboration via Heterogeneous-Agent Lyapunov Policy Optimization

ICML 2026oral

To improve generalization and resilience in human–robot collaboration (HRC), robots must handle the combinatorial diversity of human behaviors and contexts, motivating multi-agent reinforcement learning (MARL). However, inherent heterogeneity between robots and humans creates a rationality gap (RG) …

Cited by 0SourceScholar
2026

Light of Normals: Unified Feature Representation for Universal Photometric Stereo

ICLR 2026poster

Universal photometric stereo (PS) is defined by two factors: it must (i) operate under arbitrary, unknown lighting conditions and (ii) avoid reliance on specific illumination models. Despite progress (e.g., SDM UniPS), two challenges remain. First, current encoders cannot guarantee that illumination…

Cited by 0SourcecodeScholar
2026

PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion

ICML 2026spotlight

Achieving a complete and explorable 360-degree visual world is a cornerstone of immersive content creation. While recent advances in video generation have achieved impressive results, they follow a 2D paradigm that treats content generation as transitions of 2D pixels, lacking an intrinsic understan…

Cited by 0SourceScholar
2026

SceneTransporter: Optimal Transport-Guided Compositional Latent Diffusion for Single-Image Structured 3D Scene Generation

ICLR 2026poster

We introduce SceneTransporter, an end-to-end framework for structured 3D scene generation from a single image. While existing methods generate part-level 3D objects, they often fail to organize these parts into distinct instances in open-world scenes. Through a debiased clustering probe, we reveal a…

Cited by 0SourcecodeScholar
2026

Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation

ICLR 2026poster

Camera-centric understanding and generation are two cornerstones of spatial intelligence, yet they are typically studied in isolation. We present Puffin, a unified camera-centric multimodal model that extends spatial awareness along the camera dimension. Puffin integrates language regression and dif…

Cited by 0SourcecodeScholar
2026

Yo'City: Personalized and Boundless 3D Realistic City Scene Generation via Self-Critic Expansion

CVPR 2026

Realistic 3D city generation is fundamental to a wide range of applications, including virtual reality and digital twins. However, most existing methods rely on training a single diffusion model, which limits their ability to generate personalized and boundless city-scale scenes. In this paper, we p

Cited by 0SourceScholar
2025

3D StreetUnveiler with Semantic-aware 2DGS - a simple baseline

ICLR 2025poster

Unveiling an empty street from crowded observations captured by in-car cameras is crucial for autonomous driving. However, removing all temporarily static objects, such as stopped vehicles and standing pedestrians, presents a significant challenge. Unlike object-centric 3D inpainting, which relies o…

Cited by 0SourcePDFScholar
2025

AR-1-to-3: Single Image to Consistent 3D Object via Next-View Prediction

ICCV 2025poster

Novel view synthesis (NVS) is a cornerstone for image-to-3d creation. However, existing works still struggle to maintain consistency between the generated views and the input views, especially when there is a significant camera pose difference, leading to poor-quality 3D geometries and textures. We…

2025

Adaptive Pruning of Pretrained Transformer via Differential Inclusions

ICLR 2025poster

Large transformers have demonstrated remarkable success, making it necessary to compress these models to reduce inference costs while preserving their performance. Current compression algorithms prune transformers at fixed compression ratios, requiring a unique pruning process for each ratio, which…

Cited by 0SourcePDFScholar
2025

AnimateAnything: Consistent and Controllable Animation for Video Generation

CVPR 2025poster

We propose a unified approach for video-controlled generation, enabling text-based guidance and manual annotations to control the generation of videos, similar to camera direction guidance. Specifically, we designed a two-stage algorithm. In the first stage, we convert all control information into f…

Cited by 10SourcePDFScholar
2025

DeepMesh: Auto-Regressive Artist-mesh Creation with Reinforcement Learning

ICCV 2025poster

Triangle meshes play a crucial role in 3D applications for efficient manipulation and rendering. While auto-regressive methods generate structured meshes by predicting discrete vertex tokens, they are often constrained by limited face counts and mesh incompleteness. To address these challenges, we p…

Cited by 0SourcePDFScholar
2025

DimensionX: Create Any 3D and 4D Scenes from a Single Image with Decoupled Video Diffusion

ICCV 2025poster

In this paper, we introduce DimensionX, a framework designed to generate photorealistic 3D and 4D scenes from just a single image with video diffusion. Our approach begins with the insight that both the spatial structure of a 3D scene and the temporal evolution of a 4D scene can be effectively repre…

Cited by 0SourcePDFScholar
2025

EMOS: Embodiment-aware Heterogeneous Multi-robot Operating System with LLM Agents

ICLR 2025poster

Heterogeneous multi-robot systems (HMRS) have emerged as a powerful ap- proach for tackling complex tasks that single robots cannot manage alone. Current large-language-model-based multi-agent systems (LLM-based MAS) have shown success in areas like software development and operating systems, but ap…

Cited by 1SourcePDFScholar
2025

FlexWorld: Progressively Expanding 3D Scenes for Flexible-View Exploration

NeurIPS 2025poster

Generating flexible-view 3D scenes, including 360° rotation and zooming, from single images is challenging due to a lack of 3D data. To this end, we introduce FlexWorld, a novel framework that progressively constructs a persistent 3D Gaussian splatting representation by synthesizing and integrating…

Cited by 0SourcecodeScholar
2025

Human-Imperceptible, Machine-Recognizable Images

IJCAI 2025

Massive human-related data is collected to train neural networks for computer vision tasks. A major conflict is exposed relating to software engineers between better developing AI systems and distancing from the sensitive training data. To reconcile this conflict, the paper proposes an efficient pri

2025

Human2LocoMan: Learning Versatile Quadrupedal Manipulation with Human Pretraining

RSS 2025poster

Quadrupedal robots have demonstrated impressive locomotion capabilities in complex environments, but equipping them with autonomous versatile manipulation skills in a scalable way remains a significant challenge. In this work, we introduce a system that integrates data collection and imitation learn…

Cited by 0PDFcodeScholar
2025

LocoTouch: Learning Dynamic Quadrupedal Transport with Tactile Sensing

CoRL 2025poster

Quadrupedal robots have demonstrated remarkable agility and robustness in traversing complex terrains. However, they struggle with dynamic object interactions, where contact must be precisely sensed and controlled. To bridge this gap, we present LocoTouch, a system that equips quadrupedal robots wit…

Cited by 0SourceScholar
2025

MeshAnything V2: Artist-Created Mesh Generation with Adjacent Mesh Tokenization

ICCV 2025poster

Meshes are the de facto 3D representation in the industry but are labor-intensive to produce. Recently, a line of research has focused on autoregressively generating meshes. This approach processes meshes into a sequence composed of vertices and then generates them vertex by vertex, similar to how a…

2025

MeshGen: Generating PBR Textured Mesh with Render-Enhanced Auto-Encoder and Generative Data Augmentation

CVPR 2025highlight

In this paper, we introduce MeshGen, an advanced image-to-3D pipeline that generates high-quality 3D meshes with detailed geometry and physically based rendering (PBR) textures. Addressing the challenges faced by existing 3D native diffusion models, such as suboptimal auto-encoder performance, limit…

2025

PivotMesh: Generic 3D Mesh Generation via Pivot Vertices Guidance

ICLR 2025poster

Generating compact and sharply detailed 3D meshes poses a significant challenge for current 3D generative models. Different from extracting dense meshes from neural representation, some recent works try to model the native mesh distribution (i.e., a set of triangles), which generates more compact re…

Cited by 11SourcePDFScholar
2025

ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and Reasoning

CVPR 2025poster

Open-vocabulary 3D visual grounding and reasoning aim to localize objects in a scene based on implicit language descriptions, even when they are occluded. This ability is crucial for tasks such as vision-language navigation and autonomous robotics. However, current methods struggle because they rely…

Cited by 0SourcePDFScholar
2025

Spatial-Temporal Aware Visuomotor Diffusion Policy Learning

ICCV 2025poster

Visual imitation learning is effective for robots to learn versatile tasks. However, many existing methods rely on behavior cloning with supervised historical trajectories, limiting their 3D spatial and 4D spatiotemporal awareness. Consequently, these methods struggle to capture the 3D structures an…

Cited by 0SourcePDFScholar
2025

Towards Enhanced Image Inpainting: Mitigating Unwanted Object Insertion and Preserving Color Consistency

CVPR 2025highlight

Recent advances in image inpainting increasingly use generative models to handle large irregular masks. However, these models can create unrealistic inpainted images due to two main issues: (1) Unwanted object insertion: Even with unmasked areas as context, generative models may still generate arbit…

2024

CRM: Single Image to 3D Textured Mesh with Convolutional Reconstruction Model

ECCV 2024poster

"Feed-forward 3D generative models like the Large Reconstruction Model (LRM) [?] have demonstrated exceptional generation speed. However, the transformer-based methods do not leverage the geometric priors of the triplane component in their architecture, often leading to sub-optimal quality given the…

2024

DreamReward: Aligning Human Preference in Text-to-3D Generation

ECCV 2024poster

"3D content creation from text prompts has shown remarkable success recently. However, current text-to-3D methods often generate 3D results that do not align well with human preferences. In this paper, we present a comprehensive framework, coined DreamReward, to learn and improve text-to-3D models f…

2024

GaussianEditor: Swift and Controllable 3D Editing with Gaussian Splatting

CVPR 2024poster

3D editing plays a crucial role in many areas such as gaming and virtual reality. Traditional 3D editing methods which rely on representations like meshes and point clouds often fall short in realistically depicting complex scenes. On the other hand methods based on implicit 3D representations like…

2024

InstructPix2NeRF: Instructed 3D Portrait Editing from a Single Image

ICLR 2024poster

With the success of Neural Radiance Field (NeRF) in 3D-aware portrait editing, a variety of works have achieved promising results regarding both quality and 3D consistency. However, these methods heavily rely on per-prompt optimization when handling natural language as editing instructions. Due to t…

2024

LeftRefill: Filling Right Canvas based on Left Reference through Generalized Text-to-Image Diffusion Model

CVPR 2024poster

This paper introduces LeftRefill an innovative approach to efficiently harness large Text-to-Image (T2I) diffusion models for reference-guided image synthesis. As the name implies LeftRefill horizontally stitches reference and target views together as a whole input. The reference image occupies the…

2024

NeuroPictor: Refining fMRI-to-Image Reconstruction via Multi-individual Pretraining and Multi-level Modulation

ECCV 2024poster

"Recent fMRI-to-image approaches mainly focused on associating fMRI signals with specific conditions of pre-trained diffusion models. These approaches, while producing high-quality images, capture only a limited aspect of the complex information in fMRI signals and offer little detailed control over…

2024

Test-Time Linear Out-of-Distribution Detection

CVPR 2024poster

Out-of-Distribution (OOD) detection aims to address the excessive confidence prediction by neural networks by triggering an alert when the input sample deviates significantly from the training distribution (in-distribution) indicating that the output may not be reliable. Current OOD detection approa…

2024

Towards Global Optimal Visual In-Context Learning Prompt Selection

NeurIPS 2024poster

Visual In-Context Learning (VICL) is a prevailing way to transfer visual foundation models to new tasks by leveraging contextual information contained in in-context examples to enhance learning and prediction of query sample. The fundamental problem in VICL is how to select the best prompt to activa…

Cited by 4SourcePDFScholar
2024

Unified Lexical Representation for Interpretable Visual-Language Alignment

NeurIPS 2024poster

Visual-Language Alignment (VLA) has gained a lot of attention since CLIP's groundbreaking work. Although CLIP performs well, the typical direct latent feature alignment lacks clarity in its representation and similarity scores. On the other hand, lexical representation, a vector whose element repr…

2024

Vidu4D: Single Generated Video to High-Fidelity 4D Reconstruction with Dynamic Gaussian Surfels

NeurIPS 2024poster

Video generative models are receiving particular attention given their ability to generate realistic and imaginative frames. Besides, these models are also observed to exhibit strong 3D consistency, significantly enhancing their potential to act as world simulators. In this work, we present Vidu4D,…

Cited by 18SourcePDFScholar
2023

Benchmarking Robustness of 3D Object Detection to Common Corruptions

CVPR 2023poster

3D object detection is an important task in autonomous driving to perceive the surroundings. Despite the excellent performance, the existing 3D detectors lack the robustness to real-world corruptions caused by adverse weathers, sensor noises, etc., provoking concerns about the safety and reliability…

2023

Coarse-to-Fine Amodal Segmentation with Shape Prior

ICCV 2023poster

Amodal object segmentation is a challenging task that involves segmenting both visible and occluded parts of an object. In this paper, we propose a novel approach, called Coarse-to-Fine Segmentation (C2F-Seg), that addresses this problem by progressively modeling the amodal segmentation. C2F-Seg…

Cited by 24PDFcodeScholar
2023

Compacting Binary Neural Networks by Sparse Kernel Selection

CVPR 2023poster

Binary Neural Network (BNN) represents convolution weights with 1-bit values, which enhances the efficiency of storage and computation. This paper is motivated by a previously revealed phenomenon that the binary kernels in successful BNNs are nearly power-law distributed: their values are mostly clu…

Cited by 7SourcePDFScholar
2023

ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation

NeurIPS 2023spotlight

Score distillation sampling (SDS) has shown great promise in text-to-3D generation by distilling pretrained large-scale text-to-image diffusion models, but suffers from over-saturation, over-smoothing, and low-diversity problems. In this work, we propose to model the 3D parameter as a random variabl…

2023

Root Pose Decomposition Towards Generic Non-rigid 3D Reconstruction with Monocular Videos

ICCV 2023poster

This work focuses on the 3D reconstruction of non-rigid objects based on monocular RGB video sequences. Concretely, we aim at building high-fidelity models for generic object categories and casually captured scenes. To this end, we do not assume known root poses of objects, and do not utilize catego…

Cited by 9PDFcodeScholar
2023

Towards Effective Adversarial Textured 3D Meshes on Physical Face Recognition

CVPR 2023highlight

Face recognition is a prevailing authentication solution in numerous biometric applications. Physical adversarial attacks, as an important surrogate, can identify the weaknesses of face recognition systems and evaluate their robustness before deployed. However, most existing physical attacks are eit…

2022

Bridged Transformer for Vision and Point Cloud 3D Object Detection

CVPR 2022poster

3D object detection is a crucial research topic in computer vision, which usually uses 3D point clouds as input in conventional setups. Recently, there is a trend of leveraging multiple sources of input data, such as complementing the 3D point cloud with 2D images that often have richer color and fe…

Cited by 53PDFScholar
2022

Multimodal Token Fusion for Vision Transformers

CVPR 2022poster

Many adaptations of transformers have emerged to address the single-modal vision tasks, where self-attention modules are stacked to handle input sources like images. Intuitively, feeding multiple modalities of data to vision transformers could improve the performance, yet the inner-modal attentive w…

Cited by 217PDFcodeScholar
2022

Sound Adversarial Audio-Visual Navigation

ICLR 2022poster

Audio-visual navigation task requires an agent to find a sound source in a realistic, unmapped 3D environment by utilizing egocentric audio-visual observations. Existing audio-visual navigation works assume a clean environment that solely contains the target sound, which, however, would not be suita…

2021

Sub-Bit Neural Networks: Learning To Compress and Accelerate Binary Neural Networks

ICCV 2021poster

In the low-bit quantization field, training Binarized Neural Networks (BNNs) is the extreme solution to ease the deployment of deep models on resource-constrained devices, having the lowest storage cost and significantly cheaper bit-wise operations compared to 32-bit floating-point counterparts. In…

Cited by 20PDFcodeScholar
2020

Deep Multimodal Fusion by Channel Exchanging

NeurIPS 2020poster

Deep multimodal fusion by using multiple sources of data for classification or regression has exhibited a clear advantage over the unimodal counterpart on various applications. Yet, current methods including aggregation-based and alignment-based fusion are still inadequate in balancing the trade-off…

2020

Resolution Switchable Networks for Runtime Efficient Image Recognition

ECCV 2020poster

We propose a general method to train a single convolutional neural network which is capable of switching image resolutions at inference. Thus the running speed can be selected to meet various computational resource limits. Networks trained with the proposed method are named Resolution Switchable Net…

2017

Average SCR loss analysis for polarimetric STAP with Kronecker structured covariance matrix

ICASSP 2017accepted

The paper presents the average signal-to-clutter loss (SCRL) analysis for polarimetric space-time adaptive processing by exploiting the Kronecker structure of the clutter covariance matrix (CM). An expression for the average SCRL as a function of the mean square error of the corresponding CM estimat…

Cited by 0SourceScholar