← Search

Wen Liu

24 accepted papers

2026

Pointer-CAD: Unifying B-Rep and Command Sequences via Pointer-based Edges & Faces Selection

CVPR 2026

Constructing computer-aided design (CAD) models is labor-intensive but essential for engineering and manufacturing. Recent advances in Large Language Models (LLMs) have inspired the LLM-based CAD generation by representing CAD as command sequences. But these methods struggle in practical scenarios b

Cited by 0SourcecodeScholar
2025

DUNE: Sim2Real Transfer for Depth-based Navigation in Unstructured Dynamic Indoor Environments

ICASSP 2025accepted

Collision-free navigation in dynamic environments, especially with moving pedestrians, is crucial for mobile robots. This paper introduces DUNE, a depth-based policy trained in simulation for collision-free navigation of Ackermann mobile robots in unstructured indoor environments. DUNE uses a CNN-LS…

Cited by 0SourceScholar
2025

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

CVPR 2025poster

We introduce Janus, an autoregressive framework that unifies multimodal understanding and generation. Prior research often relies on a single visual encoder for both tasks, such as Chameleon. However, due to the differing levels of information granularity required by multimodal understanding and gen…

2025

JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

CVPR 2025poster

We present JanusFlow, a powerful framework that unifies image understanding and generation in a single model.JanusFlow introduces a minimalist architecture that integrates autoregressive language models with rectified flow, a state-of-the-art method in generative modeling.Our key finding demonstrate…

2024

DreamCraft3D: Hierarchical 3D Generation with Bootstrapped Diffusion Prior

ICLR 2024poster

We present DreamCraft3D, a hierarchical 3D content generation method that produces high-fidelity and coherent 3D objects. We tackle the problem by leveraging a 2D reference image to guide the stages of geometry sculpting and texture boosting. A central focus of this work is to address the consistenc…

2024

OneRestore: A Universal Restoration Framework for Composite Degradation

ECCV 2024poster

"In real-world scenarios, image impairments often manifest as composite degradations, presenting a complex interplay of elements such as low light, haze, rain, and snow. Despite this reality, existing restoration methods typically target isolated degradation types, thereby falling short in environme…

2024

PM-INR: Prior-Rich Multi-Modal Implicit Large-Scale Scene Neural Representation

AAAI 2024technical

Recent advancements in implicit neural representations have contributed to high-fidelity surface reconstruction and photorealistic novel view synthesis. However, with the expansion of the scene scale, such as block or city level, existing methods will encounter challenges because traditional samplin…

Cited by 2SourcePDFScholar
2024

Paint3D: Paint Anything 3D with Lighting-Less Texture Diffusion Models

CVPR 2024poster

This paper presents Paint3D a novel coarse-to-fine generative framework that is capable of producing high-resolution lighting-less and diverse 2K UV texture maps for untextured 3D meshes conditioned on text or image inputs. The key challenge addressed is generating high-quality textures without embe…

2024

Unbounded-GS: Extending 3D Gaussian Splatting With Hybrid Representation for Unbounded Large-Scale Scene Reconstruction

RA-L 2024

Modeling large-scale scenes from multi-view images is challenging due to the trade-off dilemma between visual quality and computational cost. Existing NeRF-based methods have made advancements in neural implicit representation through volumetric ray-marching, but still struggle to deal with cubicall

Cited by 7SourceScholar
2023

A Large-Scale Outdoor Multi-Modal Dataset and Benchmark for Novel View Synthesis and Implicit Scene Reconstruction

ICCV 2023poster

Neural Radiance Fields (NeRF) has achieved impressive results in single object scene reconstruction and novel view synthesis, as demonstrated on many single modality and single object focused indoor scene datasets like DTU, BMVS, and NeRF Synthetic. However, the study of NeRF on large-scale outdoor…

Cited by 29PDFScholar
2023

Executing Your Commands via Motion Diffusion in Latent Space

CVPR 2023poster

We study a challenging task, conditional human motion generation, which produces plausible human motion sequences according to various conditional inputs, such as action classes or textual descriptors. Since human motions are highly diverse and have a property of quite different distribution from co…

2023

Michelangelo: Conditional 3D Shape Generation based on Shape-Image-Text Aligned Latent Representation

NeurIPS 2023poster

We present a novel alignment-before-generation approach to tackle the challenging task of generating general 3D shapes based on 2D images or texts. Directly learning a conditional generative model from images or texts to 3D shapes is prone to producing inconsistent results with the conditions becaus…

2023

MotionGPT: Human Motion as a Foreign Language

NeurIPS 2023poster

Though the advancement of pre-trained large language models unfolds, the exploration of building a unified model for language and other multimodal data, such as motion, remains challenging and untouched so far. Fortunately, human motion displays a semantic coupling akin to human language, often perc…

2023

PDF: Point Diffusion Implicit Function for Large-scale Scene Neural Representation

NeurIPS 2023poster

Recent advances in implicit neural representations have achieved impressive results by sampling and fusing individual points along sampling rays in the sampling space. However, due to the explosively growing sampling space, finely representing and synthesizing detailed textures remains a challenge f…

Cited by 5SourcePDFScholar
2022

Coordinates Are NOT Lonely - Codebook Prior Helps Implicit Neural 3D representations

NeurIPS 2022accept

Implicit neural 3D representation has achieved impressive results in surface or scene reconstruction and novel view synthesis, which typically uses the coordinate-based multi-layer perceptrons (MLPs) to learn a continuous scene representation. However, existing approaches, such as Neural Radiance Fi…

2021

Appearance-Motion Memory Consistency Network for Video Anomaly Detection

AAAI 2021technical

Abnormal event detection in the surveillance video is an essential but challenging task, and many methods have been proposed to deal with this problem. The previous methods either only consider the appearance information or directly integrate the results of appearance and motion information without…

2021

Speech Drives Templates: Co-Speech Gesture Synthesis With Learned Templates

ICCV 2021poster

Co-speech gesture generation is to synthesize a gesture sequence that not only looks real but also matches with the input speech audio. Our method generates the movements of a complete upper body, including arms, hands, and the head. Although recent data-driven methods achieve great success, challen…

Cited by 82PDFcodeScholar
2020

Encoding Structure-Texture Relation with P-Net for Anomaly Detection in Retinal Images

ECCV 2020poster

Anomaly detection in retinal image refers to the identification of abnormality caused by various retinal diseases/lesions, by only leveraging normal images in training phase. Normal images from healthy subjects often have regular structures (e.g., the structured blood vessels in the fundus image, or…

2019

Liquid Warping GAN: A Unified Framework for Human Motion Imitation, Appearance Transfer and Novel View Synthesis

ICCV 2019poster

We tackle the human motion imitation, appearance transfer, and novel view synthesis within a unified framework, which means that the model once being trained can be used to handle all these tasks. The existing task-specific methods mainly use 2D keypoints (pose) to estimate the human body structure.…

Cited by 333PDFcodeScholar
2018

Future Frame Prediction for Anomaly Detection – A New Baseline

CVPR 2018poster

Anomaly detection in videos refers to the identification of events that do not conform to expected behavior. However, almost all existing methods tackle the problem by minimizing the reconstruction errors of training data, which cannot guarantee a larger reconstruction error for an abnormal event. I…

2018

Geolocation of Unknown Emitters Using Tdoa of Path Rays Through the Ionosphere by Multiple Coordinated Distant Receivers

ICASSP 2018accepted

We consider the problem of unknown emitter geolocation using the time difference of arrival (TDOA) of the path rays through the ionosphere by multiple coordinated distant receivers. We formulate the geolocation in the sense of maximum likelihood with the exact ray expressions for the quasi-parabolic…

Cited by 0SourceScholar