← Search

Jingwei Huang

30 accepted papers

2026

ArtLLM: Generating Articulated Assets via 3D LLM

CVPR 2026

Creating interactive digital environments for gaming, robotics, and simulation relies on articulated 3D objects whose functionality emerges from their part geometry and kinematic structure. However, existing approaches remain fundamentally limited: optimization-based reconstruction methods require s

Cited by 0SourceScholar
2026

LATTICE: Democratize High-Fidelity 3D Generation at Scale

CVPR 2026

We present LATTICE, a new framework for high-fidelity 3D asset generation that bridges the quality and scalability gap between 3D and 2D generative models. While 2D image synthesis benefits from fixed spatial grids and well-established transformer architectures, 3D generation remains fundamentally m

Cited by 0SourcecodeScholar
2026

NaTex: Seamless Texture Generation as Latent Color Diffusion

CVPR 2026

We present NaTex, a native texture generation framework that predicts texture color directly in 3D space. In contrast to previous approaches that rely on baking 2D multi-view images synthesized by geometry-conditioned Multi-View Diffusion models (MVDs), NaTex avoids several inherent limitations of t

Cited by 8SourcecodeScholar
2026

PoseMaster: A Unified 3D Native Framework for Stylized Pose Generation

CVPR 2026

Pose stylization, which aims to synthesize stylized content aligning with target poses, serves as a fundamental task across 2D, 3D, and video domains. In the 3D realm, prevailing approaches typically rely on a cascade pipeline: first manipulating the image pose via 2D foundation models and subsequen

Cited by 0SourceScholar
2025

ClinBench: A Standardized Multi-Domain Framework for Evaluating Large Language Models in Clinical Information Extraction

NeurIPS 2025poster

Large Language Models (LLMs) offer substantial promise for clinical natural language processing (NLP); however, a lack of standardized benchmarking methodologies limits their objective evaluation and practical translation. To address this gap, we introduce ClinBench, an open-source, multi-model, mul…

Cited by 0SourcecodeScholar
2025

Epona: Autoregressive Diffusion World Model for Autonomous Driving

ICCV 2025poster

Diffusion models have demonstrated exceptional visual quality in video generation, making them promising for autonomous driving world modeling. However, existing video diffusion-based world models struggle with flexible-length, long-horizon predictions and integrating trajectory planning. This is be…

2025

LLGS: Unsupervised Gaussian Splatting for Image Enhancement and Reconstruction in Pure Dark Environment

ICRA 2025

D Gaussian Splatting has shown remarkable capabilities in novel view rendering tasks and exhibits significant potential for multi-view optimization. However, the original 3D Gaussian Splatting lacks color representation for inputs in lowlight environments. Simply using enhanced images as inputs woul

Cited by 3SourceScholar
2025

MaRI: Material Retrieval Integration across Domains

CVPR 2025poster

Accurate material retrieval is critical for creating realistic 3D assets. Existing methods rely on datasets that capture shape-invariant and lighting-varied representations of materials, which are scarce and face challenges due to limited diversity and inadequate real-world generalization. Most curr…

Cited by 1SourcePDFScholar
2025

SVG-Head: Hybrid Surface-Volumetric Gaussians for High-Fidelity Head Reconstruction and Real-Time Editing

ICCV 2025poster

Creating high-fidelity and editable head avatars is a pivotal challenge in computer vision and graphics, boosting many AR/VR applications. While recent advancements have achieved photorealistic renderings and plausible animation, head editing, especially real-time appearance editing, remains challen…

Cited by 0SourcePDFScholar
2025

Unleashing Vecset Diffusion Model for Fast Shape Generation

ICCV 2025poster

3D shape generation has greatly flourished through the development of so-called "native" 3D diffusion, particularly through the Vectset Diffusion Model (VDM). While recent advancements have shown promising results in generating high-resolution 3D shapes, VDM still struggles at high-speed generation.…

2024

CN-RMA: Combined Network with Ray Marching Aggregation for 3D Indoor Object Detection from Multi-view Images

CVPR 2024poster

This paper introduces CN-RMA a novel approach for 3D indoor object detection from multi-view images. We observe the key challenge as the ambiguity of image and 3D correspondence without explicit geometry to provide occlusion information. To address this issue CN-RMA leverages the synergy of 3D recon…

2024

NGP-RT: Fusing Multi-Level Hash Features with Lightweight Attention for Real-Time Novel View Synthesis

ECCV 2024poster

"This paper presents NGP-RT, a novel approach for enhancing the rendering speed of Instant-NGP to achieve real-time novel view synthesis. As a classic NeRF-based method, Instant-NGP stores implicit features in multi-level grids or hash tables and applies a shallow MLP to convert the implicit feature…

Cited by 0SourcePDFScholar
2024

SAI3D: Segment Any Instance in 3D Scenes

CVPR 2024poster

Advancements in 3D instance segmentation have traditionally been tethered to the availability of annotated datasets limiting their application to a narrow spectrum of object categories. Recent efforts have sought to harness vision-language models like CLIP for open-set semantic reasoning yet these m…

2022

Point Primitive Transformer for Long-Term 4D Point Cloud Video Understanding

ECCV 2022poster

"This paper proposes a 4D backbone for long-term point cloud video understanding. A typical way to capture spatial-temporal context is using 4Dconv or transformer without hierarchy. However, those methods are neither effective nor efficient enough due to camera motion, scene changes, sampling patter…

2021

DeepLM: Large-Scale Nonlinear Least Squares on Deep Learning Frameworks Using Stochastic Domain Decomposition

CVPR 2021poster

We propose a novel approach for large-scale nonlinear least squares problems based on deep learning frameworks. Nonlinear least squares are commonly solved with the Levenberg-Marquardt (LM) algorithm for fast convergence. We implement a general and efficient LM solver on a deep learning framework by…

Cited by 20PDFcodeScholar
2021

EPP-MVSNet: Epipolar-Assembling Based Depth Prediction for Multi-View Stereo

ICCV 2021poster

In this paper, we proposed EPP-MVSNet, a novel deep learning network for 3D reconstruction from multi-view stereo (MVS). EPP-MVSNet can accurately aggregate features at high resolution to a limited cost volume with an optimal depth range, thus, leads to effective and efficient 3D construction. Disti…

Cited by 137PDFScholar
2021

PrimitiveNet: Primitive Instance Segmentation With Local Primitive Embedding Under Adversarial Metric

ICCV 2021poster

We present PrimitiveNet, a novel approach for high-resolution primitive instance segmentation from point clouds on a large scale. Our key idea is to transform the global segmentation problem into easier local tasks. We train a high-resolution primitive embedding network to predict explicit geometry…

Cited by 29PDFcodeScholar
2020

Adversarial Texture Optimization From RGB-D Scans

CVPR 2020poster

Realistic color texture generation is an important step in RGB-D surface reconstruction, but remains challenging in practice due to inaccuracies in reconstructed geometry, misaligned camera poses, and view-dependent imaging artifacts. In this work, we present a novel approach for color texture gener…

Cited by 59PDFcodeScholar
2020

Deformation-Aware 3D Model Embedding and Retrieval

ECCV 2020poster

We introduce a new problem of mph{retrieving} 3D models that are mph{deformable} to a given query shape and present a novel deep mph{deformation-aware} embedding to solve this retrieval task. 3D model retrieval is a fundamental operation for recovering a clean and complete 3D model from a noisy and…

2020

Local Implicit Grid Representations for 3D Scenes

CVPR 2020poster

Shape priors learned from data are commonly used to reconstruct 3D objects from partial or noisy data. Yet no such shape priors are available for indoor scenes, since typical 3D autoencoders cannot handle their scale, complexity, or diversity. In this paper, we introduce Local Implicit Grid Represen…

Cited by 659PDFcodeScholar
2020

ShapeFlow: Learnable Deformation Flows Among 3D Shapes

NeurIPS 2020spotlight

We present ShapeFlow, a flow-based model for learning a deformation space for entire classes of 3D shapes with large intra-class variations. ShapeFlow allows learning a multi-template deformation space that is agnostic to shape topology, yet preserves fine geometric details. Different from a generat…

Cited by 101SourcePDFScholar
2019

Convolutional Neural Networks on Non-uniform Geometrical Signals Using Euclidean Spectral Transformation

ICLR 2019poster

Convolutional Neural Networks (CNN) have been successful in processing data signals that are uniformly sampled in the spatial domain (e.g., images). However, most data signals do not natively exist on a grid, and in the process of being sampled onto a uniform physical grid suffer significant aliasin…

Cited by 16SourcePDFScholar
2019

FrameNet: Learning Local Canonical Frames of 3D Surfaces From a Single RGB Image

ICCV 2019poster

In this work, we introduce the novel problem of identifying dense canonical 3D coordinate frames from a single RGB image. We observe that each pixel in an image corresponds to a surface in the underlying 3D geometry, where a canonical frame can be identified as represented by three orthogonal axes,…

Cited by 53PDFScholar
2019

NeurVPS: Neural Vanishing Point Scanning via Conic Convolution

NeurIPS 2019poster

We present a simple yet effective end-to-end trainable deep network with geometry-inspired convolutional operators for detecting vanishing points in images. Traditional convolutional neural networks rely on aggregating edge features and do not have mechanisms to directly exploit the geometric proper…

2019

Normalized Object Coordinate Space for Category-Level 6D Object Pose and Size Estimation

CVPR 2019oral

The goal of this paper is to estimate the 6D pose and dimensions of unseen object instances in an RGB-D image. Contrary to "instance-level" 6D pose estimation tasks, our problem assumes that no exact object CAD models are available during either training or testing time. To handle different and unse…

Cited by 870PDFcodeScholar
2019

Spherical CNNs on Unstructured Grids

ICLR 2019poster

We present an efficient convolution kernel for Convolutional Neural Networks (CNNs) on unstructured grids using parameterized differential operators while focusing on spherical signals such as panorama images or planetary signals. To this end, we replace conventional convolution kernels with linear…

2019

TextureNet: Consistent Local Parametrizations for Learning From High-Resolution Signals on Meshes

CVPR 2019oral

We introduce, TextureNet, a neural network architecture designed to extract features from high-resolution signals associated with 3D surface meshes (e.g., color texture maps). The key idea is to utilize a 4-rotational symmetric(4-RoSy) field to define a domain for convolution on a surface. Thou…

Cited by 141PDFScholar
2015

Automatic Thumbnail Generation Based on Visual Representativeness and Foreground Recognizability

ICCV 2015poster

We present an automatic thumbnail generation technique based on two essential considerations: how well they visually represent the original photograph, and how well the foreground can be recognized after the cropping and downsizing steps of thumbnailing. These factors, while important for the image…

Cited by 25PDFScholar