← Search

Jie Guo

39 accepted papers

2026

BrepGaussian: CAD reconstruction from Multi-View Images with Gaussian Splatting

CVPR 2026

The boundary representation (B-Rep) models a 3D solid as its explicit boundaries: trimmed corners, edges, and faces. Recovering B-Rep representation from unstructured data is a challenging and valuable task of computer vision and graphics. Recent advances in deep learning have greatly improved the r

Cited by 0SourceScholar
2026

Combinatorial Bandit Bayesian Optimization for Tensor Outputs

ICLR 2026poster

Bayesian optimization (BO) has been widely used to optimize expensive and black-box functions across various domains. Existing BO methods have not addressed tensor-output functions. To fill this gap, we propose a novel tensor-output BO method. Specifically, we first introduce a tensor-output Gaussia…

Cited by 0SourceScholar
2026

LoG3D: Ultra-High-Resolution 3D Shape Modeling via Local-to-Global Partitioning

CVPR 2026

Generating high-fidelity 3D contents remains a fundamental challenge due to the complexity of representing arbitrary topologies--such as open surfaces and intricate internal structures--while preserving geometric details. Prevailing methods based on signed distance fields (SDFs) are hampered by cost

Cited by 0SourceScholar
2026

MatMart: Material Reconstruction of 3D Objects via Diffusion

CVPR 2026

Applying diffusion models to physically-based material estimation and generation has recently gained prominence. In this paper, we propose MatMart, a novel material reconstruction framework for 3D objects, offering the following advantages. First, MatMart adopts a two-stage reconstruction, starting

Cited by 0SourcecodeScholar
2025

Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models

NeurIPS 2025poster

Recent advances in Multimodal Large Language Models (MLLMs) have significantly improved 2D visual understanding, prompting interest in their application to complex 3D reasoning tasks. However, it remains unclear whether these models can effectively capture the detailed spatial information required f…

Cited by 0SourceScholar
2025

CombatVLA: An Efficient Vision-Language-Action Model for Combat Tasks in 3D Action Role-Playing Games

ICCV 2025poster

Recent advances in Vision-Language-Action models (VLAs) have expanded the capabilities of embodied intelligence. However, significant challenges remain in real-time decision-making in complex 3D environments, which demand second-level responses, high-resolution perception, and tactical reasoning und…

2025

EPCPE: A Real-time End-to-End Pipeline for RGB-based Category-level 6D Pose Estimation

ICASSP 2025accepted

RGB-based category-level 6D pose estimation methods have faced significant challenges in achieving real-time performance, primarily due to the design of two-stage pipeline. To address this issue, we propose a novel end-to-end pipeline named EPCPE. In detail, we first extract implicit rotation featur…

Cited by 0SourceScholar
2025

EdgeMovingNet: Edge-preserving Point Cloud Reconstruction via Joint Geometry Features

CVPR 2025poster

Point cloud reconstruction is a critical process in 3D representation and reverse engineering. When it comes to CAD models, edges are significant features that play a crucial role in characterizing the geometry of 3D shapes. However, few points are exactly sampled on edges during acquisition, result…

Cited by 0SourcePDFScholar
2025

GaRe: Relightable 3D Gaussian Splatting for Outdoor Scenes from Unconstrained Photo Collections

ICCV 2025poster

We propose a 3D Gaussian splatting-based framework for outdoor relighting that leverages intrinsic image decomposition to precisely integrate sunlight, sky radiance, and indirect lighting from unconstrained photo collections. Unlike prior methods that compress the per-image global illumination into…

Cited by 0SourcePDFScholar
2025

High-quality Point Cloud Oriented Normal Estimation via Hybrid Angular and Euclidean Distance Encoding

CVPR 2025poster

The proliferation of Light Detection and Ranging (LiDAR) technology has facilitated the acquisition of three-dimensional point clouds, which are integral to applications in VR, AR, and Digital Twin. Oriented normals, critical for 3D reconstruction and scene analysis, cannot be directly extracted fro…

Cited by 0SourcePDFScholar
2025

IRMamba: Pixel Difference Mamba with Layer Restoration for Infrared Small Target Detection

AAAI 2025technical

Infrared small target detection (IRSTD) focuses on identifying small targets in infrared images. Despite advancements with deep learning, challenges persist due to the IR long-range imaging mechanism, where targets are small, dim, and easily lost in noise and background clutter. Current deep learnin…

Cited by 0SourcePDFScholar
2025

MOCID: Motion Context and Displacement Information Learning for Moving Infrared Small Target Detection

AAAI 2025technical

In the field of Moving Infrared Small Target Detection (MIRSTD), current methods typically use sequential modeling with two individual modules for spatial and temporal processing. However, such a modeling strategy lacks clear guidance on the motion and displacement difference between moving targets…

2025

Multimodal Prior Learning with Double Constraint Alignment for Snapshot Spectral Compressive Imaging

IJCAI 2025

The objective of snapshot spectral compressive imaging reconstruction is to recover the 3D hyperspectral image (HSI) from a 2D measurement. Existing methods either focus on network architecture design or simply introduce image-level prior to the model. However, these methods lack guiding information

Cited by 0SourcePDFScholar
2025

Real-Time Neural Denoising with Render-Aware Knowledge Distillation

AAAI 2025technical

Real-time Monte Carlo (MC) ray tracing with low sampling rates demands a denoising algorithm that adeptly balances the trade-off between quality and efficiency. Previous works have paid much attention on designing delicate denoising architecture while ignoring model compression. In this work, we pre…

Cited by 0SourcePDFScholar
2025

SAIST: Segment Any Infrared Small Target Model Guided by Contrastive Language-Image Pretraining

CVPR 2025poster

Infrared Small Target Detection (IRSTD) aims to identify low signal-to-noise ratio small targets in infrared images with complex backgrounds, which is crucial for various applications. However, existing IRSTD methods typically rely solely on image modalities for processing, which fail to fully captu…

Cited by 0SourcePDFScholar
2025

SGCR: Spherical Gaussians for Efficient 3D Curve Reconstruction

CVPR 2025poster

Neural rendering techniques have made substantial progress in generating photo-realistic 3D scenes. The latest 3D Gaussian Splatting technique has achieved high quality novel view synthesis as well as fast rendering speed. However, 3D Gaussians lack proficiency in defining accurate 3D geometric stru…

2025

Sparse Point Cloud Patches Rendering via Splitting 2D Gaussians

CVPR 2025poster

Current learning-based methods predict NeRF or 3D Gaussians from point clouds to achieve photo-realistic rendering but still depend on categorical priors, dense point clouds, or additional refinements. Hence, we introduce a novel point cloud rendering method by predicting 2D Gaussians from point clo…

2024

Exploring Multi-Modal Control in Music-Driven Dance Generation

ICASSP 2024accepted

Existing music-driven 3D dance generation methods mainly concentrate on high-quality dance generation, but lack sufficient control during the generation process. To address these issues, we propose a unified framework capable of generating high-quality dance movements and supporting multi-modal cont…

Cited by 0SourceScholar
2024

IRPruneDet: Efficient Infrared Small Target Detection via Wavelet Structure-Regularized Soft Channel Pruning

AAAI 2024technical

Infrared Small Target Detection (IRSTD) refers to detecting faint targets in infrared images, which has achieved notable progress with the advent of deep learning. However, the drive for improved detection accuracy has led to larger, intricate models with redundant parameters, causing storage and co…

2024

LiDAR-Net: A Real-scanned 3D Point Cloud Dataset for Indoor Scenes

CVPR 2024poster

In this paper we present LiDAR-Net a new real-scanned indoor point cloud dataset containing nearly 3.6 billion precisely point-level annotated points covering an expansive area of 30000m^2. It encompasses three prevalent daily environments including learning scenes working scenes and living scenes.…

Cited by 8SourcePDFScholar
2024

Lodge: A Coarse to Fine Diffusion Network for Long Dance Generation Guided by the Characteristic Dance Primitives

CVPR 2024poster

We propose Lodge a network capable of generating extremely long dance sequences conditioned on given music. We design Lodge as a two-stage coarse to fine diffusion architecture and propose the characteristic dance primitives that possess significant expressiveness as intermediate representations bet…

2024

Practical Measurements of Translucent Materials with Inter-Pixel Translucency Prior

CVPR 2024poster

Material appearance is a key component of photorealism with a pronounced impact on human perception. Although there are many prior works targeting at measuring opaque materials using light-weight setups (e.g. consumer-level cameras) little attention is paid on acquiring the optical properties of tra…

Cited by 1SourcePDFScholar
2024

Prompt3D: Random Prompt Assisted Weakly-Supervised 3D Object Detection

CVPR 2024poster

The prohibitive cost of annotations for fully supervised 3D indoor object detection limits its practicality. In this work we propose Random Prompt Assisted Weakly-supervised 3D Object Detection termed as Prompt3D a weakly-supervised approach that leverages position-level labels to overcome this chal…

2024

Semantic Human Mesh Reconstruction with Textures

CVPR 2024poster

The field of 3D detailed human mesh reconstruction has made significant progress in recent years. However current methods still face challenges when used in industrial applications due to unstable results low-quality meshes and a lack of UV unwrapping and skinning weights. In this paper we present S…

2023

ESSAformer: Efficient Transformer for Hyperspectral Image Super-resolution

ICCV 2023poster

Single hyperspectral image super-resolution (single-HSI-SR) aims to restore a high-resolution hyperspectral image from a low-resolution observation. However, the prevailing CNN-based approaches have shown limitations in building long-range dependencies and capturing interaction information between s…

Cited by 83PDFcodeScholar
2023

LED: Label Correlation Enhanced Decoder for Multi-Label Text Classification

ICASSP 2023accepted

Multi-label text classification, which aims to predict the relevant labels for each given document, is one of the fundamental tasks of natural language processing. Recent studies have utilized Transformer, which embeds texts and class labels into a joint space to capture the label correlation. Howev…

Cited by 0SourceScholar
2023

Support or Refute: Analyzing the Stance of Evidence to Detect Out-of-Context Mis- and Disinformation

EMNLP 2023long main

Mis- and disinformation online have become a major societal problem as major sources of online harms of different kinds. One common form of mis- and disinformation is out-of-context (OOC) information, where different pieces of information are falsely associated, e.g., a real image combined with a fa…

Cited by 28SourcecodeScholar
2023

Symmetric Shape-Preserving Autoencoder for Unsupervised Real Scene Point Cloud Completion

CVPR 2023poster

Unsupervised completion of real scene objects is of vital importance but still remains extremely challenging in preserving input shapes, predicting accurate results, and adapting to multi-category data. To solve these problems, we propose in this paper an Unsupervised Symmetric Shape-Preserving Auto…

Cited by 18SourcePDFScholar
2022

ISNet: Shape Matters for Infrared Small Target Detection

CVPR 2022poster

Infrared small target detection (IRSTD) refers to extracting small and dim targets from blurred backgrounds, which has a wide range of applications such as traffic management and marine rescue. Due to the low signal-to-noise ratio and low contrast, infrared targets are easily submerged in the backgr…

Cited by 353PDFcodeScholar
2022

SAR-to-Optical Image Translation via Neural Partial Differential Equations

IJCAI 2022poster

Synthetic Aperture Radar (SAR) becomes prevailing in remote sensing while SAR images are challenging to interpret by human visual perception due to the active imaging mechanism and speckle noise. Recent researches on SAR-to-optical image translation provide a promising solution and have attracted in…

Cited by 11SourcePDFScholar
2022

Unsupervised Point Cloud Completion and Segmentation by Generative Adversarial Autoencoding Network

NeurIPS 2022accept

Most existing point cloud completion methods assume the input partial point cloud is clean, which is not practical in practice, and are Most existing point cloud completion methods assume the input partial point cloud is clean, which is not the case in practice, and are generally based on supervised…

Cited by 9SourcePDFScholar
2021

GLAVNet: Global-Local Audio-Visual Cues for Fine-Grained Material Recognition

CVPR 2021poster

In this paper, we aim to recognize materials with combined use of auditory and visual perception. To this end, we construct a new dataset named GLAudio that consists of both the geometry of the object being struck and the sound captured from either modal sound synthesis (for virtual objects) or real…

Cited by 8PDFScholar
2021

Hierarchical Disentangled Representation Learning for Outdoor Illumination Estimation and Editing

ICCV 2021poster

Data-driven sky models have gained much attention in outdoor illumination prediction recently, showing superior performance against analytical models. However, naively compressing an outdoor panorama into a low-dimensional latent vector, as existing models have done, causes two major problems. One i…

Cited by 19PDFScholar
2020

Deep Surface Normal Estimation on the 2-Sphere with Confidence Guided Semantic Attention

ECCV 2020poster

We propose a deep convolutional neural network (CNN) to estimate surface normal from a single color image accompanied with a low-quality depth channel. Unlike most previous works, we predict the normal on the 2-sphere rather than the 3D Euclidean space, which produces naturally normalized values and…

Cited by 3SourcePDFScholar
2019

Learning Actor Relation Graphs for Group Activity Recognition

CVPR 2019poster

Modeling relation between actors is important for recognizing group activity in a multi-person scene. This paper aims at learning discriminative relation between actors efficiently using deep models. To this end, we propose to build a flexible and efficient \rm Actor Relation Graph (ARG) to simult…

Cited by 333PDFcodeScholar
2019

Toward an Efficient Hybrid Interaction Paradigm for Object Manipulation in Optical See-Through Mixed Reality

IROS 2019poster

Human-computer interaction (HCI) plays an important role in the near-field mixed reality, in which the hand-based interaction is one of the most widely-used interaction modes, especially in the applications based on optical see-through head-mounted displays (OST-HMDs). In this paper, such interactio…

Cited by 5SourceScholar