← Search

Yueqi Duan

45 accepted papers

2026

CFG-Ctrl: Control-Based Classifier-Free Diffusion Guidance

CVPR 2026

Classifier-Free Guidance (CFG) has emerged as a central approach for enhancing semantic alignment in flow-based diffusion models. In this paper, we explore a unified framework called **CFG-Ctrl**, which reinterprets CFG as a control applied to the first-order continuous-time generative flow, using t

Cited by 2SourcecodeScholar
2026

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling

ICLR 2026poster

Videos inherently represent 2D projections of a dynamic 3D world. However, our analysis suggests that video diffusion models trained solely on raw video data often fail to capture meaningful geometric-aware structure in their learned representations. To bridge this gap between video diffusion models…

Cited by 0SourcecodeScholar
2026

SimRecon: SimReady Compositional Scene Reconstruction from Real Videos

CVPR 2026

Compositional scene reconstruction seeks to create object-centric representations rather than holistic scenes from real-world videos, which is natively applicable for simulation and interaction. Conventional compositional reconstruction approaches primarily emphasize on visual appearance and show li

Cited by 0SourcecodeScholar
2026

Spatial-Aware Reduction Framework: Towards Efficient and Faithful Visual State Space Models

ICML 2026poster

Mamba demonstrates strong efficiency in modeling long visual sequences. However, when token reduction is applied to structurally enhanced Mamba variants, these models exhibit a severe performance collapse. We attribute this degradation to the spatially agnostic nature of existing reduction methods, …

Cited by 0SourceScholar
2025

D3QE: Learning Discrete Distribution Discrepancy-aware Quantization Error for Autoregressive-Generated Image Detection

ICCV 2025poster

The emergence of visual autoregressive (AR) models has revolutionized image generation while presenting new challenges for synthetic image detection. Unlike previous GAN or diffusion-based methods, AR models generate images through discrete token prediction, exhibiting both marked improvements in im…

2025

DimensionX: Create Any 3D and 4D Scenes from a Single Image with Decoupled Video Diffusion

ICCV 2025poster

In this paper, we introduce DimensionX, a framework designed to generate photorealistic 3D and 4D scenes from just a single image with video diffusion. Our approach begins with the insight that both the spatial structure of a 3D scene and the temporal evolution of a 4D scene can be effectively repre…

Cited by 0SourcePDFScholar
2025

LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion

ICCV 2025poster

Recovering 3D structures with open-vocabulary scene understanding from 2D images is a fundamental but daunting task. Recent developments have achieved this by performing per-scene optimization with embedded language information. However, they heavily rely on the calibrated dense-view reconstruction…

Cited by 0SourcePDFScholar
2025

Learning Efficient and Generalizable Human Representation with Human Gaussian Model

ICCV 2025poster

Modeling animatable human avatars from videos is a long-standing and challenging problem. While conventional methods require per-instance optimization, recent feed-forward methods have been proposed to generate 3D Gaussians with a learnable network.However, these methods predict independent Gaussian…

2025

Scene Splatter: Momentum 3D Scene Generation from Single Image with Video Diffusion Model

CVPR 2025poster

In this paper, we propose Scene Splatter, a momentum-based paradigm for video diffusion to generate generic scenes from single image. Existing methods, which employ video generation models to synthesize novel views, suffer from limited video length and scene inconsistency, leading to artifacts and d…

Cited by 1SourcePDFScholar
2025

ScenePainter: Semantically Consistent Perpetual 3D Scene Generation with Concept Relation Alignment

ICCV 2025poster

Perpetual 3D scene generation aims to produce long-range and coherent 3D view sequences, which is applicable for long-term video synthesis and 3D scene reconstruction. Existing methods follow a "navigate-and-imagine" fashion and rely on outpainting for successive view expansion. However, the generat…

Cited by 0SourcePDFScholar
2025

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence

NeurIPS 2025spotlight

Recent advancements in Multimodal Large Language Models (MLLMs) have significantly enhanced performance on 2D visual tasks. However, improving their spatial intelligence remains a challenge. Existing 3D MLLMs always rely on additional 3D or 2.5D data to incorporate spatial awareness, restricting the…

Cited by 0SourceScholar
2025

SpectralAR: Spectral Autoregressive Visual Generation

ICCV 2025poster

Autoregressive visual generation has garnered increasing attention due to its scalability and compatibility with other modalities compared with diffusion models. Most existing methods construct visual sequences as spatial patches for autoregressive generation. However, image patches are inherently p…

2025

SurfelSplat: Learning Efficient and Generalizable Gaussian Surfel Representations for Sparse-View Surface Reconstruction

NeurIPS 2025poster

3D Gaussian Splatting (3DGS) has demonstrated impressive performance in 3D scene reconstruction. Beyond novel view synthesis, it shows great potential for multi-view surface reconstruction. Existing methods employ optimization-based reconstruction pipelines that achieve precise and complete surface…

Cited by 0SourceScholar
2025

Video-T1: Test-time Scaling for Video Generation

ICCV 2025poster

With the scale capability of increasing training data, model size, and computational cost, video generation has achieved impressive results in digital creation, enabling users to express creativity across various domains. Recently, researchers in Large Language Models (LLMs) have expanded the scalin…

Cited by 0SourcePDFScholar
2024

Category-Level Multi-Part Multi-Joint 3D Shape Assembly

CVPR 2024poster

Shape assembly composes complex shapes geometries by arranging simple part geometries and has wide applications in autonomous robotic assembly and CAD modeling. Existing works focus on geometry reasoning and neglect the actual physical assembly process of matching and fitting joints which are the co…

Cited by 15SourcePDFScholar
2024

Gaussian Graph Network: Learning Efficient and Generalizable Gaussian Representations from Multi-view Images

NeurIPS 2024poster

3D Gaussian Splatting (3DGS) has demonstrated impressive novel view synthesis performance. While conventional methods require per-scene optimization, more recently several feed-forward methods have been proposed to generate pixel-aligned Gaussian representations with a learnable network, which are g…

Cited by 1SourcePDFScholar
2024

GeoAuxNet: Towards Universal 3D Representation Learning for Multi-sensor Point Clouds

CVPR 2024poster

Point clouds captured by different sensors such as RGB-D cameras and LiDAR possess non-negligible domain gaps. Most existing methods design different network architectures and train separately on point clouds from various sensors. Typically point-based methods achieve outstanding performances on eve…

2024

Memory-based Adapters for Online 3D Scene Perception

CVPR 2024poster

In this paper we propose a new framework for online 3D scene perception. Conventional 3D scene perception methods are offline i.e. take an already reconstructed 3D scene geometry as input which is not applicable in robotic applications where the input data is streaming RGB-D videos rather than a com…

Cited by 5SourcePDFScholar
2024

MirageRoom: 3D Scene Segmentation with 2D Pre-trained Models by Mirage Projection

CVPR 2024highlight

Nowadays leveraging 2D images and pre-trained models to guide 3D point cloud feature representation has shown a remarkable potential to boost the performance of 3D fundamental models. While some works rely on additional data such as 2D real-world images and their corresponding camera poses recent st…

Cited by 2SourcePDFScholar
2024

OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving

ECCV 2024poster

"Understanding how the 3D scene evolves is vital for making decisions in autonomous driving. Most existing methods achieve this by predicting the movements of object boxes, which cannot capture more fine-grained scene information. In this paper, we explore a new framework of learning a world model,…

2024

Semantic Flow: Learning Semantic Fields of Dynamic Scenes from Monocular Videos

ICLR 2024poster

In this work, we pioneer Semantic Flow, a neural semantic representation of dynamic scenes from monocular videos. In contrast to previous NeRF methods that reconstruct dynamic scenes from the colors and volume densities of individual points, Semantic Flow learns semantics from continuous flows that…

Cited by 5SourcePDFScholar
2024

Unique3D: High-Quality and Efficient 3D Mesh Generation from a Single Image

NeurIPS 2024poster

In this work, we introduce Unique3D, a novel image-to-3D framework for efficiently generating high-quality 3D meshes from single-view images, featuring state-of-the-art generation fidelity and strong generalizability. Previous methods based on Score Distillation Sampling (SDS) can produce diversifie…

2023

HOI-aware Adaptive Network for Weakly-supervised Action Segmentation

IJCAI 2023poster

In this paper, we propose an HOI-aware adaptive network named AdaAct for weakly-supervised action segmentation. Most existing methods learn a fixed network to predict the action of each frame with the neighboring frames. However, this would result in ambiguity when estimating similar actions, such a…

Cited by 6SourcePDFScholar
2023

MonoNeRF: Learning a Generalizable Dynamic Radiance Field from Monocular Videos

ICCV 2023poster

In this paper, we target at the problem of learning a generalizable dynamic radiance field from monocular videos. Different from most existing NeRF methods that are based on multiple views, monocular videos only contain one view at each timestamp, thereby suffering from ambiguity along the view dire…

Cited by 43PDFcodeScholar
2023

SEFormer: Structure Embedding Transformer for 3D Object Detection

AAAI 2023technical

Effectively preserving and encoding structure features from objects in irregular and sparse LiDAR points is a crucial challenge to 3D object detection on the point cloud. Recently, Transformer has demonstrated promising performance on many 2D and even 3D vision tasks. Compared with the fixed and ri…

2022

Bridge-Prompt: Towards Ordinal Action Understanding in Instructional Videos

CVPR 2022poster

Action recognition models have shown a promising capability to classify human actions in short video clips. In a real scenario, multiple correlated human actions commonly occur in particular orders, forming semantically meaningful human activities. Conventional action recognition approaches focus on…

Cited by 88PDFcodeScholar
2022

Learning Dynamic Facial Radiance Fields for Few-Shot Talking Head Synthesis

ECCV 2022poster

"Talking head synthesis is an emerging technology with wide applications in film dubbing, virtual avatars and online education. Recent NeRF-based methods generate more natural talking videos, as they better capture the 3D structural information of faces. However, a specific model needs to be trained…

2022

Learning Transferable Human-Object Interaction Detector With Natural Language Supervision

CVPR 2022poster

It is difficult to construct a data collection including all possible combinations of human actions and interacting objects due to the combinatorial nature of human-object interactions (HOI). In this work, we aim to develop a transferable HOI detector for unseen interactions. Existing HOI detectors…

Cited by 66PDFcodeScholar
2022

Object Pursuit: Building a Space of Objects via Discriminative Weight Generation

ICLR 2022poster

We propose a framework to continuously learn object-centric representations for visual learning and understanding. Existing object-centric representations either rely on supervisions that individualize objects in the scene, or perform unsupervised disentanglement that can hardly deal with complex sc…

2022

Uncertainty-Aware Representation Learning for Action Segmentation

IJCAI 2022poster

In this paper, we propose an uncertainty-aware representation Learning (UARL) method for action segmentation. Most existing action segmentation methods exploit continuity information of the action period to predict frame-level labels, which ignores the temporal ambiguity of the transition region bet…

Cited by 17SourcePDFScholar
2021

CAPTRA: CAtegory-Level Pose Tracking for Rigid and Articulated Objects From Point Clouds

ICCV 2021poster

In this work, we tackle the problem of category-level online pose tracking for objects from point cloud sequences. For the first time, we propose a unified framework that can handle 9DoF object pose tracking for novel rigid object instances as well as per-part pose tracking for articulated objects f…

Cited by 114PDFcodeScholar
2021

Vector Neurons: A General Framework for SO(3)-Equivariant Networks

ICCV 2021poster

Invariance and equivariance to the rotation group have been widely discussed in the 3D deep learning community for pointclouds. Yet most proposed methods either use complex mathematical tools that may limit their accessibility, or are tied to specific input data types and network architectures. In t…

Cited by 345PDFcodeScholar
2018

GraphBit: Bitwise Interaction Mining via Deep Reinforcement Learning

CVPR 2018poster

In this paper, we propose a GraphBit method to learn deep binary descriptors in a directed acyclic graph unsupervisedly, representing bitwise interactions as edges between the nodes of bits. Conventional binary representation learning methods enforce each element to be binarized into zero or one. Ho…

Cited by 37SourcePDFScholar