← Search

Ruqi Huang

20 accepted papers

2026

Animator-Centric Skeleton Generation on Objects with Fine-Grained Details

CVPR 2026

Skeleton generation is essential for animating 3D assets, but current deep learning methods remain limited: they cannot handle the growing structural complexity of modern models and offer minimal controllability, creating a major bottleneck for real-world animation workflows. To address this, we pro

Cited by 0SourceScholar
2026

DanceTogether: Generating Interactive Multi-Person Video without Identity Drifting

ICLR 2026poster

Controllable video generation (CVG) has advanced rapidly, yet current systems falter when more than one actor must move, interact, and exchange positions under noisy control signals. We address this gap with DanceTogether, the first end-to-end diffusion framework that turns a single reference image…

Cited by 0SourceScholar
2026

GS-Playground: A High-Throughput Photorealistic Simulator for Vision-Informed Robot Learning

RSS 2026poster

Embodied AI research is undergoing a shift toward vision-centric perceptual paradigms. While massively parallel simulators have catalyzed breakthroughs in proprioception-based locomotion, their potential remains largely untapped for vision-centric tasks due to the prohibitive computational overhead …

Cited by 0SourceScholar
2026

LottieGPT: Tokenizing Vector Animation for Autoregressive Generation

CVPR 2026

Despite rapid progress in video generation, existing models are incapable of producing vector animation, a dominant and highly expressive form of multimedia on the Internet. Vector animations offer resolution-independence, compactness, semantic structure, and editable parametric motion representatio

Cited by 0SourceScholar
2026

OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention

ICML 2026poster

Humans perceive the world through diverse modalities that operate synergistically to support a holistic understanding of their surroundings. However, existing omnimodal models still exhibit substantial performance degradation on visual tasks when the audio modality is incorporated. We identify this …

Cited by 0SourceScholar
2026

SIFThinker: Spatially-Aware Image Focus for Visual Reasoning

AAAI 2026technical

Current multimodal large language models (MLLMs) still face significant challenges in complex visual tasks (e.g., spatial understanding, fine-grained perception). Prior methods have tried to incorporate visual reasoning, however, they fail to leverage attention correction with spatial cues to iterat

Cited by 0SourcePDFScholar
2026

Seeing Space and Motion: Enhancing Latent Actions with Geometric and Dynamic Awareness for Vision-Language-Action Models

ICRA 2026poster

Latent Action Models (LAMs) enable Vision-Language-Action (VLA) systems to learn semantic action representations from large-scale unannotated data. Yet, we identify two bottlenecks of LAMs: 1) the commonly adopted end-to-end trained image encoder suffers from poor spatial understanding; 2) LAMs can …

2026

Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views

CVPR 2026

Though recent advances in vision-language models (VLMs) have achieved remarkable progress across a wide range of multimodal tasks, understanding 3D spatial relationships from limited views remains a significant challenge. Previous reasoning methods typically rely on pure text (e.g., topological cogn

Cited by 0SourcecodeScholar
2025

ARMO: Autoregressive Rigging for Multi-Category Objects

ICCV 2025poster

Recent advancements in large-scale generative models have significantly improved the quality and diversity of 3D shape generation. However, most existing methods focus primarily on generating static 3D models, overlooking the potential dynamic nature of certain shapes, such as humanoids, animals, an…

Cited by 0SourcePDFScholar
2025

DISCOVERSE: Efficient Robot Simulation in Complex High-Fidelity Environments

IROS 2025

We present Discoverse, the first unified, modular, open-source 3DGS-based simulation framework for Real2Sim2Real robot learning. It features a holistic Real2Sim pipeline that synthesizes hyper-realistic geometry and appearance of complex real-world scenarios, paving the way for analyzing and bridgin

Cited by 14SourcecodeScholar
2025

DRiVE: Diffusion-based Rigging Empowers Generation of Versatile and Expressive Characters

CVPR 2025poster

Recent advances in generative models have enabled high-quality 3D character reconstruction from multi-modal. However, animating these generated characters remains a challenging task, especially for complex elements like garments and hair, due to the lack of large-scale datasets and effective rigging…

2025

DV-Matcher: Deformation-based Non-rigid Point Cloud Matching Guided by Pre-trained Visual Features

CVPR 2025poster

In this paper, we present DV-Matcher, a novel learning-based framework for estimating dense correspondences between non-rigidly deformable point clouds. Learning directly from unstructured point clouds without meshing or manual labelling, our framework delivers high-quality dense correspondences, wh…

2024

OmniSeg3D: Omniversal 3D Segmentation via Hierarchical Contrastive Learning

CVPR 2024poster

Towards holistic understanding of 3D scenes a general 3D segmentation method is needed that can segment diverse objects without restrictions on object quantity or categories while also reflecting the inherent hierarchical structure. To achieve this we propose OmniSeg3D an omniversal segmentation met…

2024

XScale-NVS: Cross-Scale Novel View Synthesis with Hash Featurized Manifold

CVPR 2024poster

We propose XScale-NVS for high-fidelity cross-scale novel view synthesis of real-world large-scale scenes. Existing representations based on explicit surface suffer from discretization resolution or UV distortion while implicit volumetric representations lack scalability for large scenes due to the…

2023

RealGraph: A Multiview Dataset for 4D Real-world Context Graph Generation

ICCV 2023poster

In this paper, we propose a brand new scene understanding paradigm called "Context Graph Generation (CGG)", aiming at abstracting holistic semantic information in the complicated 4D world. The CGG task capitalizes on the calibrated multiview videos of a dynamic scene, and targets at recovering seman…

Cited by 1PDFcodeScholar
2023

Spatially and Spectrally Consistent Deep Functional Maps

ICCV 2023poster

Cycle consistency has long been exploited as a powerful prior for jointly optimizing maps within a collection of shapes. In this paper, we investigate its utility in the approaches of Deep Functional Maps, which are considered state-of-the-art in non-rigid shape matching. We first justify that under…

Cited by 27PDFcodeScholar
2022

ElasticMVS: Learning elastic part representation for self-supervised multi-view stereopsis

NeurIPS 2022accept

Self-supervised multi-view stereopsis (MVS) attracts increasing attention for learning dense surface predictions from only a set of images without onerous ground-truth 3D training data for supervision. However, existing methods highly rely on the local photometric consistency, which fails to identif…

Cited by 9SourcePDFScholar
2019

OperatorNet: Recovering 3D Shapes From Difference Operators

ICCV 2019poster

This paper proposes a learning-based framework for reconstructing 3D shapes from functional operators, compactly encoded as small-sized matrices. To this end we introduce a novel neural architecture, called OperatorNet, which takes as input a set of linear operators representing a shape and produces…

Cited by 18PDFcodeScholar