← Search

Junting Dong

19 accepted papers

2026

DanceTogether: Generating Interactive Multi-Person Video without Identity Drifting

ICLR 2026poster

Controllable video generation (CVG) has advanced rapidly, yet current systems falter when more than one actor must move, interact, and exchange positions under noisy control signals. We address this gap with DanceTogether, the first end-to-end diffusion framework that turns a single reference image…

Cited by 0SourceScholar
2026

IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation

CVPR 2026

Recent research on medical MLLMs has shifted its focus from image-level understanding to fine-grained, pixel-level comprehension. Although segmentation serves as the foundation for pixel-level understanding, existing approaches face two major challenges. First, they introduce implicit segmentation t

Cited by 0SourcecodeScholar
2025

ARMO: Autoregressive Rigging for Multi-Category Objects

ICCV 2025poster

Recent advancements in large-scale generative models have significantly improved the quality and diversity of 3D shape generation. However, most existing methods focus primarily on generating static 3D models, overlooking the potential dynamic nature of certain shapes, such as humanoids, animals, an…

Cited by 0SourcePDFScholar
2025

DRiVE: Diffusion-based Rigging Empowers Generation of Versatile and Expressive Characters

CVPR 2025poster

Recent advances in generative models have enabled high-quality 3D character reconstruction from multi-modal. However, animating these generated characters remains a challenging task, especially for complex elements like garments and hair, due to the lack of large-scale datasets and effective rigging…

2025

GAS: Generative Avatar Synthesis from a Single Image

ICCV 2025poster

We present a unified and generalizable framework for synthesizing view-consistent and temporally coherent avatars from a single image, addressing the challenging task of single-image avatar generation. Existing diffusion-based methods often condition on sparse human templates (e.g., depth or normal…

2025

Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data

ICCV 2025poster

Generating diverse and natural human motion sequences based on textual descriptions constitutes a fundamental and challenging research area within the domains of computer vision, graphics, and robotics. Despite significant advancements in this field, current methodologies often face challenges regar…

2025

HOMIE: Humanoid Loco-Manipulation with Isomorphic Exoskeleton Cockpit

RSS 2025poster

Current humanoid teleoperation systems either lack reliable low-level control policies, or struggle to acquire accurate whole-body control commands, making it difficult to teleoperate humanoids for loco-manipulation tasks. To solve these issues, we propose HOMIE, a novel humanoid teleoperation syste…

Cited by 9PDFScholar
2025

Horizon-GS: Unified 3D Gaussian Splatting for Large-Scale Aerial-to-Ground Scenes

CVPR 2025poster

Seamless integration of both aerial and street view images remains a significant challenge in neural scene reconstruction and rendering. Existing methods predominantly focus on single domain, limiting their applications in immersive environments, which demand extensive free view exploration with lar…

Cited by 1SourcePDFScholar
2025

SIGMAN: Scaling 3D Human Gaussian Generation with Millions of Assets

ICCV 2025poster

3D human digitization has long been a highly pursued yet challenging task. Existing methods aim to generate high-quality 3D digital humans from single or multiple views, but remain primarily constrained by current paradigms and the scarcity of 3D human assets. Specifically, recent approaches fall in…

Cited by 0SourcePDFScholar
2025

ScaMo: Exploring the Scaling Law in Autoregressive Motion Generation Model

CVPR 2025poster

The scaling law has been validated in various domains, such as natural language processing (NLP) and massive computer vision tasks; however, its application to motion generation remains largely unexplored. In this paper, we introduce a scalable motion generation framework that includes the motion to…

Cited by 6SourcePDFScholar
2024

Capturing Closely Interacted Two-Person Motions with Reaction Priors

CVPR 2024poster

In this paper we focus on capturing closely interacted two-person motions from monocular videos an important yet understudied topic. Unlike less-interacted motions closely interacted motions contain frequently occurring inter-human occlusions which pose significant challenges to existing capturing a…

Cited by 1SourcePDFScholar
2024

EpiDiff: Enhancing Multi-View Synthesis via Localized Epipolar-Constrained Diffusion

CVPR 2024poster

Generating multiview images from a single view facilitates the rapid generation of a 3D mesh conditioned on a single image. Recent methods that introduce 3D global representation into diffusion models have shown the potential to generate consistent multiviews but they have reduced generation speed a…

2023

NaviNeRF: NeRF-based 3D Representation Disentanglement by Latent Semantic Navigation

ICCV 2023poster

3D representation disentanglement aims to identify, decompose, and manipulate the underlying explanatory factors of 3D data, which helps AI fundamentally understand our 3D world. This task is currently under-explored and poses great challenges: (i) the 3D representations are complex and in general c…

Cited by 11PDFcodeScholar
2023

iVS-Net: Learning Human View Synthesis from Internet Videos

ICCV 2023poster

Recent advances in implicit neural representations make it possible to generate free-viewpoint videos of the human from sparse view images. To avoid the expensive training for each person, previous methods adopt the generalizable human model and demonstrate impressive results. However, these methods…

Cited by 6PDFScholar
2022

TotalSelfScan: Learning Full-body Avatars from Self-Portrait Videos of Faces, Hands, and Bodies

NeurIPS 2022accept

Recent advances in implicit neural representations make it possible to reconstruct a human-body model from a monocular self-rotation video. While previous works present impressive results of human body reconstruction, the quality of reconstructed face and hands are relatively low. The main reason i…

2021

Animatable Neural Radiance Fields for Modeling Dynamic Human Bodies

ICCV 2021poster

This paper addresses the challenge of reconstructing an animatable human model from a multi-view video. Some recent works have proposed to decompose a non-rigidly deforming scene into a canonical neural radiance field and a set of deformation fields that map observation-space points to the canonical…

Cited by 514PDFcodeScholar
2020

Motion Capture from Internet Videos

ECCV 2020poster

Recent advances in image-based human pose estimation make it possible to capture 3D human motion from a single RGB video. However, the inherent depth ambiguity and self-occlusion in a single view prohibit the recovery of as high-quality motion as multi-view reconstruction. While multi-view videos ar…

2019

Fast and Robust Multi-Person 3D Pose Estimation From Multiple Views

CVPR 2019poster

This paper addresses the problem of 3D pose estimation for multiple people in a few calibrated camera views. The main challenge of this problem is to find the cross-view correspondences among noisy and incomplete 2D pose predictions. Most previous methods address this challenge by directly reasoning…

Cited by 266PDFScholar