← Search

Jianqi Chen

6 accepted papers

2026

PoseGAM: Robust Unseen Object Pose Estimation via Geometry-Aware Multi-View Reasoning

CVPR 2026

6D object pose estimation, which predicts the transformation of an object relative to the camera, remains challenging for unseen objects. Existing approaches typically rely on explicitly constructing feature correspondences between the query image and either the object model or template images. In t

Cited by 0SourcecodeScholar
2026

ShapeGen4D: Towards High Quality 4D Shape Generation from Videos

ICLR 2026poster

Video-conditioned 4D shape generation aims to recover time-varying 3D geometry and view-consistent appearance directly from an input video. In this work, we introduce a native video-to-4D shape generation framework that synthesizes a single dynamic 3D representation end-to-end from the video. Our…

Cited by 0SourceScholar
2025

Sitcom-Crafter: A Plot-Driven Human Motion Generation System in 3D Scenes

ICLR 2025poster

Recent advancements in human motion synthesis have focused on specific types of motions, such as human-scene interaction, locomotion or human-human interaction, however, there is a lack of a unified system capable of generating a diverse combination of motion types. In response, we introduce *Sitcom…

2025

V2M4: 4D Mesh Animation Reconstruction from a Single Monocular Video

ICCV 2025poster

We present V2M4, a novel 4D reconstruction method that directly generates a usable 4D mesh animation asset from a single monocular video. Unlike existing approaches that rely on priors from multi-view image and video generation models, our method is based on native 3D mesh generation models. Naively…

Cited by 0SourcePDFScholar
2024

Prototypical Information Bottlenecking and Disentangling for Multimodal Cancer Survival Prediction

ICLR 2024spotlight

Multimodal learning significantly benefits cancer survival prediction, especially the integration of pathological images and genomic data. Despite advantages of multimodal learning for cancer survival prediction, massive redundancy in multimodal data prevents it from extracting discriminative and co…

2023

OvarNet: Towards Open-Vocabulary Object Attribute Recognition

CVPR 2023poster

In this paper, we consider the problem of simultaneously detecting objects and inferring their visual attributes in an image, even for those with no manual annotations provided at the training stage, resembling an open-vocabulary scenario. To achieve this goal, we make the following contributions: (…