← Search

Xiaojuan Wang

8 accepted papers

2025

Generative Inbetweening: Adapting Image-to-Video Models for Keyframe Interpolation

ICLR 2025poster

We present a method for generating video sequences with coherent motion between a pair of input keyframes. We adapt a pretrained large-scale image-to-video diffusion model (originally trained to generate videos moving forward in time from a single input image) for keyframe interpolation, i.e., to pr…

Cited by 7SourcePDFScholar
2025

VideoMiner: Iteratively Grounding Key Frames of Hour-Long Videos via Tree-based Group Relative Policy Optimization

ICCV 2025poster

Understanding hour-long videos with multi-modal large language models (MM-LLMs) enriches the landscape of human-centered AI applications. However, for end-to-end video understanding with LLMs, uniformly sampling video frames results in LLMs being overwhelmed by a vast amount of irrelevant informatio…

2024

Generative Powers of Ten

CVPR 2024highlight

We present a method that uses a text-to-image model to generate consistent content across multiple image scales enabling extreme semantic zooms into a scene e.g. ranging from a wide-angle landscape view of a forest to a macro shot of an insect sitting on one of the tree branches. We achieve this thr…

Cited by 5SourcePDFScholar
2022

QueryPose: Sparse Multi-Person Pose Regression via Spatial-Aware Part-Level Query

NeurIPS 2022accept

We propose a sparse end-to-end multi-person pose regression framework, termed QueryPose, which can directly predict multi-person keypoint sequences from the input image. The existing end-to-end methods rely on dense representations to preserve the spatial detail and structure for precise keypoint lo…

2022

Single-Stage Is Enough: Multi-Person Absolute 3D Pose Estimation

CVPR 2022poster

The existing multi-person absolute 3D pose estimation methods are mainly based on two-stage paradigm, i.e., top-down or bottom-up, leading to redundant pipelines with high computation cost. We argue that it is more desirable to simplify such two-stage paradigm to a single-stage one to promote both e…

Cited by 56PDFScholar
2015

Multi-Scale Learning for Low-Resolution Person Re-Identification

ICCV 2015poster

In real world person re-identification (re-id), images of people captured at very different resolutions from different locations need be matched. Existing re-id models typically normalise all person images to the same size. However, a low-resolution (LR) image contains much less information about a…

Cited by 196PDFScholar