← Search

Jie Min

4 accepted papers

2025

VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation

CVPR 2025poster

Current large multimodal models (LMMs) face significant challenges in processing and comprehending long-duration or high-resolution videos, which is mainly due to the lack of high-quality datasets. To address this issue from a data-centric perspective, we propose VISTA, a simple yet effective video…

Cited by 4SourcePDFScholar
2020

Learning Object Placement by Inpainting for Compositional Data Augmentation

ECCV 2020poster

We study the problem of common sense placement of visual objects in an image.  This involves multiple aspects of visual recognition: the instance segmentation of the scene, 3D layout, and common knowledge of how objects are placed and where objects are moving in the 3D scene. This seemingly simple t…

Cited by 78SourcePDFScholar
2020

Nested Scale-Editing for Conditional Image Synthesis

CVPR 2020poster

We propose an image synthesis approach that provides stratified navigation in the latent code space. With a tiny amount of partial or very low-resolution image, our approach can consistently out-perform state-of-the-art counterparts in terms of generating the closest sampled image to the ground trut…

Cited by 7PDFScholar
2019

Liquid Warping GAN: A Unified Framework for Human Motion Imitation, Appearance Transfer and Novel View Synthesis

ICCV 2019poster

We tackle the human motion imitation, appearance transfer, and novel view synthesis within a unified framework, which means that the model once being trained can be used to handle all these tasks. The existing task-specific methods mainly use 2D keypoints (pose) to estimate the human body structure.…

Cited by 333PDFcodeScholar