← Search

Jingyu Wu

3 accepted papers

2026

OmniGround: A Comprehensive Spatio-Temporal Grounding Benchmark for Real-World Complex Scenarios

CVPR 2026

Spatio-Temporal Video Grounding (STVG) aims to localize target objects in videos based on natural language descriptions. While Multimodal Large Language Models have shown promise, a significant gap remains between current models and real-world demands involving diverse objects and complex queries. W

Cited by 0SourceScholar
2023

Preserving Structural Consistency in Arbitrary Artist and Artwork Style Transfer

AAAI 2023technical

Deep generative models are effective in style transfer. Previous methods learn one or several specific artist-style from a collection of artworks. These methods not only homogenize the artist-style of different artworks of the same artist but also lack generalization for the unseen artists. To solv…

Cited by 5SourcePDFScholar
2021

Image Synthesis From Layout With Locality-Aware Mask Adaption

ICCV 2021poster

This paper is concerned with synthesizing images conditioned on a layout (a set of bounding boxes with object categories). Existing works construct a layout-mask-image pipeline. Object masks are generated separately and mapped to bounding boxes to form a whole semantic segmentation mask (layout-to-m…

Cited by 80PDFcodeScholar