← Search

Zhonghua Zhai

4 accepted papers

2025

Universal Video Temporal Grounding with Generative Multi-modal Large Language Models

NeurIPS 2025poster

This paper presents a computational model for universal video temporal grounding, which accurately localizes temporal moments in videos based on natural language queries (e.g., questions or descriptions). Unlike existing methods that are often limited to specific video domains or durations, we prop…

Cited by 0SourcecodeScholar
2024

Turbo: Informativity-Driven Acceleration Plug-In for Vision-Language Large Models

ECCV 2024oral

"Vision-Language Large Models (VLMs) recently become primary backbone of AI, due to the impressive performance. However, their expensive computation costs, i.e., throughput and delay, impede potentials in the real-world scenarios. To achieve acceleration for VLMs, most existing methods focus on the…

Cited by 9SourcePDFScholar
2024

Wear-Any-Way: Manipulable Virtual Try-on via Sparse Correspondence Alignment

ECCV 2024poster

"This paper introduces a novel framework for virtual try-on, termed . Different from previous methods, is a customizable solution. Besides generating high-fidelity results, our method supports users to precisely manipulate the wearing style. To achieve this goal, we first construct a strong pipeline…

2021

Demodalizing Face Recognition with Synthetic Samples

AAAI 2021technical

Using data generated by generative adversarial networks or three-dimensional (3D) technology for face recognition training is a theoretically reasonable solution to the problems of unbalanced data distributions and data scarcity. However, due to the modal difference between synthetic data and real d…