← Search

Liming Zhao

12 accepted papers

2025

ContextHOI: Spatial Context Learning for Human-Object Interaction Detection

AAAI 2025technical

Spatial contexts, such as the backgrounds and surroundings, are considered critical in Human-Object Interaction (HOI) recognition, especially when the instance-centric foreground is blurred or occluded. Recent advancements in HOI detectors are usually built upon detection transformer pipelines. Whil…

Cited by 1SourcePDFScholar
2025

Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models

CVPR 2025poster

Recent Multi-modal Large Language Models (MLLMs) have been challenged by the computational overhead resulting from massive video frames, often alleviated through compression strategies. However, the visual content is not equally contributed to user instructions, existing strategies (e.g., average po…

2025

Improved Video VAE for Latent Video Diffusion Model

CVPR 2025poster

Variational Autoencoder (VAE) aims to compress pixel data into low-dimensional latent space, playing an important role in OpenAI's Sora and other latent video diffusion generation models. While most existing video VAEs inflate a pre-trained image VAE into the 3D causal structure for temporal-spatial…

2025

Orchestrating the Symphony of Prompt Distribution Learning for Human-Object Interaction Detection

AAAI 2025technical

Human-object interaction (HOI) detectors with popular query-transformer architecture have achieved promising performance. However, accurately identifying uncommon visual patterns and distinguishing between ambiguous HOIs continue to be difficult for them. We observe that these difficulties may arise…

Cited by 1SourcePDFScholar
2024

FuseTeacher: Modality-fused Encoders are Strong Vision Supervisors

ECCV 2024poster

"Learning visual representation with image-text datasets attracts a lot of attention in recent years. Existing approaches primarily rely on cross-modality supervision, and incorporate intra-modality supervision if necessary. They overlook the potential benefits of modality-fused supervision. Since m…

2024

Large Brain Model for Learning Generic Representations with Tremendous EEG Data in BCI

ICLR 2024spotlight

The current electroencephalogram (EEG) based deep learning models are typically designed for specific datasets and applications in brain-computer interaction (BCI), limiting the scale of the models and thus diminishing their perceptual capabilities and generalizability. Recently, Large Language Mode…

2023

MomentDiff: Generative Video Moment Retrieval from Random to Real

NeurIPS 2023poster

Video moment retrieval pursues an efficient and generalized solution to identify the specific temporal segments within an untrimmed video that correspond to a given language description. To achieve this goal, we provide a generative diffusion-based framework called MomentDiff, which simulates a typi…

2023

Progressive Spatio-Temporal Prototype Matching for Text-Video Retrieval

ICCV 2023oral

The performance of text-video retrieval has been significantly improved by vision-language cross-modal learning schemes. The typical solution is to directly align the global video-level and sentence-level features during learning, which would ignore the intrinsic video-text relations, i.e., a text…

Cited by 46PDFcodeScholar
2023

RLEG: Vision-Language Representation Learning with Diffusion-based Embedding Generation

ICML 2023poster

Vision-language representation learning models (e.g., CLIP) have achieved state-of-the-art performance on various downstream tasks, which usually need large-scale training data to learn discriminative representation. Recent progress on generative diffusion models (e.g., DALL-E 2) has demonstrated th…

Cited by 12SourcePDFScholar
2020

Weakly Supervised Learning with Side Information for Noisy Labeled Images

ECCV 2020poster

In many real-world datasets, like WebVision, the performance of DNN based classier is often limited by the noisy labeled data. To tackle this problem, some image related side information, such as captions and tags, often reveal underlying relationships across images. In this paper, we present an eff…

Cited by 60SourcePDFScholar
2018

Geometry-Aware Scene Text Detection With Instance Transformation Network

CVPR 2018poster

Localizing text in the wild is challenging in the situations of complicated geometric layout of the targets like random orientation and large aspect ratio. In this paper, we propose a geometry-aware modeling approach tailored for scene text representation with an end-to-end learning scheme. In our a…

Cited by 111SourcePDFScholar
2017

Deeply-Learned Part-Aligned Representations for Person Re-Identification

ICCV 2017poster

In this paper, we address the problem of person re-identification, which refers to associating the persons captured from different cameras. We propose a simple yet effective human part-aligned representation for handling the body part misalignment problem. Our approach decomposes the human body into…

Cited by 951PDFScholar