← Search

Zimo Liu

12 accepted papers

2026

DMGINE: Day-Memory Guided Nighttime Image Enhancement for Dynamic Traffic Scenes

AAAI 2026technical

We introduce Daytime-Memory Guided Nighttime Image Enhancement (DMGNIE) framework, the first framework that turns long-running daytime surveillance videos of a single intersection into persistent “daytime memory” to guide nighttime image enhancement in traffic scenes. Our key insight is simple yet p

Cited by 0SourcePDFScholar
2026

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining

ICML 2026poster

Large language models (LLMs) have achieved remarkable breakthroughs across various applications. However, their architectures remain inefficient in pretraining due to two main limitations: (i) self-attention lacks an explicit inductive bias for locality, leading to redundant modeling of sequence-int…

Cited by 0SourceScholar
2025

An Exploration with Entropy Constrained 3D Gaussians for 2D Video Compression

ICLR 2025poster

3D Gaussian Splatting (3DGS) has witnessed its rapid development in novel view synthesis, which attains high quality reconstruction and real-time rendering. At the same time, there is still a gap before implicit neural representation (INR) can become a practical compressor due to the lack of stream…

2025

DM-Adapter: Domain-Aware Mixture-of-Adapters for Text-Based Person Retrieval

AAAI 2025technical

Text-based person retrieval (TPR) has gained significant attention as a fine-grained and challenging task that closely aligns with practical applications. Tailoring CLIP to person domain is now a emerging research topic due to the abundant knowledge of vision-language pretraining, but challenges sti…

2025

Pre-Trained Vision-Language Models as Noisy Partial Annotators

AAAI 2025technical

In noisy partial label learning, each training sample is associated with a set of candidate labels, and the ground-truth label may be contained within this set. With the emergence of powerful pre-trained vision-language models, e.g. CLIP, it is natural to consider using these models to automatically…

2024

Clip-Based Synergistic Knowledge Transfer for text-based Person Retrieval

ICASSP 2024accepted

Text-based Person Retrieval (TPR) aims to retrieve the target person images given a textual query. The primary challenge lies in bridging the substantial gap between vision and language modalities, especially when dealing with limited large-scale datasets. In this paper, we introduce a CLIP-based Sy…

Cited by 0SourceScholar
2024

Controller-Guided Partial Label Consistency Regularization with Unlabeled Data

AAAI 2024technical

Partial label learning (PLL) learns from training examples each associated with multiple candidate labels, among which only one is valid. In recent years, benefiting from the strong capability of dealing with ambiguous supervision and the impetus of modern data augmentation methods, consistency regu…

Cited by 3SourcePDFScholar
2024

M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation

NeurIPS 2024poster

This paper presents M$^3$GPT, an advanced $\textbf{M}$ultimodal, $\textbf{M}$ultitask framework for $\textbf{M}$otion comprehension and generation. M$^3$GPT operates on three fundamental principles. The first focuses on creating a unified representation space for various motion-relevant modalities…

2023

Lifelong Person Re-identification via Knowledge Refreshing and Consolidation

AAAI 2023technical

Lifelong person re-identification (LReID) is in significant demand for real-world development as a large amount of ReID data is captured from diverse locations over time and cannot be accessed at once inherently. However, a key challenge for LReID is how to incrementally preserve old knowledge and g…

2019

Deep Reinforcement Active Learning for Human-in-the-Loop Person Re-Identification

ICCV 2019oral

Most existing person re-identification(Re-ID) approaches achieve superior results based on the assumption that a large amount of pre-labelled data is usually available and can be put into training phrase all at once. However, this assumption is not applicable to most real-world deployment of the Re-…

Cited by 118PDFScholar