← Search

Liulei Li

14 accepted papers

2026

Beyond Frequency: Scoring-Driven Debiasing for Object Detection via Blueprint-Prompted Image Synthesis

ICLR 2026poster

This paper presents a generation-based debiasing framework for object detection. Prior debiasing methods are often limited by the representation diversity of samples, while naive generative augmentation often preserves the biases it aims to solve. Moreover, our analysis reveals that simply generatin…

Cited by 0SourcecodeScholar
2025

UNIALIGN: Scaling Multimodal Alignment within One Unified Model

CVPR 2025poster

We present UNIALIGN, a unified model to align an arbitrary number of modalities (\text e.g. , image, text, audio, 3D point cloud, etc.) through one encoder and a single training phase. Existing solutions typically employ distinct encoders for each modality, resulting in increased parameters as the…

2024

Clustering Propagation for Universal Medical Image Segmentation

CVPR 2024poster

Prominent solutions for medical image segmentation are typically tailored for automatic or interactive setups posing challenges in facilitating progress achieved in one task to another. This also necessitates separate models for each task duplicating both training time and parameters. To address abo…

2024

Human-Object Interaction Detection Collaborated with Large Relation-driven Diffusion Models

NeurIPS 2024poster

Prevalent human-object interaction (HOI) detection approaches typically leverage large-scale visual-linguistic models to help recognize events involving humans and objects. Though promising, models trained via contrastive learning on text-image pairs often neglect mid/low-level visual cues and strug…

Cited by 7SourcePDFScholar
2023

Boosting Video Object Segmentation via Space-Time Correspondence Learning

CVPR 2023poster

Current top-leading solutions for video object segmentation (VOS) typically follow a matching-based regime: for each query frame, the segmentation mask is inferred according to its correspondence to previously processed and the first annotated frames. They simply exploit the supervisory signals from…

2023

Large-Scale Person Detection and Localization Using Overhead Fisheye Cameras

ICCV 2023oral

Location determination finds wide applications in daily life. Instead of existing efforts devoted to localizing tourist photos captured by perspective cameras, in this article, we focus on developing person positioning solutions using overhead fisheye cameras. Such solutions are advantageous in larg…

Cited by 25PDFScholar
2023

Unified Mask Embedding and Correspondence Learning for Self-Supervised Video Segmentation

CVPR 2023poster

The objective of this paper is self-supervised learning of video object segmentation. We develop a unified framework which simultaneously models cross-frame dense correspondence for locally discriminative feature learning and embeds object-level context for target-mask decoding. As a result, it is a…

2022

Locality-Aware Inter- and Intra-Video Reconstruction for Self-Supervised Correspondence Learning

CVPR 2022poster

Our target is to learn visual correspondence from unlabeled videos. We develop LIIR, a locality-aware inter-and intra-video reconstruction framework that fills in three missing pieces, i.e., instance discrimination, location awareness, and spatial compactness, of self-supervised correspondence learn…

Cited by 54PDFcodeScholar