← Search

Zhong-Yu Li

5 accepted papers

2025

Towards RAW Object Detection in Diverse Conditions

CVPR 2025highlight

Existing object detection methods often consider sRGB input, which was compressed from RAW data using ISP originally designed for visualization. However, such compression might lose crucial information for detection, especially under complex light and weather conditions. We introduce the AODRaw data…

2025

VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning

ICCV 2025poster

Recent advances in diffusion models have significantly advanced image generation; however, existing models remain task-specific, limiting their efficiency and generalizability. While universal models attempt to address these limitations, they face critical challenges, including generalizable instruc…

2024

Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation

ICML 2024poster

Pre-trained vision-language models, e.g., CLIP, have been successfully applied to zero-shot semantic segmentation. Existing CLIP-based approaches primarily utilize visual features from the last layer to align with text embeddings, while they neglect the crucial information in intermediate layers tha…

2024

DFormer: Rethinking RGBD Representation Learning for Semantic Segmentation

ICLR 2024poster

We present DFormer, a novel RGB-D pretraining framework to learn transferable representations for RGB-D segmentation tasks. DFormer has two new key innovations: 1) Unlike previous works that encode RGB-D information with RGB pretrained backbone, we pretrain the backbone using image-depth pairs from…

Cited by 58SourcePDFScholar
2021

Global2Local: Efficient Structure Search for Video Action Segmentation

CVPR 2021poster

Temporal receptive fields of models play an important role in action segmentation. Large receptive fields facilitate the long-term relations among video clips while small receptive fields help capture the local details. Existing methods construct models with hand-designed receptive fields in layers.…

Cited by 97PDFcodeScholar