← Search

jiazhou zhou

6 accepted papers

2026

T-Rex-Omni: Integrating Negative Visual Prompt in Generic Object Detection

AAAI 2026technical

Object detection methods have evolved from closed-set to open-set paradigms over the years. Current open-set object detectors, however, remain constrained by their exclusive reliance on positive indicators based on given prompts like text descriptions or visual exemplars. This positive-only paradigm

Cited by 0SourcePDFScholar
2024

Centering the Value of Every Modality: Towards Efficient and Resilient Modality-agnostic Semantic Segmentation

ECCV 2024poster

"Fusing an arbitrary number of modalities is vital for achieving robust multi-modal fusion of semantic segmentation yet remains less explored to date. Recent endeavors regard RGB modality as the center and the others as the auxiliary, yielding an asymmetric architecture with two branches. However, t…

Cited by 11SourcePDFScholar
2024

ExACT: Language-guided Conceptual Reasoning and Uncertainty Estimation for Event-based Action Recognition and More

CVPR 2024highlight

Event cameras have recently been shown beneficial for practical vision tasks such as action recognition thanks to their high temporal resolution power efficiency and reduced privacy concerns. However current research is hindered by 1) the difficulty in processing events because of their prolonged du…

Cited by 23SourcePDFScholar
2024

LaSe-E2V: Towards Language-guided Semantic-aware Event-to-Video Reconstruction

NeurIPS 2024poster

Event cameras harness advantages such as low latency, high temporal resolution, and high dynamic range (HDR), compared to standard cameras. Due to the distinct imaging paradigm shift, a dominant line of research focuses on event-to-video (E2V) reconstruction to bridge event-based and standard comput…

Cited by 2SourcePDFScholar
2024

UniBind: LLM-Augmented Unified and Balanced Representation Space to Bind Them All

CVPR 2024poster

We present UniBind a flexible and efficient approach that learns a unified representation space for seven diverse modalities-- images text audio point cloud thermal video and event data. Existing works eg. ImageBind treat the image as the central modality and build an image-centered representation s…

Cited by 13SourcePDFScholar