← Search

Xiaojuan Qi*

3 accepted papers

2024

Can OOD Object Detectors Learn from Foundation Models?

ECCV 2024poster

"Out-of-distribution (OOD) object detection is a challenging task due to the absence of open-set OOD data. Inspired by recent advancements in text-to-image generative models, such as Stable Diffusion, we study the potential of generative models trained on large-scale open-set data to synthesize OOD…

2024

Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models

ECCV 2024poster

"We introduce Groma, a Multimodal Large Language Model (MLLM) with grounded and fine-grained visual perception ability. Beyond holistic image understanding, Groma is adept at region-level tasks such as region captioning and visual grounding. Such capabilities are built upon a localized visual tokeni…

2024

Let the Avatar Talk using Texts without Paired Training Data

ECCV 2024poster

"This paper introduces text-driven talking avatar generation, a task that uses text to instruct both the generation and animation of an avatar. One significant obstacle in this task is the absence of paired text and talking avatar data for model training, limiting data-driven methodologies. To this…

Cited by 0SourcePDFScholar