← Search

Junjie Fei

5 accepted papers

2025

Document Haystacks: Vision-Language Reasoning Over Piles of 1000+ Documents

CVPR 2025poster

Large multimodal models (LMMs) have achieved impressive progress in vision-language understanding, yet they face limitations in real-world applications requiring complex reasoning over a large number of images. Existing benchmarks for multi-image question-answering are limited in scope, each questio…

2025

Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description

ICCV 2025poster

In this paper, we introduce Part-Aware Point Grounded Description (PaPGD), a challenging task aimed at advancing 3D multimodal learning for fine-grained, part-aware segmentation grounding and detailed explanation of 3D objects. Existing 3D datasets largely focus on either vision-only part segmentati…

Cited by 0SourcePDFScholar
2025

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks

NeurIPS 2025poster

Large multimodal models (LMMs) have shown remarkable progress in audiovisual understanding, yet they struggle with real-world scenarios that require complex reasoning across extensive video collections. Existing benchmarks for video question answering remain limited in scope, typically involving one…

Cited by 0SourceScholar
2025

WikiAutoGen: Towards Multi-Modal Wikipedia-Style Article Generation

ICCV 2025poster

Knowledge discovery and collection are intelligence-intensive tasks that traditionally require significant human effort to ensure high-quality outputs. Recent research has explored multi-agent frameworks for automating Wikipedia-style article generation by retrieving and synthesizing information fro…

2023

Transferable Decoding with Visual Entities for Zero-Shot Image Captioning

ICCV 2023poster

Image-to-text generation aims to describe images using natural language. Recently, zero-shot image captioning based on pre-trained vision-language models (VLMs) and large language models (LLMs) has made significant progress. However, we have observed and empirically demonstrated that these methods a…

Cited by 53PDFcodeScholar