← Search

Jiajin Tang

12 accepted papers

2026

Chart Deep Research in LVLMs via Parallel Relative Policy Optimization

ICLR 2026poster

With the rapid advancement of data science, charts have evolved from simple numerical presentation tools to essential instruments for insight discovery and decision-making support. However, current chart data intelligence exhibits significant limitations in deep research capabilities, with existing…

Cited by 0SourceScholar
2025

Discovering Compositional Hallucinations in LVLMs

NeurIPS 2025poster

Large language models (LLMs) and vision-language models (LVLMs) have driven the paradigm shift towards general-purpose foundation models. However, both of them are prone to hallucinations, which compromise their factual accuracy and reliability. While existing research primarily focuses on isolated…

Cited by 0SourceScholar
2025

Rethinking Query-based Transformer for Continual Image Segmentation

CVPR 2025poster

Class-incremental/Continual image segmentation (CIS) aims to train an image segmenter in stages, where the set of available categories differs at each stage. To leverage the built-in objectness of query-based transformers, which mitigates catastrophic forgetting of mask proposals, current methods of…

2025

Sim-DETR: Unlock DETR for Temporal Sentence Grounding

ICCV 2025poster

Temporal sentence grounding aims to identify exact moments in a video that correspond to a given textual query, typically addressed with detection transformer (DETR) solutions. However, we find that typical strategies designed to enhance DETR do not improve, and may even degrade, its performance in…

Cited by 0SourcePDFScholar
2025

Why LVLMs Are More Prone to Hallucinations in Longer Responses: The Role of Context

ICCV 2025poster

Large Vision-Language Models (LVLMs) have made significant progress in recent years but are also prone to hallucination issues. They exhibit more hallucinations in longer, free-form responses, often attributed to accumulated uncertainties. In this paper, we ask: Does increased hallucination result s…

2024

Part2Object: Hierarchical Unsupervised 3D Instance Segmentation

ECCV 2024poster

"Unsupervised 3D instance segmentation aims to segment objects from a 3D point cloud without any annotations. Existing methods face the challenge of either too loose or too tight clustering, leading to under-segmentation or over-segmentation. To address this issue, we propose Part2Object, hierarchic…

2023

Contrastive Grouping With Transformer for Referring Image Segmentation

CVPR 2023poster

Referring image segmentation aims to segment the target referent in an image conditioning on a natural language expression. Existing one-stage methods employ per-pixel classification frameworks, which attempt straightforwardly to align vision and language at the pixel level, thus failing to capture…

2023

DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multimodal Reasoning in Language Models

NeurIPS 2023poster

A long-standing goal of AI systems is to perform complex multimodal reasoning like humans. Recently, large language models (LLMs) have made remarkable strides in such multi-step reasoning on the language modality solely by leveraging the chain of thought (CoT) to mimic human thinking. However, the t…

Cited by 100SourcePDFScholar