← Search

Xiangteng He

10 accepted papers

2026

From Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific Literature

CVPR 2026

There is growing interest in biomedical vision--language models trained on scientific literature. However, most pipelines compress rich multi-panel figures and long captions into coarse figure-level pairs, discarding the fine-grained correspondences clinicians rely on when zooming into local structu

Cited by 0SourceScholar
2026

InvAD: Inversion-based Reconstruction-Free Anomaly Detection with Diffusion Models

CVPR 2026

Despite the remarkable success, recent reconstruction-based anomaly detection (AD) methods via diffusion modeling still involve fine-grained noise-strength tuning and computationally expensive multi-step denoising, leading to a fundamental tension between fidelity and efficiency. In this paper, we p

Cited by 0SourcecodeScholar
2026

To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models

ICLR 2026poster

Large Vision Language Models (LVLMs) have recently emerged as powerful architectures capable of understanding and reasoning over both visual and textual information. These models typically rely on two key components: a Vision Transformer (ViT) and a Large Language Model (LLM). ViT encodes visual con…

Cited by 0SourceScholar
2025

ICE-Bench: A Unified and Comprehensive Benchmark for Image Creating and Editing

ICCV 2025poster

Image generation has witnessed significant advancements in the past few years. However, evaluating the performance of image generation models remains a formidable challenge. In this paper, we propose ICE-Bench, a unified and comprehensive benchmark designed to rigorously assess image generation mode…

2025

VarCMP: Adapting Cross-Modal Pre-Training Models for Video Anomaly Retrieval

AAAI 2025technical

Video anomaly retrieval (VAR) aims to retrieve pertinent abnormal or normal videos from collections of untrimmed and long videos through cross-modal requires such as textual descriptions and synchronized audios. Cross-modal pre-training (CMP) models, by pre-training on large-scale cross-modal pairs,…

Cited by 0SourcePDFScholar
2024

FashionERN: Enhance-and-Refine Network for Composed Fashion Image Retrieval

AAAI 2024technical

The goal of composed fashion image retrieval is to locate a target image based on a reference image and modified text. Recent methods utilize symmetric encoders (e.g., CLIP) pre-trained on large-scale non-fashion datasets. However, the input for this task exhibits an asymmetric nature, where the ref…

Cited by 5SourcePDFScholar
2024

FineFMPL: Fine-grained Feature Mining Prompt Learning for Few-Shot Class Incremental Learning

IJCAI 2024poster

Few-Shot Class Incremental Learning (FSCIL) aims to continually learn new classes with few training samples without forgetting already learned old classes. Existing FSCIL methods generally fix the backbone network in incremental sessions to achieve a balance between suppressing forgetting old classe…

2023

PosterLayout: A New Benchmark and Approach for Content-Aware Visual-Textual Presentation Layout

CVPR 2023poster

Content-aware visual-textual presentation layout aims at arranging spatial space on the given canvas for pre-defined elements, including text, logo, and underlay, which is a key to automatic template-free creative graphic design. In practical applications, e.g., poster designs, the canvas is origina…

2023

Scanning Only Once: An End-to-end Framework for Fast Temporal Grounding in Long Videos

ICCV 2023poster

Video temporal grounding aims to pinpoint a video segment that matches the query description. Despite the recent advance in short-form videos (e.g., in minutes), temporal grounding in long videos (e.g., in hours) is still at its early stage. To address this challenge, a common practice is to employ…

Cited by 18PDFcodeScholar