← Search

Qiuyu Kong

1 accepted papers

2026

StructXLIP: Enhancing Vision-language Models with Multimodal Structural Cues

CVPR 2026

Edge-based representations are fundamental cues for visual understanding, a principle rooted in early vision research and still central today. We extend this principle to vision-language alignment, showing that isolating and aligning structural cues across modalities can greatly benefit fine-tuning

Cited by 0SourcecodeScholar