← Search

Jinxia Zhang

8 accepted papers

2026

RefChess: Monte-Carlo Move Selection for Zero-Shot Referring Image Segmentation

ICML 2026poster

Recent advances in zero-shot referring image segmentation (RIS), driven by foundation models such as SAM and CLIP, have improved cross-modal alignment between visual regions and natural language expressions. Nevertheless, selecting the correct segmentation proposal remains challenging, as existing m…

Cited by 0SourceScholar
2025

CoF: Coarse to Fine-Grained Image Understanding for Multi-modal Large Language Models

ICASSP 2025accepted

The impressive performance of Large Language Model (LLM) has prompted researchers to develop Multi-modal LLM (MLLM), which has shown great potential for various multi-modal tasks. However, current MLLM often struggles to effectively address fine-grained multi-modal challenges. We argue that this lim…

Cited by 0SourceScholar
2025

Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples

AAAI 2025technical

Existing Vision-Language Pretraining (VLP) methods have achieved remarkable improvements across a variety of vision-language tasks, confirming their effectiveness in capturing coarse-grained semantic correlations. However, their capability for fine-grained understanding, which is critical for many…

Cited by 2SourcePDFScholar
2025

FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training

ICASSP 2025accepted

Large-scale Vision-Language Pre-training (VLP) has demonstrated remarkable success in the general domain. However, in the fashion domain, items are distinguished by fine-grained attributes such as texture and material, which are crucial for tasks such as retrieval. Existing models often fail to take…

Cited by 0SourceScholar
2025

MADiff: Text-Guided Fashion Image Editing with Mask Prediction and Attention-Enhanced Diffusion

ICASSP 2025accepted

Text-guided image editing model has achieved great success in general domain. However, directly applying these models to the fashion domain may encounter two issues: (1) Inaccurate localization of editing region; (2) Weak editing magnitude. To address these issues, the MADiff model is proposed. Spec…

Cited by 0SourceScholar
2025

MQM-APE: Toward High-Quality Error Annotation Predictors with Automatic Post-Editing in LLM Translation Evaluators

COLING 2025main

Large Language Models (LLMs) have shown significant potential as judges for Machine Translation (MT) quality assessment, providing both scores and fine-grained feedback. Although approaches such as GEMBA-MQM have shown state-of-the-art performance on reference-free evaluation, the predicted errors d…

2025

Runaway is Ashamed, But Helpful: On the Early-Exit Behavior of Large Language Model-based Agents in Embodied Environments

EMNLP 2025

Agents powered by large language models (LLMs) have demonstrated strong planning and decision-making capabilities in complex embodied environments. However, such agents often suffer from inefficiencies in multi-turn interactions, frequently trapped in repetitive loops or issuing ineffective commands

2025

Shifting Spotlight for Co-supervision: A Simple yet Efficient Single-branch Network to See Through Camouflage

ICASSP 2025accepted

Camouflaged object detection (COD) remains a challenging task in computer vision. Existing methods often resort to additional branches for edge supervision, incurring substantial computational costs. To address this, we propose the Co-Supervised Spotlight Shifting Network (CS<sup xmlns:mml="http://w…

Cited by 0SourceScholar