← Search

Mengxue Kang

3 accepted papers

2025

Enhancing Image Editing with Chain-of-Thought Reasoning and Multimodal Large Language Models

ICASSP 2025accepted

Image editing in our daily lives often requires models to first understand user’s intention and then proceed with the editing. Despite significant advancements in image editing technology, understanding and executing complex instructions remains a substantial challenge. Existing image editing models…

Cited by 0SourceScholar
2024

Unifying Visual and Vision-Language Tracking via Contrastive Learning

AAAI 2024technical

Single object tracking aims to locate the target object in a video sequence according to the state specified by different modal references, including the initial bounding box (BBOX), natural language (NL), or both (NL+BBOX). Due to the gap between different modalities, most existing trackers are des…

2023

Alleviating Catastrophic Forgetting of Incremental Object Detection via Within-Class and Between-Class Knowledge Distillation

ICCV 2023poster

Incremental object detection (IOD) task requires a model to learn continually from newly added data. However, directly fine-tuning a well-trained detection model on a new task will sharply decrease the performance on old tasks, which is known as catastrophic forgetting. Knowledge distillation, inclu…

Cited by 15PDFScholar