← Search

Junshi Huang

15 accepted papers

2026

UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text Editing

AAAI 2026technical

With the rapid advancement of image generation, visual text editing using natural language instructions has received increasing attention. The main challenge of this task is to fully understand the instruction and reference image, and thus generate visual text that is style-consistent with the imag

Cited by 0SourcePDFScholar
2024

Tuning-Free Inversion-Enhanced Control for Consistent Image Editing

AAAI 2024technical

Consistent editing of real images is a challenging task, as it requires performing non-rigid edits (e.g., changing postures) to the main objects in the input image without changing their identity or attributes. To guarantee consistent attributes, some existing methods fine-tune the entire model or t…

Cited by 12SourcePDFScholar
2023

Bridging Search Region Interaction With Template for RGB-T Tracking

CVPR 2023poster

RGB-T tracking aims to leverage the mutual enhancement and complement ability of RGB and TIR modalities for improving the tracking process in various scenarios, where cross-modal interaction is the key component. Some previous methods concatenate the RGB and TIR search region features directly to pe…

2023

Divide and Adapt: Active Domain Adaptation via Customized Learning

CVPR 2023highlight

Active domain adaptation (ADA) aims to improve the model adaptation performance by incorporating the active learning (AL) techniques to label a maximally-informative subset of target samples. Conventional AL methods do not consider the existence of domain shift, and hence, fail to identify the truly…

2023

Enriching Phrases with Coupled Pixel and Object Contexts for Panoptic Narrative Grounding

IJCAI 2023poster

Panoptic narrative grounding (PNG) aims to segment things and stuff objects in an image described by noun phrases of a narrative caption. As a multimodal task, an essential aspect of PNG is the visual-linguistic interaction between image and caption. The previous two-stage method aggregates visual c…

Cited by 5SourcePDFScholar
2023

Masked Auto-Encoders Meet Generative Adversarial Networks and Beyond

CVPR 2023poster

Masked Auto-Encoder (MAE) pretraining methods randomly mask image patches and then train a vision Transformer to reconstruct the original pixels based on the unmasked patches. While they demonstrates impressive performance for downstream vision tasks, it generally requires a large amount of training…

Cited by 20SourcePDFScholar
2023

Uncertainty-Aware Image Captioning

AAAI 2023technical

It is well believed that the higher uncertainty in a word of the caption, the more inter-correlated context information is required to determine it. However, current image captioning methods usually consider the generation of all words in a sentence sequentially and equally. In this paper, we propos…

Cited by 19SourcePDFScholar
2022

Adaptive Spatial-BCE Loss for Weakly Supervised Semantic Segmentation

ECCV 2022poster

"For Weakly-Supervised Semantic Segmentation (WSSS) with image-level annotation, mostly relies on the classification network to generate initial segmentation pseudo-labels. However, the optimization target of classification networks usually neglects the discrimination between different pixels, like…

2022

Language-Bridged Spatial-Temporal Interaction for Referring Video Object Segmentation

CVPR 2022poster

Referring video object segmentation aims to predict foreground labels for objects referred by natural language expressions in videos. Previous methods either depend on 3D ConvNets or incorporate additional 2D ConvNets as encoders to extract mixed spatial-temporal features. However, these methods suf…

Cited by 74PDFcodeScholar
2021

Embedded Discriminative Attention Mechanism for Weakly Supervised Semantic Segmentation

CVPR 2021poster

Weakly Supervised Semantic Segmentation (WSSS) with image-level annotation uses class activation maps from the classifier as pseudo-labels for semantic segmentation. However, such activation maps usually highlight the local discriminative regions rather than the whole object, which deviates from the…

Cited by 178PDFcodeScholar
2021

Rethinking BiSeNet for Real-Time Semantic Segmentation

CVPR 2021poster

BiSeNet has been proved to be a popular two-stream network for real-time segmentation. However, its principle of adding an extra path to encode spatial information is time-consuming, and the backbones borrowed from pretrained tasks, e.g., image classification, may be inefficient for image segmentati…

Cited by 814PDFcodeScholar
2017

More Is Less: A More Complicated Network With Less Inference Complexity

CVPR 2017poster

In this paper, we present a novel and general network structure towards accelerating the inference process of convolutional neural networks, which is more complicated in network structure yet with less inference complexity. The core idea is to equip each original convolutional layer with another low…

Cited by 390PDFcodeScholar
2015

Cross-Domain Image Retrieval With a Dual Attribute-Aware Ranking Network

ICCV 2015poster

We address the problem of cross-domain image retrieval, considering the following practical application: given a user photo depicting a clothing image, our goal is to retrieve the same or attribute-similar clothing items from online shopping stores. This is a challenging problem due to the large dis…

Cited by 541PDFScholar
2015

Deep Domain Adaptation for Describing People Based on Fine-Grained Clothing Attributes

CVPR 2015poster

We address the problem of describing people based on fine-grained clothing attributes. This is an important problem for many practical applications, such as identifying target suspects or finding missing people based on detailed clothing descriptions in surveillance videos or consumer photos. We app…

Cited by 338SourcePDFScholar