← Search

Wenjian Huang

9 accepted papers

2026

TripleFDS: Triple Feature Disentanglement and Synthesis for Scene Text Editing

AAAI 2026technical

Scene Text Editing (STE) aims to naturally modify text in images while preserving visual consistency, the decisive factors of which can be divided into three parts, i.e., text style, text content, and background. Previous methods have struggled with incomplete disentanglement of editable attributes,

Cited by 0SourcePDFScholar
2025

Open-Det: An Efficient Learning Framework for Open-Ended Detection

ICML 2025poster

Open-Ended object Detection (OED) is a novel and challenging task that detects objects and generates their category names in a free-form manner, without requiring additional vocabularies during inference. However, the existing OED models, such as GenerateU, require large-scale datasets for training,…

2025

Unsupervised Part Discovery via Descriptor-Based Masked Image Restoration with Optimized Constraints

ICCV 2025poster

Part-level features are crucial for image understanding, but few studies focus on them because of the lack of fine-grained labels. Although unsupervised part discovery can eliminate the reliance on labels, most of them cannot maintain robustness across various categories and scenarios, which restric…

2024

MLP-DINO: Category Modeling and Query Graphing with Deep MLP for Object Detection

IJCAI 2024poster

Popular transformer-based detectors detect objects in a one-to-one manner, where both the bounding box and category of each object are predicted only by the single query, leading to the box-sensitive category predictions. Additionally, the initialization of positional queries solely based on the pre…

2023

Cross Contrasting Feature Perturbation for Domain Generalization

ICCV 2023poster

Domain generalization (DG) aims to learn a robust model from source domains that generalize well on unseen target domains. Recent studies focus on generating novel domain samples or features to diversify distributions complementary to source domains. Yet, these approaches can hardly deal with the re…

Cited by 26PDFcodeScholar
2023

Rethinking Alignment and Uniformity in Unsupervised Image Semantic Segmentation

AAAI 2023technical

Unsupervised image segmentation aims to match low-level visual features with semantic-level representations without outer supervision. In this paper, we address the critical properties from the view of feature alignments and feature uniformity for UISS models. We also make a comparison between UISS…

Cited by 22SourcePDFScholar
2023

Strip-MLP: Efficient Token Interaction for Vision MLP

ICCV 2023poster

Token interaction operation is one of the core modules in MLP-based models to exchange and aggregate information between different spatial locations. However, the power of token interaction on the spatial dimension is highly dependent on the spatial resolution of the feature maps, which limits the m…

Cited by 13PDFcodeScholar
2022

Density-driven Regularization for Out-of-distribution Detection

NeurIPS 2022accept

Detecting out-of-distribution (OOD) samples is essential for reliably deploying deep learning classifiers in open-world applications. However, existing detectors relying on discriminative probability suffer from the overconfident posterior estimate for OOD data. Other reported approaches either impo…

Cited by 17SourcePDFScholar
2022

Sparse Local Patch Transformer for Robust Face Alignment and Landmarks Inherent Relation Learning

CVPR 2022poster

Heatmap regression methods have dominated face alignment area in recent years while they ignore the inherent relation between different landmarks. In this paper, we propose a Sparse Local Patch Transformer (SLPT) for learning the inherent relation. The SLPT generates the representation of each singl…

Cited by 62PDFcodeScholar