← Search

Wei Su

11 accepted papers

2025

Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-time Open-Vocabulary Object Detection

ICCV 2025poster

The Mixture of Experts (MoE) architecture has excelled in Large Vision-Language Models (LVLMs), yet its potential in real-time open-vocabulary object detectors, which also leverage large-scale vision-language datasets but smaller models, remains unexplored. This work investigates this domain, reveal…

2025

ReFF: Reinforcing Format Faithfulness in Language Models Across Varied Tasks

AAAI 2025technical

Following formatting instructions to generate well-structured content is a fundamental yet often unmet capability for large language models (LLMs). To study this capability, which we refer to as format faithfulness, we present FormatBench, a comprehensive format-related benchmark. Compared to previo…

2025

TransBench: Breaking Barriers for Transferable Graphical User Interface Agents in Dynamic Digital Environments

ACL 2025finding

Graphical User Interface (GUI) agents, which autonomously operate on digital interfaces through natural language instructions, hold transformative potential for accessibility, automation, and user experience. A critical aspect of their functionality is grounding — the ability to map linguistic inten…

2024

ScanFormer: Referring Expression Comprehension by Iteratively Scanning

CVPR 2024poster

Referring Expression Comprehension (REC) aims to localize the target objects specified by free-form natural language descriptions in images. While state-of-the-art methods achieve impressive performance they perform a dense perception of images which incorporates redundant visual regions unrelated t…

Cited by 10SourcePDFScholar
2023

GaitGCI: Generative Counterfactual Intervention for Gait Recognition

CVPR 2023poster

Gait is one of the most promising biometrics that aims to identify pedestrians from their walking patterns. However, prevailing methods are susceptible to confounders, resulting in the networks hardly focusing on the regions that reflect effective walking patterns. To address this fundamental proble…

Cited by 59SourcePDFScholar
2023

Language Adaptive Weight Generation for Multi-Task Visual Grounding

CVPR 2023poster

Although the impressive performance in visual grounding, the prevailing approaches usually exploit the visual backbone in a passive way, i.e., the visual backbone extracts features with fixed weights without expression-related hints. The passive perception may lead to mismatches (e.g., redundant and…

2023

Referring Expression Comprehension Using Language Adaptive Inference

AAAI 2023technical

Different from universal object detection, referring expression comprehension (REC) aims to locate specific objects referred to by natural language expressions. The expression provides high-level concepts of relevant visual and contextual patterns, which vary significantly with different expressions…

Cited by 20SourcePDFScholar
2022

MetaGait: Learning to Learn an Omni Sample Adaptive Representation for Gait Recognition

ECCV 2022poster

"Gait recognition, which aims at identifying individuals by their walking patterns, has recently drawn increasing research attention. However, gait recognition still suffers from the conflicts between the limited binary visual clues of the silhouette and numerous covariates with diverse scales, whic…

Cited by 46SourcePDFScholar
2020

Efficient Dynamic Scene Deblurring Using Spatially Variant Deconvolution Network With Optical Flow Guided Training

CVPR 2020poster

In order to remove the non-uniform blur of images captured from dynamic scenes, many deep learning based methods design deep networks for large receptive fields and strong fitting capabilities, or use multi-scale strategy to deblur image on different scales gradually. Restricted by the fixed structu…

Cited by 121PDFScholar
2020

The Open Brands Dataset: Unified Brand Detection and Recognition at Scale

ICASSP 2020accepted

Intellectual property protection(IPP) have received more and more attention recently due to the development of the global e-commerce platforms. brand recognition plays a significant role in IPP. Recent studies for brand recognition and detection are based on small-scale datasets that are not compreh…

Cited by 0SourceScholar