← Search

Cong Hu

12 accepted papers

2026

FB-CLIP: Fine-Grained Zero-Shot Anomaly Detection with Foreground-Background Disentanglement

CVPR 2026

Fine-grained anomaly detection is crucial in industrial and medical applications, but labeled anomalies are often scarce, making zero-shot detection challenging. While vision-language models like CLIP offer promising solutions, they struggle with foreground-background feature entanglement and coarse

Cited by 1SourcecodeScholar
2026

Multi-Paradigm Collaborative Adversarial Attack Against Multi-Modal Large Language Models

CVPR 2026

The rapid progress of Multi-Modal Large Language Models (MLLMs) has significantly advanced downstream applications. However, this progress also exposes serious transferable adversarial vulnerabilities. In general, existing adversarial attacks against MLLMs typically rely on surrogate models trained

Cited by 0SourcecodeScholar
2026

Towards Highly Transferable Vision-Language Attack via Semantic-Augmented Dynamic Contrastive Interaction

CVPR 2026

With the rapid advancement and widespread application of vision-language pre-training (VLP) models, their vulnerability to adversarial attacks has become a critical concern. In general, the adversarial examples can typically be designed to exhibit transferable power, attacking not only different mod

Cited by 0SourcecodeScholar
2025

Improving Multilingual Sign Language Translation with Automatically Clustered Language Family Information

COLING 2025main

Sign Language Translation (SLT) bridges the communication gap between deaf and hearing individuals by converting sign language videos into spoken language texts. While most SLT research has focused on bilingual translation models, the recent surge in interest has led to the exploration of Multilingu…

2025

beta-FFT: Nonlinear Interpolation and Differentiated Training Strategies for Semi-Supervised Medical Image Segmentation

CVPR 2025poster

Co-training has achieved significant success in the field of semi-supervised learning; however, the *homogenization phenomenon*, which arises from multiple models tending towards similar decision boundaries, remains inadequately addressed. To tackle this issue, we propose a novel algorithm called **…

2024

Adaptive Simultaneous Sign Language Translation with Confident Translation Length Estimation

COLING 2024main

Traditional non-simultaneous Sign Language Translation (SLT) methods, while effective for pre-recorded videos, face challenges in real-time scenarios due to inherent inference delays. The emerging field of simultaneous SLT aims to address this issue by progressively translating incrementally receive…

2024

An Explicit Multi-Modal Fusion Method for Sign Language Translation

ICASSP 2024accepted

Sign Language Translation (SLT) aims to convert sign language videos into corresponding spoken text sequences. However, the inherent modality gap between sign language video and text hinders the development of SLT. Motivated by the linguistic consistency between gloss <sup xmlns:mml="http://www.w3.o…

Cited by 0SourceScholar
2024

Conditional Variational Autoencoder for Sign Language Translation with Cross-Modal Alignment

AAAI 2024technical

Sign language translation (SLT) aims to convert continuous sign language videos into textual sentences. As a typical multi-modal task, there exists an inherent modality gap between sign language videos and spoken language text, which makes the cross-modal alignment between visual and textual modalit…

2023

A Token-Level Contrastive Framework for Sign Language Translation

ICASSP 2023accepted

Sign Language Translation (SLT) is a promising technology to bridge the communication gap between the deaf and the hearing people. Recently, researchers have adopted Neural Machine Translation (NMT) methods, which usually require large-scale corpus for training, to achieve SLT. However, the publicly…

Cited by 0SourceScholar
2023

An Improved Koopman-MPC Framework for Data-Driven Modeling and Control of Soft Actuators

RA-L 2023

The challenge of achieving precise control of soft actuators with strong nonlinearity is mainly due to the difficulty of deriving models suitable for model-based control techniques. Fortunately, Koopman operator provides a data-driven method for constructing control-oriented models of nonlinear syst

Cited by 49SourceScholar
2022

Where to Attack: A Dynamic Locator Model for Backdoor Attack in Text Classifications

COLING 2022main

Nowadays, deep-learning based NLP models are usually trained with large-scale third-party data which can be easily injected with malicious backdoors. Thus, BackDoor Attack (BDA) study has become a trending research to help promote the robustness of an NLP system. Text-based BDA aims to train a poiso…

2018

Continuous Shared Control for Robotic Arm Reaching Driven by a Hybrid Gaze-Brain Machine Interface

IROS 2018poster

The brain-machine interface (BMI) has been reported to offer the potential for controlling the assistive robot for the motor impaired people, using the non-invasively obtained electroencephalogram (EEG) signals. However, the EEG based BMI may not be sufficient and stable to drive the robot moving fr…

Cited by 12SourceScholar