← Search

Wen Huang

20 accepted papers

2026

3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models

ICML 2026oral

Current Large Language Models have achieved Olympiad-level logic, yet Vision-Language Models paradoxically falter on elementary spatial tasks like block counting. This capability mismatch reveals a critical "spatial intelligence gap," where models fail to construct coherent 3D mental representations…

Cited by 0SourceScholar
2026

ATA: Bridging Implicit Reasoning with Attention-Guided and Action-Guided Inference for Vision-Language Action Models

ICRA 2026poster

Vision-Language-Action (VLA) models rely on current observations, including images, language instructions, and robot states, to predict actions and complete tasks. While accurate visual perception is crucial for precise action prediction and execution, recent work has attempted to further improve pe…

2026

DeCoDe: Decoupling Binding Position and Molecular Conformation in 3D Ligand Diffusion for Structure-Based Drug Design

ICML 2026spotlight

Recent advances in diffusion models show promise for Structure-Based Drug Design (SBDD), which aims to generate 3D ligand molecules that bind tightly to specific protein targets. This involves jointly optimizing the ligand's 3D conformation and its binding position within the protein pocket. However…

Cited by 0SourceScholar
2026

Improving Deepfake Detection with Reinforcement Learning-Based Adaptive Data Augmentation

AAAI 2026technical

The generalization capability of deepfake detectors is crucial for real-world applications. Data augmentation to generate synthetic fake faces has served as an effective strategy to enhance generalization. Interestingly, current state-of-the-art (SoTA) methods rely on fixed augmentation strategies,

Cited by 0SourcePDFScholar
2026

Rejoining Precious Artifacts: Efficiently Bone Stick Rejoining Based Massive Fragment Images by Contour, Script, and Texture

AAAI 2026technical

Rejoining fragment images of precious artifacts is a meaningful task because complete artifacts could provide valuable clues for the research of human civilization. However, existing rejoining methods face several challenges including time-consuming manual annotation, insufficient rejoining accuracy

Cited by 0SourcePDFScholar
2026

RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation Localization

ICLR 2026poster

Visual manipulation localization (VML) aims to identify tampered regions in images and videos, a task that has become increasingly challenging with the rise of advanced editing tools. Existing methods face two main issues: resolution diversity, where resizing or padding distorts forensic traces and…

Cited by 0SourcecodeScholar
2025

Data-Efficient Low-Complexity Acoustic Scene Classification via Distilling and Progressive Pruning

ICASSP 2025accepted

The goal of the acoustic scene classification (ASC) task is to classify recordings into one of the predefined acoustic scene classes. However, in real-world scenarios, ASC systems often encounter challenges such as recording device mismatch, low-complexity constraints, and the limited availability o…

Cited by 0SourceScholar
2025

Generalizable Audio Deepfake Detection via Latent Space Refinement and Augmentation

ICASSP 2025accepted

Advances in speech synthesis technologies, like text-to-speech (TTS) and voice conversion (VC), have made detecting deepfake speech increasingly challenging. Spoofing countermeasures often struggle to generalize effectively, particularly when faced with unseen attacks. To address this, we propose a…

Cited by 0SourceScholar
2025

Improving Federated Domain Generalization Through Dynamical Weights Calculated from Data Influences on Global Model Update

AAAI 2025technical

With the popularity of federated learning, federated domain generalization (FedDG) has attracted more and more attentions. Existing works of federated learning indicate that the generalization performance of the global model can be improved when the global model is obtained by aggregating local mode…

Cited by 0SourcePDFScholar
2025

SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods

ACL 2025long

As speech generation technology advances, the risk of misuse through deepfake audio has become a pressing concern, which underscores the critical need for robust detection systems. However, many existing speech deepfake datasets are limited in scale and diversity, making it challenging to train mode…

2024

Data-Free Hard-Label Robustness Stealing Attack

AAAI 2024technical

The popularity of Machine Learning as a Service (MLaaS) has led to increased concerns about Model Stealing Attacks (MSA), which aim to craft a clone model by querying MLaaS. Currently, most research on MSA assumes that MLaaS can provide soft labels and that the attacker has a proxy dataset with a si…

2024

Exploring Large Scale Pre-Trained Models for Robust Machine Anomalous Sound Detection

ICASSP 2024accepted

Machine anomalous sound detection is a useful technique for various applications, but it often suffers from poor generalization due to the challenges of data collection and complex acoustic environment. To address this issue, we propose a robust machine anomalous sound detection model that leverages…

Cited by 0SourceScholar
2024

Robust Cross-Domain Speaker Verification with Multi-Level Domain Adapters

ICASSP 2024accepted

Speaker verification encounters significant challenges when confronted with diverse domain data, often resulting in performance degradation due to domain mismatch. To enhance performance in cross-domain scenarios, we introduce the Domain Adapter, an adaptable module designed for specific domains. Th…

Cited by 0SourceScholar
2024

Robustly Improving Bandit Algorithms with Confounded and Selection Biased Offline Data: A Causal Approach

AAAI 2024technical

This paper studies bandit problems where an agent has access to offline data that might be utilized to potentially improve the estimation of each arm’s reward distribution. A major obstacle in this setting is the existence of compound biases from the observational data. Ignoring these biases and bli…

Cited by 2SourcePDFScholar
2024

Visual Hallucinations of Multi-modal Large Language Models

ACL 2024findings

Visual hallucination (VH) means that a multi-modal LLM (MLLM) imagines incorrect details about an image in visual question answering. Existing studies find VH instances only in existing image datasets, which results in biased understanding of MLLMs’ performance under VH due to limited diversity of s…

2022

Solving Stackelberg Prediction Game with Least Squares Loss via Spherically Constrained Least Squares Reformulation

ICML 2022oral

The Stackelberg prediction game (SPG) is popular in characterizing strategic interactions between a learner and an attacker. As an important special case, the SPG with least squares loss (SPG-LS) has recently received much research attention. Although initially formulated as a difficult bi-level opt…

Cited by 15SourcePDFScholar