← Search

Yu Huang

43 accepted papers

2026

Acoustic Interference: A New Paradigm Weaponizing Acoustic Latent Semantic for Universal Jailbreak against Large Audio Language Models

ICML 2026poster

The integration of audio modality into Large Audio Language Models (LALMs) significantly expands their attack surface. Existing jailbreak paradigms predominantly treat audio as a carrier for malicious payloads, relying on semantic optimization, acoustic parameter control, or additive perturbation to…

Cited by 1SourceScholar
2026

Beyond Conservation: Flexible Molecular Assembly with Unbalanced Diffusion Bridge

AAAI 2026technical

Molecular assembly (MA) has long been a fundamental task in chemistry and biology, with the potential to create new materials and enable novel functions beyond the molecular scale. However, its vast conformational search space poses substantial challenges, and current generative models remain limite

Cited by 0SourcePDFScholar
2026

Efficient Segmentation with Multimodal Large Language Model via Token Routing

AAAI 2026technical

Recent advances in multimodal large language models (MLLMs) have demonstrated strong capabilities in addressing open-world segmentation tasks. However, the substantial computational cost of the LLM components presents a significant challenge, especially in segmentation tasks, where efficiency has lo

Cited by 0SourcePDFScholar
2026

MedMKEB: A Comprehensive Knowledge Editing Benchmark for Medical Multimodal Large Language Models

AAAI 2026technical

Recent advances in multimodal large language models (MLLMs) have significantly improved medical AI, enabling it to unify the understanding of visual and textual information. However, as medical knowledge continues to evolve, it is critical to allow these models to efficiently update outdated or inco

Cited by 0SourcePDFScholar
2026

Multitasks-based Deep Evidential Fusion Network for Blind Image Quality Assessment

AAAI 2026technical

Blind image quality assessment (BIQA) methods often incorporate auxiliary tasks to improve performance. However, existing approaches face limitations due to insufficient integration and a lack of flexible uncertainty estimation, leading to suboptimal performance. To address these challenges, we prop

Cited by 0SourcePDFScholar
2026

On the Learning Dynamics of RLVR at the Edge of Competence

ICML 2026poster

Reinforcement Learning with Verifiable Rewards (RLVR) has been a main driver of recent breakthroughs in large reasoning models. Yet it remains a mystery how rewards based solely on final outcomes can help overcome the long-horizon barrier to extended reasoning. To understand this, we develop a theor…

Cited by 0SourceScholar
2026

Scene Experts: Specializing in 3D Gaussian Splatting with Adaptive Decomposition

AAAI 2026technical

Anchor-based 3D Gaussian Splatting (GS), exemplified by Scaffold-GS, achieves remarkable storage efficiency through a hybrid explicit-implicit representation. However, their reliance on a single, monolithic network to decode anchor features imposes a severe bottleneck on model capacity, often result

Cited by 0SourcePDFScholar
2026

Steering Where to Diffuse: Generative Modeling of Phenotypic Response Simulation with Steered Diffusion Bridge

CVPR 2026

Simulation of cellular morphology change has long been a fundamental task in quantitative biology and high-throughput screening, with the potential to accelerate therapeutic development and elucidate disease mechanisms beyond empirical clinical practice. However, the vast perturbation space poses ch

Cited by 0SourceScholar
2026

SysMoBench: Evaluating AI on Formally Specifying Complex Real-World Systems

ICLR 2026poster

Formal models are essential to specifying large, complex computer systems and verifying their correctness, but are notoriously expensive to write and maintain. Recent advances in generative AI show promise in generating certain forms of specifications. However, existing work mostly targets small cod…

Cited by 0SourcecodeScholar
2026

TIME: Tensor-Factorized Mixture-of-Experts with Intrinsic Routing for Lifelong Multimodal Knowledge Editing

ICML 2026poster

Lifelong multimodal knowledge editing allows vision language models to continuously adapt to dynamic updates to avoid catastrophic forgetting. To mitigate interference between sequential updates, recent paradigms have shifted towards modular parameter isolation. However, this strategy faces a critic…

Cited by 0SourceScholar
2026

Temporally Detailed Hypergraph Neural ODE for Disease Progression Modeling

ICLR 2026poster

Disease progression modeling aims to characterize and predict how a patient's disease complications worsen over time based on longitudinal electronic health records (EHRs). Accurate modeling of disease progression, such as type 2 diabetes, can enhance patient sub-phenotyping and inform effective and…

Cited by 0SourceScholar
2026

VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation

AAAI 2026technical

Visual generative models have achieved remarkable progress in synthesizing photorealistic images and videos, yet aligning their outputs with human preferences across critical dimensions remains a persistent challenge. Though reinforcement learning from human feedback offers promise for preference al

Cited by 0SourcePDFScholar
2026

XLinear: A Lightweight and Accurate MLP-Based Model for Long-Term Time Series Forecasting with Exogenous Inputs

AAAI 2026technical

Despite the prevalent assumption of uniform variable importance in long-term time series forecasting models, real-world applications often exhibit asymmetric causal relationships and varying data acquisition costs. Specifically, cost‐effective exogenous data (e.g., local weather) can unilaterally in

Cited by 0SourcePDFScholar
2025

A Theoretical Analysis of Self-Supervised Learning for Vision Transformers

ICLR 2025poster

Self-supervised learning has become a cornerstone in computer vision, primarily divided into reconstruction-based methods like masked autoencoders (MAE) and discriminative methods such as contrastive learning (CL). Recent empirical observations reveal that MAE and CL capture different types of repr…

Cited by 0SourcePDFScholar
2025

DatawiseAgent: A Notebook-Centric LLM Agent Framework for Adaptive and Robust Data Science Automation

EMNLP 2025

Existing large language model (LLM) agents for automating data science show promise, but they remain constrained by narrow task scopes, limited generalization across tasks and models, and over-reliance on state-of-the-art (SOTA) LLMs. We introduce DatawiseAgent, a notebook-centric LLM agent framewor

Cited by 0SourcePDFScholar
2025

Does Chain-of-Thought Reasoning Really Reduce Harmfulness from Jailbreaking?

ACL 2025finding

Jailbreak attacks have been observed to largely fail against recent reasoning models enhanced by Chain-of-Thought (CoT) reasoning. However, the underlying mechanism remains underexplored, and relying solely on reasoning capacity may raise security concerns. In this paper, we try to answer the questi…

Cited by 0SourcePDFScholar
2025

Domain Generalization in CLIP via Learning with Diverse Text Prompts

CVPR 2025poster

Domain generalization (DG) aims to train a model on source domains that can generalize well to unseen domains. Recent advances in Vision-Language Models (VLMs), such as CLIP, exhibit remarkable generalization capabilities across a wide range of data distributions, benefiting tasks like DG. However,…

Cited by 0SourcePDFScholar
2025

Exploit Your Latents: Coarse-Grained Protein Backmapping with Latent Diffusion Models

AAAI 2025technical

Coarse-grained (CG) molecular dynamics of proteins is a preferred approach to studying large molecules on extended time scales by condensing the entire atomic model into a limited number of pseudo-atoms and preserving the thermodynamic properties of the system. However, the significantly increased e…

Cited by 0SourcePDFScholar
2025

MIMO: A Medical Vision Language Model with Visual Referring Multimodal Input and Pixel Grounding Multimodal Output

CVPR 2025poster

Currently, medical vision language models are widely used in medical vision question answering tasks. However, existing models are confronted with two issues: for input, the model only relies on text instructions and lacks direct understanding of visual clues in the image; for output, the model only…

2025

MoleBridge: Synthetic Space Projecting with Discrete Markov Bridges

NeurIPS 2025poster

Molecular synthetic space projecting is a critical technique in de novo molecular design, which aims to rectify molecules without synthesizability guarantee by converting them into synthetic postfix notations. However, the vast synthesizable chemical space and the discrete data modalities involved p…

Cited by 0SourceScholar
2025

Multi-head Transformers Provably Learn Symbolic Multi-step Reasoning via Gradient Descent

NeurIPS 2025poster

Transformers have demonstrated remarkable capabilities in multi-step reasoning tasks. However, understandings of the underlying mechanisms by which they acquire these abilities through training remain limited, particularly from a theoretical standpoint. This work investigates how transformers learn…

Cited by 0SourceScholar
2025

Prototypical Graph Alignment for Text-based Person Search

ICASSP 2025accepted

Text-based Person Search is one of the downstream tasks of cross-modal retrieval. The key challenge is aligning features of two extremely irrelavant modalities into the same latent space. Recent works within Prototype Learning introduce a few learnable parameters to map heterogeneous features into t…

Cited by 0SourceScholar
2025

STAMPsy: Towards SpatioTemporal-Aware Mixed-Type Dialogues for Psychological Counseling

AAAI 2025technical

Online psychological counseling dialogue systems are trending, offering a convenient and accessible alternative to traditional in-person therapy. However, existing psychological counseling dialogue systems mainly focus on basic empathetic dialogue or QA with minimal professional knowledge and withou…

2025

Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization

NeurIPS 2025poster

The ability to reason lies at the core of artificial intelligence (AI), and challenging problems usually call for deeper and longer reasoning to tackle. A crucial question about AI reasoning is whether models can extrapolate learned reasoning patterns to solve harder tasks that require longer chain…

Cited by 0SourceScholar
2025

Understanding Fine-tuning CLIP for Open-vocabulary Semantic Segmentation in Hyperbolic Space

CVPR 2025poster

CLIP, a foundational vision-language model, has emerged as a powerful tool for open-vocabulary semantic segmentation. While freezing the text encoder preserves its powerful embeddings, recent studies show that fine-tuning both the text and image encoders jointly significantly enhances segmentation p…

2024

A Learnable Discrete-Prior Fusion Autoencoder with Contrastive Learning for Tabular Data Synthesis

AAAI 2024technical

The actual collection of tabular data for sharing involves confidentiality and privacy constraints, leaving the potential risks of machine learning for interventional data analysis unsafely averted. Synthetic data has emerged recently as a privacy-protecting solution to address this challenge. Howev…

Cited by 7SourcePDFScholar
2024

A Novel Multi-Atlas Fusion Model Based On Contrastive Learning For Functional Connectivity Graph Diagnosis

ICASSP 2024accepted

Functional connectivity (FC) graph analysis is an important method for diagnosing brain disorders using functional magnetic resonance imaging (fMRI). Existing FC graph diagnosis approaches preprocess the brain by dividing it into specific regions using atlases. However, relying on a single atlas exc…

Cited by 0SourceScholar
2024

Accelerating Convergence of Score-Based Diffusion Models, Provably

ICML 2024poster

Score-based diffusion models, while achieving remarkable empirical performance, often suffer from low sampling speed, due to extensive function evaluations needed during the sampling phase. Despite a flurry of recent activities towards speeding up diffusion generative modeling in practice, theoretic…

Cited by 80SourcePDFScholar
2024

BetterV: Controlled Verilog Generation with Discriminative Guidance

ICML 2024poster

Due to the growing complexity of modern Integrated Circuits (ICs), there is a need for automated circuit design methods. Recent years have seen increasing research in hardware design language generation to facilitate the design process. In this work, we propose a Verilog generation framework, Better…

Cited by 40SourcePDFScholar
2024

Detection, Diagnosis, and Explanation: A Benchmark for Chinese Medial Hallucination Evaluation

COLING 2024main

Large Language Models (LLMs) have made significant progress recently. However, their practical use in healthcare is hindered by their tendency to generate hallucinations. One specific type, called snowballing hallucination, occurs when LLMs encounter misleading information, and poses a security thre…

2024

In-Context Learning with Representations: Contextual Generalization of Trained Transformers

NeurIPS 2024poster

In-context learning (ICL) refers to a remarkable capability of pretrained large language models, which can learn a new task given a few examples during inference. However, theoretical understanding of ICL is largely under-explored, particularly whether transformers can be trained to generalize to un…

Cited by 8SourcePDFScholar
2024

MLeVLM: Improve Multi-level Progressive Capabilities based on Multimodal Large Language Model for Medical Visual Question Answering

ACL 2024findings

Medical visual question answering (MVQA) requires in-depth understanding of medical images and questions to provide reliable answers. We summarize multi-level progressive capabilities that models need to focus on in MVQA: recognition, details, diagnosis, knowledge, and reasoning. Existing MVQA model…

2023

CSS: A Large-scale Cross-schema Chinese Text-to-SQL Medical Dataset

ACL 2023findings

The cross-domain text-to-SQL task aims to build a system that can parse user questions into SQL on complete unseen databases, and the single-domain text-to-SQL task evaluates the performance on identical databases. Both of these setups confront unavoidable difficulties in real-world applications. To…

2023

ClimSim: A large multi-scale dataset for hybrid physics-ML climate emulation

NeurIPS 2023oral

Modern climate projections lack adequate spatial and temporal resolution due to computational constraints. A consequence is inaccurate and imprecise predictions of critical processes such as storms. Hybrid methods that combine physics with machine learning (ML) have introduced a new generation of hi…

2022

Modality Competition: What Makes Joint Training of Multi-modal Network Fail in Deep Learning? (Provably)

ICML 2022spotlight

Despite the remarkable success of deep multi-modal learning in practice, it has not been well-explained in theory. Recently, it has been observed that the best uni-modal network outperforms the jointly trained multi-modal network across different combinations of modalities on various tasks, which is…

Cited by 123SourcePDFScholar
2022

Provable Generalization of Overparameterized Meta-learning Trained with SGD

NeurIPS 2022accept

Despite the empirical success of deep meta-learning, theoretical understanding of overparameterized meta-learning is still limited. This paper studies the generalization of a widely used meta-learning approach, Model-Agnostic Meta-Learning (MAML), which aims to find a good initialization for fast ad…

Cited by 11SourcePDFScholar
2021

GAEN: Graph Attention Evolving Networks

IJCAI 2021poster

Real-world networked systems often show dynamic properties with continuously evolving network nodes and topology over time. When learning from dynamic networks, it is beneficial to correlate all temporal networks to fully capture the similarity/relevance between nodes. Recent work for dynamic networ…

2021

What Makes Multi-Modal Learning Better than Single (Provably)

NeurIPS 2021poster

The world provides us with data of multiple modalities. Intuitively, models fusing data from different modalities outperform their uni-modal counterparts, since more information is aggregated. Recently, joining the success of deep learning, there is an influential line of work on deep multi-modal le…

Cited by 340SourcePDFScholar