← Search

Bin Hu

49 accepted papers

2026

A Supervised Multi-task Framework for Joint cryo-ET Restoration Enabled by Generative Physical Simulation

CVPR 2026

Cryo-electron tomography (cryo-ET) enables in-situ visualization of cellular ultrastructure, but reconstructions are severely degraded by extremely low SNR and missing-wedge artifacts due to dose limits and restricted tilt angles. Existing learning-based approaches are further constrained by inaccur

Cited by 0SourceScholar
2026

Clarify Before You Draw: Proactive Agents for Robust Text-to-CAD Generation

ICML 2026poster

Large language models have recently enabled text-to-CAD systems that synthesize parametric CAD programs (e.g., CadQuery) from natural-language prompts. In practice, however, geometric descriptions can be under-specified or internally inconsistent: critical dimensions may be missing and constraints m…

Cited by 0SourceScholar
2026

Cyto-SSL: A Self-Supervised Pretraining Framework for Cytology Foundation Model

AAAI 2026technical

Cytological images originate from exfoliated cells, collected via liquid-based slides and digitized into whole slide images (WSIs). Unlike histological WSIs that exhibit continuous and well-structured tissue, cytological WSIs are sparse in spatial distribution and unstructured in cellular relationsh

Cited by 0SourcePDFScholar
2026

Doxing via the Lens: Revealing Location-related Privacy Leakage on Multi-modal Large Reasoning Models

ICLR 2026poster

Recent advances in multi-modal large reasoning models (MLRMs) have shown significant ability to interpret complex visual content. While these models possess impressive reasoning capabilities, they also introduce novel and underexplored privacy risks. In this paper, we identify a novel category of pr…

Cited by 0SourcecodeScholar
2026

ST-LLM: Spatial Transcriptomics Embedding with Large Language Models

AAAI 2026technical

Spatial transcriptomics provides unprecedented opportunities to analyze gene patterns while preserving spatial tissue architecture. However, traditional deep learning methods for spatial transcriptomics analysis face significant challenges in multi-modal data integration, spatial dependency modeling

Cited by 0SourcePDFScholar
2026

TSRBench: A Comprehensive Multi-task Multi-modal Time Series Reasoning Benchmark for Generalist Models

ICML 2026poster

Time series data is ubiquitous in real-world scenarios and crucial for critical applications ranging from energy management to traffic control. Consequently, the ability to reason over time series is a fundamental skill for generalist models to solve complex problems. However, current benchmarks for…

Cited by 0SourceScholar
2026

Venom: Liquid Diffusion-Guided Gradient Inversion for Breaking Differential Privacy in Federated Learning

AAAI 2026technical

Gradient perturbation mechanisms, such as differential privacy (DP), aim to defend against gradient inversion attacks (GIA) by injecting noise into the shared gradients. Recent studies have shown that DP-based defenses lack robustness against advanced GIAs. However, existing gradient inversion metho

Cited by 0SourcePDFScholar
2025

A Framework Based on Data Augmentation for Knowledge Graph Entity Typing

ICASSP 2025accepted

The task of knowledge graph entity typing (KGET) aims to infer the missing types for entities in knowledge graphs, which is a significant subtask of knowledge graph completion (KGC). In despite of its progress, we observe that the sparsity of the dataset greatly affects the task itself as well as do…

Cited by 0SourceScholar
2025

Adjoint Sampling: Highly Scalable Diffusion Samplers via Adjoint Matching

ICML 2025poster

We introduce Adjoint Sampling, a highly scalable and efficient algorithm for learning diffusion processes that sample from unnormalized densities, or energy functions. It is the first on-policy approach that allows significantly more gradient updates than the number of energy evaluations and model s…

2025

Distributed Perception Aware Safe Leader Follower System via Control Barrier Methods

ICRA 2025

This paper addresses a distributed leader-follower formation control problem for a group of agents, each using a body-fixed camera with a limited field of view (FOV) for state estimation. The main challenge arises from the need to coordinate the agents' movements with their cameras' FOV to maintain

Cited by 1SourceScholar
2025

DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models

ICLR 2025poster

The rapid advancements in Vision-Language Models (VLMs) have shown great potential in tackling mathematical reasoning tasks that involve visual context. Unlike humans who can reliably apply solution steps to similar problems with minor modifications, we found that state-of-the-art VLMs like GPT-4o c…

Cited by 16SourcePDFScholar
2025

External Memory Matters: Generalizable Object-Action Memory for Retrieval-Augmented Long-Term Video Understanding

IJCAI 2025

Long video understanding with Large Language Models (LLMs) enables the description of objects that are not explicitly present in the training data. However, continuous changes in known objects and the emergence of new ones require up-to-date knowledge of objects and their dynamics for effective unde

Cited by 0SourcePDFScholar
2025

How LLMs Comprehend Temporal Meaning in Narratives: A Case Study in Cognitive Evaluation of LLMs

ACL 2025long

Large language models (LLMs) exihibit increasingly sophisticated linguistic capabilities, yet the extent to which these behaviors reflect human-like cognition versus advanced pattern recognition remains an open question.In this study, we investigate how LLMs process the temporal meaning of linguisti…

Cited by 0SourcePDFScholar
2025

PhysDiff: Physiology-based Dynamicity Disentangled Diffusion Model for Remote Physiological Measurement

AAAI 2025technical

Recent works on remote PhotoPlethysmoGraphy (rPPG) estimation typically use techniques like CNNs and Transformers to encode implicit features from facial videos for prediction. These methods learn to directly map facial videos to the static values of rPPG signals, overlooking the inherent dynamic ch…

2025

Toward Engineering AGI: Benchmarking the Engineering Design Capabilities of LLMs

NeurIPS 2025poster

Modern engineering, spanning electrical, mechanical, aerospace, civil, and computer disciplines, stands as a cornerstone of human civilization and the foundation of our society. However, engineering design poses a fundamentally different challenge for large language models (LLMs) compared with tradi…

Cited by 0SourceScholar
2025

Two‑Stage Learning of Stabilizing Neural Controllers via Zubov Sampling and Iterative Domain Expansion

NeurIPS 2025spotlight

Learning-based neural network (NN) control policies have shown impressive empirical performance. However, obtaining stability guarantees and estimates of the region of attraction of these learned neural controllers is challenging due to the lack of stable and scalable training and verification algor…

Cited by 0SourcecodeScholar
2025

VLASCD: A Visual Language Action Model for Simultaneous Chatting and Decision Making

EMNLP 2025

Recent large pretrained models such as LLMs (e.g., GPT series) and VLAs (e.g., OpenVLA) have achieved notable progress on multimodal tasks, yet they are built upon a multi-input single-output (MISO) paradigm. We show that this paradigm fundamentally limits performance in multi-input multi-output (MI

2024

A Supervised Information Enhanced Multi-Granularity Contrastive Learning Framework for EEG Based Emotion Recognition

ICASSP 2024accepted

This study introduces a novel Supervised Info-enhanced Contrastive Learning framework for EEG based Emotion Recognition (SI-CLEER). SI-CLEER employs multi-granularity contrastive learning to create robust EEG contextual representations, potentially improving emotion recognition effectiveness. Unlike…

Cited by 0SourceScholar
2024

COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

ICML 2024poster

Jailbreaks on large language models (LLMs) have recently received increasing attention. For a comprehensive assessment of LLM safety, it is essential to consider jailbreaks with diverse attributes, such as contextual coherence and sentiment/stylistic variations, and hence it is beneficial to study c…

2024

Deep Fusion of Shifted MLP and CNN for Medical Image Segmentation

ICASSP 2024accepted

Medical image segmentation is an important task in modern analysis of medical images. Current methods tend to extract either local features with convolutions or global features with Transformers. However, few of them are able to effectively fuse global and local features to facilitate segmentation.…

Cited by 0SourceScholar
2024

Fine-grained Local Sensitivity Analysis of Standard Dot-Product Self-Attention

ICML 2024poster

Self-attention has been widely used in various machine learning models, such as vision transformers. The standard dot-product self-attention is arguably the most popular structure, and there is a growing interest in understanding the mathematical properties of such attention mechanisms. This paper p…

2024

Large Language Model as a Policy Teacher for Training Reinforcement Learning Agents

IJCAI 2024poster

Recent studies have uncovered the potential of Large Language Models (LLMs) in addressing complex sequential decision-making tasks through the provision of high-level instructions. However, LLM-based agents lack specialization in tackling specific target problems, particularly in real-time dynamic e…

2024

Novel Quadratic Constraints for Extending LipSDP beyond Slope-Restricted Activations

ICLR 2024poster

Recently, semidefinite programming (SDP) techniques have shown great promise in providing accurate Lipschitz bounds for neural networks. Specifically, the LipSDP approach (Fazlyab et al., 2019) has received much attention and provides the least conservative Lipschitz upper bounds that can be compute…

Cited by 6SourcePDFScholar
2024

On the Scalability and Memory Efficiency of Semidefinite Programs for Lipschitz Constant Estimation of Neural Networks

ICLR 2024poster

Lipschitz constant estimation plays an important role in understanding generalization, robustness, and fairness in deep learning. Unlike naive bounds based on the network weight norm product, semidefinite programs (SDPs) have shown great promise in providing less conservative Lipschitz bounds with p…

2024

Structural Information Enhanced Graph Representation for Link Prediction

AAAI 2024technical

Link prediction is a fundamental task of graph machine learning, and Graph Neural Network (GNN) based methods have become the mainstream approach due to their good performance. However, the typical practice learns node representations through neighborhood aggregation, lacking awareness of the struct…

Cited by 5SourcePDFScholar
2023

A Unified Algebraic Perspective on Lipschitz Neural Networks

ICLR 2023top-25%

Important research efforts have focused on the design and training of neural networks with a controlled Lipschitz constant. The goal is to increase and sometimes guarantee the robustness against adversarial attacks. Recent promising techniques draw inspirations from different backgrounds to design 1…

2023

Complexity of Derivative-Free Policy Optimization for Structured $\mathcal{H}_\infty$ Control

NeurIPS 2023poster

The applications of direct policy search in reinforcement learning and continuous control have received increasing attention. In this work, we present novel theoretical results on the complexity of derivative-free policy optimization on an important class of robust control tasks, namely the structur…

Cited by 0SourcePDFScholar
2023

Daily Mental Health Monitoring from Speech: A Real-World Japanese Dataset and Multitask Learning Analysis

ICASSP 2023accepted

Translating mental health recognition from clinical research into real-world application requires extensive data, yet existing emotion datasets are impoverished in terms of daily mental health monitoring, especially when aiming for self-reported anxiety and depression recognition. We introduce the J…

Cited by 0SourceScholar
2023

Diffusion-Based Adversarial Sample Generation for Improved Stealthiness and Controllability

NeurIPS 2023poster

Neural networks are known to be susceptible to adversarial samples: small variations of natural examples crafted to deliberately mislead the models. While they can be easily generated using gradient-based techniques in digital and physical scenarios, they often differ greatly from the actual data di…

2023

Exploiting Connections between Lipschitz Structures for Certifiably Robust Deep Equilibrium Models

NeurIPS 2023poster

Recently, deep equilibrium models (DEQs) have drawn increasing attention from the machine learning community. However, DEQs are much less understood in terms of certified robustness than their explicit network counterparts. In this paper, we advance the understanding of certified robustness of DEQs…

2023

Federated Intelligent Terminals Facilitate Stuttering Monitoring

ICASSP 2023accepted

Stuttering is a complicated language disorder. The most common form of stuttering is developmental stuttering, which begins in childhood. Early monitoring and intervention are essential for the treatment of children with stuttering. Automatic speech recognition technology has shown its great potenti…

Cited by 0SourceScholar
2022

A Glance-and-Gaze Network for Respiratory Sound Classification

ICASSP 2022accepted

A plethora of great successes has been achieved by the existing convolutional neural networks (CNN) for respiratory sound classification. Nevertheless, simultaneously capturing both the local and global features can never be an easy task due to the limitation of a CNN’s structure. In this contributi…

Cited by 8SourceScholar
2022

Global Convergence of Direct Policy Search for State-Feedback $\mathcal{H}_\infty$ Robust Control: A Revisit of Nonsmooth Synthesis with Goldstein Subdifferential

NeurIPS 2022accept

Direct policy search has been widely applied in modern reinforcement learning and continuous control. However, the theoretical properties of direct policy search on nonsmooth robust control synthesis have not been fully understood. The optimal $\mathcal{H}_\infty$ control framework aims at designing…

Cited by 14SourcePDFScholar
2022

Provable Acceleration of Heavy Ball beyond Quadratics for a Class of Polyak-Lojasiewicz Functions when the Non-Convexity is Averaged-Out

ICML 2022spotlight

Heavy Ball (HB) nowadays is one of the most popular momentum methods in non-convex optimization. It has been widely observed that incorporating the Heavy Ball dynamic in gradient-based methods accelerates the training process of modern machine learning models. However, the progress on establishing i…

Cited by 24SourcePDFScholar
2021

Derivative-Free Policy Optimization for Linear Risk-Sensitive and Robust Control Design: Implicit Regularization and Sample Complexity

NeurIPS 2021poster

Direct policy search serves as one of the workhorses in modern reinforcement learning (RL), and its applications in continuous control tasks have recently attracted increasing attention. In this work, we investigate the convergence theory of policy gradient (PG) methods for learning the linear risk-…

Cited by 62SourcePDFScholar
2021

Multi-Level Graph Encoding with Structural-Collaborative Relation Learning for Skeleton-Based Person Re-Identification

IJCAI 2021poster

Skeleton-based person re-identification (Re-ID) is an emerging open topic providing great value for safety-critical applications. Existing methods typically extract hand-crafted features or model skeleton dynamics from the trajectory of body joints, while they rarely explore valuable relation inform…

2021

The Realization of Intelligent Robot System for Milk Tea Production

RA-L 2021

At present, there are many challenges for robotic systems to grasp in a dynamic and unstructured environment. The Robotic Grasping and Manipulation Competition (RGMC) aims to encourage researchers to focus on these challenges. The solution proposed in this letter was used to compete in the 2019 and

Cited by 0SourceScholar
2021

Weakly-Supervised Instance Segmentation via Class-Agnostic Learning With Salient Images

CVPR 2021poster

Humans have a strong class-agnostic object segmentation ability and can outline boundaries of unknown objects precisely, which motivates us to propose a box-supervised class-agnostic object segmentation (BoxCaseg) based solution for weakly-supervised instance segmentation. The BoxCaseg model is join…

Cited by 46PDFcodeScholar
2020

On the Stability and Convergence of Robust Adversarial Reinforcement Learning: A Case Study on Linear Quadratic Systems

NeurIPS 2020poster

Reinforcement learning (RL) algorithms can fail to generalize due to the gap between the simulation and the real world. One standard remedy is to use robust adversarial RL (RARL) that accounts for this gap during the policy training, by modeling the gap as an adversary against the training agent. In…

Cited by 61SourcePDFScholar
2020

Self-Supervised Gait Encoding with Locality-Aware Attention for Person Re-Identification

IJCAI 2020poster

Gait-based person re-identification (Re-ID) is valuable for safety-critical applications, and using only 3D skeleton data to extract discriminative gait features for person Re-ID is an emerging open topic. Existing methods either adopt hand-crafted features or learn gait features by traditional supe…

2019

Characterizing the Exact Behaviors of Temporal Difference Learning Algorithms Using Markov Jump Linear System Theory

NeurIPS 2019poster

In this paper, we provide a unified analysis of temporal difference learning algorithms with linear function approximators by exploiting their connections to Markov jump linear systems (MJLS). We tailor the MJLS theory developed in the control community to characterize the exact behaviors of the fir…

Cited by 73SourcePDFScholar
2018

Dissipativity Theory for Accelerating Stochastic Variance Reduction: A Unified Analysis of SVRG and Katyusha Using Semidefinite Programs

ICML 2018oral

Techniques for reducing the variance of gradient estimates used in stochastic programming algorithms for convex finite-sum problems have received a great deal of attention in recent years. By leveraging dissipativity theory from control, we provide a new perspective on two important variance-reducti…

Cited by 26SourcePDFScholar
2018

Learning Intrinsic Sparse Structures within Long Short-Term Memory

ICLR 2018poster

Model compression is significant for the wide adoption of Recurrent Neural Networks (RNNs) in both user devices possessing limited resources and business clusters requiring quick responses to large-scale service requests. This work aims to learn structurally-sparse Long Short-Term Memory (LSTM) by r…

Cited by 161SourcePDFScholar