← Search

Wenbing Huang

73 accepted papers

2026

CARD: Coarse-to-fine Autoregressive Modeling with Radix-based Decomposition for Transferable Free Energy Estimation

ICML 2026poster

Estimating free energy differences quantifies thermodynamic preferences in molecular interactions, which is central to chemistry and drug discovery. Despite fruitful progress, existing methods still face key limitations: classical computational approaches remain prohibitively expensive due to their …

Cited by 0SourceScholar
2026

CloDS: Visual-Only Unsupervised Cloth Dynamics Learning in Unknown Conditions

ICLR 2026poster

Deep learning has demonstrated remarkable capabilities in simulating complex dynamic systems. However, existing methods require known physical properties as supervision or inputs, limiting their applicability under unknown conditions. To explore this challenge, we introduce Cloth Dynamics Grounding…

Cited by 0SourcecodeScholar
2026

Recurrent Reasoning with Vision-Language Models for Estimating Long-Horizon Embodied Task Progress

CVPR 2026

Accurately estimating task progress is critical for embodied agents to plan and execute long-horizon, multi-step tasks. Despite promising advances, existing Vision-Language Models (VLMs) based methods primarily leverage their video understanding capabilities, while neglecting their complex reasoning

Cited by 0SourceScholar
2026

STAR-R1: Multi-View Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

CVPR 2026

Multimodal Large Language Models (MLLMs) remain far from human-level performance in multi-view spatial reasoning, where models must establish object correspondences across view and infer coherent scene semantics. We analyze this limitation through the Transformation-Driven Visual Reasoning (TVR) tas

Cited by 0SourcecodeScholar
2025

DenoiseVAE: Learning Molecule-Adaptive Noise Distributions for Denoising-based 3D Molecular Pre-training

ICLR 2025poster

Denoising learning of 3D molecules learns molecular representations by imposing noises into the equilibrium conformation and predicting the added noises to recover the equilibrium conformation, which essentially captures the information of molecular force fields. Due to the specificity of Potential…

Cited by 2SourcePDFScholar
2025

Geometric Mixture Models for Electrolyte Conductivity Prediction

NeurIPS 2025poster

Accurate prediction of ionic conductivity in electrolyte systems is crucial for advancing numerous scientific and technological applications. While significant progress has been made, current research faces two fundamental challenges: (1) the lack of high-quality standardized benchmarks, and (2) ina…

Cited by 0SourceScholar
2025

HeMeNet: Heterogeneous Multichannel Equivariant Network for Protein Multi-task Learning

AAAI 2025technical

Understanding and leveraging the 3D structures of proteins is central to a variety of biological and drug discovery tasks. While deep learning has been applied successfully for structure-based protein function prediction tasks, current methods usually employ distinct training for each task. However,…

2025

Large Language-Geometry Model: When LLM meets Equivariance

ICML 2025poster

Accurately predicting 3D structures and dynamics of physical systems is crucial in scientific applications. Existing approaches that rely on geometric Graph Neural Networks (GNNs) effectively enforce $\mathrm{E}(3)$-equivariance, but they often fail in leveraging extensive broader information. While…

Cited by 4SourcePDFScholar
2025

Latent Retrieval Augmented Generation of Cross-Domain Protein Binders

NeurIPS 2025poster

Designing protein binders targeting specific sites, which requires to generate realistic and functional interaction patterns, is a fundamental challenge in drug discovery. Current structure-based generative models are limited in generating nterfaces with sufficient rationality and interpretability.…

Cited by 0SourceScholar
2025

Learning 3D Anisotropic Noise Distributions Improves Molecular Force Fields

NeurIPS 2025poster

Coordinate denoising has emerged as a promising method for 3D molecular pretraining due to its theoretical connection to learning molecular force field. However, existing denoising methods rely on oversimplied molecular dynamics that assume atomic motions to be isotropic and homoscedastic. To addres…

Cited by 0SourcecodeScholar
2025

MOF-BFN: Metal-Organic Frameworks Structure Prediction via Bayesian Flow Networks

NeurIPS 2025poster

Metal-Organic Frameworks (MOFs) have attracted considerable attention due to their unique properties including high surface area and tunable porosity, and promising applications in catalysis, gas storage, and drug delivery. Structure prediction for MOFs is a challenging task, as these frameworks are…

Cited by 0SourceScholar
2025

ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning

EMNLP 2025

Reasoning-based large language models have excelled in mathematics and programming, yet their potential in knowledge-intensive medical question answering remains underexplored and insufficiently validated in clinical contexts. To bridge this gap, we introduce ReasonMed , the largest medical reasonin

2025

Size-Generalizable RNA Structure Evaluation by Exploring Hierarchical Geometries

ICLR 2025poster

Understanding the 3D structure of RNA is essential for deciphering its function and developing RNA-based therapeutics. Geometric Graph Neural Networks (GeoGNNs) that conform to the $\mathrm{E}(3)$-symmetry have advanced RNA structure evaluation, a crucial step toward RNA structure prediction. Howeve…

Cited by 2SourcePDFScholar
2025

StableToolBench-MirrorAPI: Modeling Tool Environments as Mirrors of 7,000+ Real-World APIs

ACL 2025finding

The rapid advancement of large language models (LLMs) has spurred significant interest in tool learning, where LLMs are augmented with external tools to tackle complex tasks. However, existing tool environments face challenges in balancing stability, scale, and realism, particularly for benchmarking…

2025

Think Then React: Towards Unconstrained Action-to-Reaction Motion Generation

ICLR 2025poster

Modeling human-like action-to-reaction generation has significant real-world applications, like human-robot interaction and games. Despite recent advancements in single-person motion generation, it is still challenging to well handle action-to-reaction generation, due to the difficulty of directly p…

Cited by 2SourcePDFScholar
2025

Towards Effective and Efficient Continual Pre-training of Large Language Models

ACL 2025long

Continual pre-training (CPT) has been an important approach for adapting language models to specific domains or tasks. In this paper, we comprehensively study its key designs to balance the new abilities while retaining the original abilities, and present an effective CPT method that can greatly imp…

2025

Two-in-One: Unified Multi-Person Interactive Motion Generation by Latent Diffusion Transformer

ICASSP 2025accepted

Multi-person interactive motion generation, a critical yet under-explored domain in computer character animation, poses significant challenges such as intricate modeling of inter-human interactions beyond individual motions and generating two motions with huge differences from one text condition. Cu…

Cited by 0SourceScholar
2025

UniMoMo: Unified Generative Modeling of 3D Molecules for De Novo Binder Design

ICML 2025poster

The design of target-specific molecules such as small molecules, peptides, and antibodies is vital for biological research and drug discovery. Existing generative methods are restricted to single-domain molecules, failing to address versatile therapeutic needs or utilize cross-domain transferability…

2025

UniSim: A Unified Simulator for Time-Coarsened Dynamics of Biomolecules

ICML 2025poster

Molecular Dynamics (MD) simulations are essential for understanding the atomic-level behavior of molecular systems, giving insights into their transitions and interactions. However, classical MD techniques are limited by the trade-off between accuracy and efficiency, while recent deep learning-based…

2025

Universally Invariant Learning in Equivariant GNNs

NeurIPS 2025poster

Equivariant Graph Neural Networks (GNNs) have demonstrated significant success across various applications. To achieve completeness---that is, the universal approximation property over the space of equivariant functions---the network must effectively capture the intricate multi-body interactions amo…

Cited by 0SourceScholar
2025

Zero-Shot Cyclic Peptide Design via Composable Geometric Constraints

ICML 2025poster

Cyclic peptides, characterized by geometric constraints absent in linear peptides, offer enhanced biochemical properties, presenting new opportunities to address unmet medical needs. However, designing target-specific cyclic peptides remains underexplored due to limited training data. To bridge the…

Cited by 0SourcePDFScholar
2024

3D Structure Prediction of Atomic Systems with Flow-based Direct Preference Optimization

NeurIPS 2024poster

Predicting high-fidelity 3D structures of atomic systems is a fundamental yet challenging problem in scientific domains. While recent work demonstrates the advantage of generative models in this realm, the exploration of different probability paths are still insufficient, and hallucinations during s…

Cited by 0SourcePDFScholar
2024

Are High-Degree Representations Really Unnecessary in Equivariant Graph Neural Networks?

NeurIPS 2024poster

Equivariant Graph Neural Networks (GNNs) that incorporate E(3) symmetry have achieved significant success in various scientific applications. As one of the most successful models, EGNN leverages a simple scalarization technique to perform equivariant message passing over only Cartesian vectors (i.e.…

2024

EquiPocket: an E(3)-Equivariant Geometric Graph Neural Network for Ligand Binding Site Prediction

ICML 2024oral

Predicting the binding sites of target proteins plays a fundamental role in drug discovery. Most existing deep-learning methods consider a protein as a 3D image by spatially clustering its atoms into voxels and then feed the voxelized protein into a 3D CNN for prediction. However, the CNN-based meth…

2024

Equivariant Diffusion for Crystal Structure Prediction

ICML 2024poster

In addressing the challenge of Crystal Structure Prediction (CSP), symmetry-aware deep learning models, particularly diffusion models, have been extensively studied, which treat CSP as a conditional generation task. However, ensuring permutation, rotation, and periodic translation equivariance durin…

Cited by 14SourcePDFScholar
2024

Full-Atom Peptide Design with Geometric Latent Diffusion

NeurIPS 2024poster

Peptide design plays a pivotal role in therapeutics, allowing brand new possibility to leverage target binding sites that are previously undruggable. Most existing methods are either inefficient or only concerned with the target-agnostic design of 1D sequences. In this paper, we propose a generative…

2024

Generalist Equivariant Transformer Towards 3D Molecular Interaction Learning

ICML 2024poster

Many processes in biology and drug discovery involve various 3D interactions between molecules, such as protein and protein, protein and small molecule, etc. Given that different molecules are usually represented in different granularity, existing methods usually encode each type of molecules indepe…

2024

Improving Equivariant Graph Neural Networks on Large Geometric Graphs via Virtual Nodes Learning

ICML 2024poster

Equivariant Graph Neural Networks (GNNs) have made remarkable success in a variety of scientific applications. However, existing equivariant GNNs encounter the efficiency issue for large geometric graphs and perform poorly if the input is reduced to sparse local graph for speed acceleration. In this…

Cited by 5SourcePDFScholar
2024

Learning Superconductivity from Ordered and Disordered Material Structures

NeurIPS 2024poster

Superconductivity is a fascinating phenomenon observed in certain materials under certain conditions. However, some critical aspects of it, such as the relationship between superconductivity and materials' chemical/structural features, still need to be understood. Recent successes of data-driven app…

Cited by 1SourcePDFScholar
2024

Rigid Protein-Protein Docking via Equivariant Elliptic-Paraboloid Interface Prediction

ICLR 2024poster

The study of rigid protein-protein docking plays an essential role in a variety of tasks such as drug design and protein engineering. Recently, several learning-based methods have been proposed for the task, exhibiting much faster docking speed than those computational methods. In this paper, we pro…

2024

Subequivariant Reinforcement Learning in 3D Multi-Entity Physical Environments

ICML 2024poster

Learning policies for multi-entity systems in 3D environments is far more complicated against single-entity scenarios, due to the exponential expansion of the global state space as the number of entities increases. One potential solution of alleviating the exponential complexity is dividing the glob…

Cited by 0SourcePDFScholar
2023

Compacting Binary Neural Networks by Sparse Kernel Selection

CVPR 2023poster

Binary Neural Network (BNN) represents convolution weights with 1-bit values, which enhances the efficiency of storage and computation. This paper is motivated by a previously revealed phenomenon that the binary kernels in successful BNNs are nearly power-law distributed: their values are mostly clu…

Cited by 7SourcePDFScholar
2023

Crystal Structure Prediction by Joint Equivariant Diffusion

NeurIPS 2023poster

Crystal Structure Prediction (CSP) is crucial in various scientific disciplines. While CSP can be addressed by employing currently-prevailing generative models (**e.g.** diffusion models), this task encounters unique challenges owing to the symmetric geometry of crystal structures---the invariance o…

2023

Energy-Motivated Equivariant Pretraining for 3D Molecular Graphs

AAAI 2023technical

Pretraining molecular representation models without labels is fundamental to various applications. Conventional methods mainly process 2D molecular graphs and focus solely on 2D tasks, making their pretrained models incapable of characterizing 3D geometry and thus defective for downstream 3D tasks.…

2023

Equivariant Spatio-Temporal Attentive Graph Networks to Simulate Physical Dynamics

NeurIPS 2023poster

Learning to represent and simulate the dynamics of physical systems is a crucial yet challenging task. Existing equivariant Graph Neural Network (GNN) based methods have encapsulated the symmetry of physics, \emph{e.g.}, translations, rotations, etc, leading to better generalization ability. Neverth…

2023

Planning Assembly Sequence with Graph Transformer

ICRA 2023poster

Assembly Sequence Planning (ASP) is the essential process for modern manufacturing, proven to be NP-complete thus its effective and efficient solution has been a challenge for researchers in the field. In this paper, we present a graph-transformer based framework for the ASP problem which is trained…

Cited by 23SourcecodeScholar
2023

Subequivariant Graph Reinforcement Learning in 3D Environments

ICML 2023oral

Learning a shared policy that guides the locomotion of different agents is of core interest in Reinforcement Learning (RL), which leads to the study of morphology-agnostic RL. However, existing benchmarks are highly restrictive in the choice of starting point and target point, constraining the movem…

2022

Benefits of Permutation-Equivariance in Auction Mechanisms

NeurIPS 2022accept

Designing an incentive-compatible auction mechanism that maximizes the auctioneer's revenue while minimizes the bidders’ ex-post regret is an important yet intricate problem in economics. Remarkable progress has been achieved through learning the optimal auction mechanism by neural networks. In this…

Cited by 12SourcePDFScholar
2022

Bridged Transformer for Vision and Point Cloud 3D Object Detection

CVPR 2022poster

3D object detection is a crucial research topic in computer vision, which usually uses 3D point clouds as input in conventional setups. Recently, there is a trend of leveraging multiple sources of input data, such as complementing the 3D point cloud with 2D images that often have richer color and fe…

Cited by 53PDFScholar
2022

Equivariant Graph Mechanics Networks with Constraints

ICLR 2022poster

Learning to reason about relations and dynamics over multiple interacting objects is a challenging topic in machine learning. The challenges mainly stem from that the interacting systems are exponentially-compositional, symmetrical, and commonly geometrically-constrained. Current methods, particular…

2022

Learning Active Camera for Multi-Object Navigation

NeurIPS 2022accept

Getting robots to navigate to multiple objects autonomously is essential yet difficult in robot applications. One of the key challenges is how to explore environments efficiently with camera sensors only. Existing navigation methods mainly focus on fixed cameras and few attempts have been made to na…

Cited by 27SourcePDFScholar
2022

Learning Physical Dynamics with Subequivariant Graph Neural Networks

NeurIPS 2022accept

Graph Neural Networks (GNNs) have become a prevailing tool for learning physical dynamics. However, they still encounter several challenges: 1) Physical laws abide by symmetry, which is a vital inductive bias accounting for model generalization and should be incorporated into the model design. Exis…

Cited by 46SourcePDFScholar
2022

Molecule Generation by Principal Subgraph Mining and Assembling

NeurIPS 2022accept

Molecule generation is central to a variety of applications. Current attention has been paid to approaching the generation task as subgraph prediction and assembling. Nevertheless, these methods usually rely on hand-crafted or external subgraph construction, and the subgraph assembling depends solel…

Cited by 62SourcePDFScholar
2022

Multimodal Token Fusion for Vision Transformers

CVPR 2022poster

Many adaptations of transformers have emerged to address the single-modal vision tasks, where self-attention modules are stacked to handle input sources like images. Intuitively, feeding multiple modalities of data to vision transformers could improve the performance, yet the inner-modal attentive w…

Cited by 217PDFcodeScholar
2022

SNAKE: Shape-aware Neural 3D Keypoint Field

NeurIPS 2022accept

Detecting 3D keypoints from point clouds is important for shape reconstruction, while this work investigates the dual question: can shape reconstruction benefit 3D keypoint detection? Existing methods either seek salient features according to statistics of different orders or learn to predict keypoi…

2022

Sim2Real Object-Centric Keypoint Detection and Description

AAAI 2022technical

Keypoint detection and description play a central role in computer vision. Most existing methods are in the form of scene-level prediction, without returning the object classes of different keypoints. In this paper, we propose the object-centric formulation, which, beyond the conventional setting, r…

Cited by 9SourcePDFScholar
2022

Sound Adversarial Audio-Visual Navigation

ICLR 2022poster

Audio-visual navigation task requires an agent to find a sound source in a realistic, unmapped 3D environment by utilizing egocentric audio-visual observations. Existing audio-visual navigation works assume a clean environment that solely contains the target sound, which, however, would not be suita…

2022

When to Update Your Model: Constrained Model-based Reinforcement Learning

NeurIPS 2022accept

Designing and analyzing model-based RL (MBRL) algorithms with guaranteed monotonic improvement has been challenging, mainly due to the interdependence between policy optimization and model learning. Existing discrepancy bounds generally ignore the impacts of model shifts, and their corresponding alg…

2021

Adversarial Option-Aware Hierarchical Imitation Learning

ICML 2021spotlight

It has been a challenge to learning skills for an agent from long-horizon unannotated demonstrations. Existing approaches like Hierarchical Imitation Learning(HIL) are prone to compounding errors or suboptimal solutions. In this paper, we propose Option-GAIL, a novel method to learn skills at long h…

2021

Knowledge Representation Learning with Contrastive Completion Coding

EMNLP 2021finding

Knowledge representation learning (KRL) has been used in plenty of knowledge-driven tasks. Despite fruitfully progress, existing methods still suffer from the immaturity on tackling potentially-imperfect knowledge graphs and highly-imbalanced positive-negative instances during training, both of whic…

2020

Deep Multimodal Fusion by Channel Exchanging

NeurIPS 2020poster

Deep multimodal fusion by using multiple sources of data for classification or regression has exhibited a clear advantage over the unimodal counterpart on various applications. Yet, current methods including aggregation-based and alignment-based fusion are still inadequate in balancing the trade-off…

2020

DropEdge: Towards Deep Graph Convolutional Networks on Node Classification

ICLR 2020poster

Over-fitting and over-smoothing are two main obstacles of developing deep Graph Convolutional Networks (GCNs) for node classification. In particular, over-fitting weakens the generalization ability on small dataset, while over-smoothing impedes model training by isolating output representations from…

Cited by 1783SourcecodeScholar
2020

Reusing Discriminators for Encoding: Towards Unsupervised Image-to-Image Translation

CVPR 2020poster

Unsupervised image-to-image translation is a central task in computer vision. Current translation frameworks will abandon the discriminator once the training process is completed. This paper contends a novel role of the discriminator by reusing it for encoding the images of the target domain. The pr…

Cited by 259PDFcodeScholar
2020

Self-Supervised Graph Transformer on Large-Scale Molecular Data

NeurIPS 2020poster

How to obtain informative representations of molecules is a crucial prerequisite in AI-driven drug design and discovery. Recent researches abstract molecules as graphs and employ Graph Neural Networks (GNNs) for molecular representation learning. Nevertheless, two issues impede the usage of GNNs in…

2019

A Fast and Accurate One-Stage Approach to Visual Grounding

ICCV 2019oral

We propose a simple, fast, and accurate one-stage approach to visual grounding, inspired by the following insight. The performances of existing propose-and-rank two-stage methods are capped by the quality of the region candidates they propose in the first stage --- if none of the candidates could co…

Cited by 437PDFcodeScholar
2019

Graph Convolutional Networks for Temporal Action Localization

ICCV 2019poster

Most state-of-the-art action localization systems process each action proposal individually, without explicitly exploiting their relations during learning. However, the relations between proposals actually play an important role in action localization, since a meaningful action always consists of mu…

Cited by 640PDFcodeScholar
2019

Imitation Learning from Observations by Minimizing Inverse Dynamics Disagreement

NeurIPS 2019spotlight

This paper studies Learning from Observations (LfO) for imitation learning with access to state-only demonstrations. In contrast to Learning from Demonstration (LfD) that involves both action and state supervisions, LfO is more practical in leveraging previously inapplicable resources (e.g., videos)…

Cited by 90SourcePDFScholar
2019

Progressive Feature Alignment for Unsupervised Domain Adaptation

CVPR 2019poster

Unsupervised domain adaptation (UDA) transfers knowledge from a label-rich source domain to a fully-unlabeled target domain. To tackle this task, recent approaches resort to discriminative domain transfer in virtue of pseudo-labels to enforce the class-level distribution alignment across the source…

Cited by 544PDFScholar
2018

Adaptive Sampling Towards Fast Graph Representation Learning

NeurIPS 2018poster

Graph Convolutional Networks (GCNs) have become a crucial tool on learning representations of graph vertices. The main challenge of adapting GCNs on large-scale graphs is the scalability issue that it incurs heavy cost both in computation and memory due to the uncontrollable neighborhood expansion a…

Cited by 640SourcePDFScholar
2018

Deep Feature Pyramid Reconfiguration for Object Detection

ECCV 2018poster

State-of-the-art object detectors usually learn multi-scale representations to get better results by employing feature pyramids. However, the current designs for feature pyramids are still inefficient to integrate the semantic information over different scales. In this paper, we begin by investigati…

2018

End-to-End Learning of Motion Representation for Video Understanding

CVPR 2018poster

Despite the recent success of end-to-end learned representations, hand-crafted optical flow features are still widely used in video analysis tasks. To fill this gap, we propose TVNet, a novel end-to-end trainable neural network, to learn optical-flow-like features from data. TVNet subsumes a specifi…

Cited by 265SourcePDFScholar
2018

Weakly Supervised Dense Event Captioning in Videos

NeurIPS 2018poster

Dense event captioning aims to detect and describe all events of interest contained in a video. Despite the advanced development in this area, existing methods tackle this task by making use of dense temporal annotations, which is dramatically source-consuming. This paper formulates a new problem: w…

2017

Efficient Optimization for Linear Dynamical Systems with Applications to Clustering and Sparse Coding

NeurIPS 2017poster

Linear Dynamical Systems (LDSs) are fundamental tools for modeling spatio-temporal data in various disciplines. Though rich in modeling, analyzing LDSs is not free of difficulty, mainly because LDSs do not comply with Euclidean geometry and hence conventional learning techniques can not be applied d…

Cited by 12SourcePDFScholar
2016

Sparse Coding and Dictionary Learning With Linear Dynamical Systems

CVPR 2016oral

Linear Dynamical Systems (LDSs) are the fundamental tools for encoding spatio-temporal data in various disciplines. To enhance the performance of LDSs, in this paper, we address the challenging issue of performing sparse coding on the space of LDSs, where both data and dictionary atoms are LDSs. Rat…

Cited by 38PDFScholar