← Search

Bing Su

43 accepted papers

2026

DrugTrail: Explainable Drug Discovery via Structured Reasoning and Druggability‑Tailored Preference Optimization

ICLR 2026poster

Machine learning promises to revolutionize drug discovery, but its "black-box" nature and narrow focus limit adoption by experts. While Large Language Models (LLMs) offer a path forward with their broad knowledge and interactivity, existing methods remain data-intensive and lack transparent reasonin…

Cited by 0SourceScholar
2026

Extending Sequence Length is Not All You Need: Effective Integration of Multimodal Signals for Gene Expression Prediction

ICLR 2026oral

Gene expression prediction, which predicts mRNA expression levels from DNA sequences, presents significant challenges. Previous works often focus on extending input sequence length to locate distal enhancers, which may influence target genes from hundreds of kilobases away. Our work first reveals th…

Cited by 0SourcecodeScholar
2026

From Holo Pockets to Electron Density: GPT-style Drug Design with Density

ICML 2026poster

Recent advances in generative modeling have enabled significant progress in structure-based drug design (SBDD). Existing methods typically condition molecule generation on empty binding pockets from holo complexes, overlooking informative components such as the filler (ligands and solvent). Here, we…

Cited by 0SourceScholar
2026

Null-Space Filtering for Data-free Continual Model Merging: Preserving Transparency, Promoting Fidelity

ICLR 2026poster

Data-free continual model merging (DFCMM) aims to fuse independently fine-tuned models into a single backbone that evolves with incoming tasks without accessing task data. This paper formulate two fundamental desiderata for DFCMM: transparency, avoiding interference with earlier tasks, and fidelity,…

Cited by 0SourceScholar
2026

Optical Flow Matching: Reframing Optical Flow as Continuous Transport Dynamics

CVPR 2026

Modern optical flow estimation, though empowered by recent deep neural architectures, remains rooted in the discrete correspondence paradigm inherited from classical vision. Most networks infer frame-to-frame displacements, capturing where pixels move but not how motion evolves continuously through

Cited by 0SourcecodeScholar
2026

Pareto-Guided Optimal Transport for Multi-Reward Alignment

ICML 2026poster

Text-to-image generation models have achieved remarkable progress in preference optimization, yet achieving robust alignment across diverse reward models remains a significant challenge. Existing multi-reward fusion approaches rely on weighted summation, which is costly to tune and insufficient for …

Cited by 0SourceScholar
2026

ReAttnCLIP: Training-Free Open-Vocabulary Remote Sensing Image Segmentation via Re-defined Attention in CLIP

CVPR 2026

Remote sensing image segmentation is critical for a range of applications, including natural disaster monitoring and precision agriculture. Open-vocabulary segmentation enhances flexibility by removing fixed category constraints, enabling more fine-grained and adaptive scene understanding. Unlike CL

Cited by 0SourceScholar
2025

A Plug-and-Play Bregman ADMM Module for Inferring Event Branches in Temporal Point Processes

AAAI 2025technical

An event sequence generated by a temporal point process is often associated with a hidden and structured event branching process that captures the triggering relations between its historical and current events. In this study, we design a new plug-and-play module based on the Bregman ADMM (BADMM) al…

2025

DenoiseVAE: Learning Molecule-Adaptive Noise Distributions for Denoising-based 3D Molecular Pre-training

ICLR 2025poster

Denoising learning of 3D molecules learns molecular representations by imposing noises into the equilibrium conformation and predicting the added noises to recover the equilibrium conformation, which essentially captures the information of molecular force fields. Due to the specificity of Potential…

Cited by 2SourcePDFScholar
2025

Enhancing Reward Models for High-quality Image Generation: Beyond Text-Image Alignment

ICCV 2025poster

Contemporary image generation systems have achieved high fidelity and superior aesthetic quality beyond basic text-image alignment. However, existing evaluation frameworks have failed to evolve in parallel. This study reveals that human preference reward models fine-tuned based on CLIP and BLIP arch…

2025

Large Language-Geometry Model: When LLM meets Equivariance

ICML 2025poster

Accurately predicting 3D structures and dynamics of physical systems is crucial in scientific applications. Existing approaches that rely on geometric Graph Neural Networks (GNNs) effectively enforce $\mathrm{E}(3)$-equivariance, but they often fail in leveraging extensive broader information. While…

Cited by 4SourcePDFScholar
2025

Rethinking the Bias of Foundation Model under Long-tailed Distribution

ICML 2025poster

Long-tailed learning has garnered increasing attention due to its practical significance. Among the various approaches, the fine-tuning paradigm has gained considerable interest with the advent of foundation models. However, most existing methods primarily focus on leveraging knowledge from these mo…

Cited by 0SourcePDFScholar
2025

STOP: Integrated Spatial-Temporal Dynamic Prompting for Video Understanding

CVPR 2025poster

Pre-trained on tremendous image-text pairs, vision-language models like CLIP have demonstrated promising zero-shot generalization across numerous image-based tasks. However, extending these capabilities to video tasks remains challenging due to limited labeled video data and high training costs. Rec…

2025

Temporal Dynamics Decoupling with Inverse Processing for Enhancing Human Motion Prediction

ICASSP 2025accepted

Exploring the bridge between historical and future motion behaviors remains a central challenge in human motion prediction. While most existing methods incorporate a reconstruction task as an auxiliary task into the decoder, thereby improving the modeling of spatio-temporal dependencies, they overlo…

Cited by 0SourceScholar
2024

Domain-Adaptive and Subgroup-Specific Cascaded Temperature Regression for Out-of-Distribution Calibration

ICASSP 2024accepted

Although deep neural networks yield high classification accuracy given sufficient training data, their predictions are typically overconfident or under-confident, i.e., the prediction confidences cannot truly reflect the accuracy. Post-hoc calibration tackles this problem by calibrating the predicti…

Cited by 0SourceScholar
2024

Dynamic Prompt Optimizing for Text-to-Image Generation

CVPR 2024poster

Text-to-image generative models specifically those based on diffusion models like Imagen and Stable Diffusion have made substantial advancements. Recently there has been a surge of interest in the delicate refinement of text prompts. Users assign weights or alter the injection time steps of certain…

2024

Unlocking the Power of Spatial and Temporal Information in Medical Multimodal Pre-training

ICML 2024poster

Medical vision-language pre-training methods mainly leverage the correspondence between paired medical images and radiological reports. Although multi-view spatial images and temporal sequences of image-report pairs are available in off-the-shelf multi-modal medical datasets, most existing methods h…

2023

Exploring Temporal Concurrency for Video-Language Representation Learning

ICCV 2023poster

Paired video and language data is naturally temporal concurrency, which requires the modeling of the temporal dynamics within each modality and the temporal alignment across modalities simultaneously. However, most existing video-language representation learning methods only focus on discrete semant…

Cited by 4PDFcodeScholar
2023

Modeling Video As Stochastic Processes for Fine-Grained Video Representation Learning

CVPR 2023highlight

A meaningful video is semantically coherent and changes smoothly. However, most existing fine-grained video representation learning methods learn frame-wise features by aligning frames across videos or exploring relevance between multiple views, neglecting the inherent dynamic process of each video.…

2023

Preformer: Predictive Transformer with Multi-Scale Segment-Wise Correlations for Long-Term Time Series Forecasting

ICASSP 2023accepted

In long-term time series forecasting, most Transformer-based methods adopt the standard point-wise attention mechanism, which not only has high complexity but also cannot explicitly capture the predictive dependencies from contexts since the corresponding key and value are transformed from the same…

Cited by 0SourceScholar
2023

Self-Supervised Action Representation Learning from Partial Spatio-Temporal Skeleton Sequences

AAAI 2023technical

Self-supervised learning has demonstrated remarkable capability in representation learning for skeleton-based action recognition. Existing methods mainly focus on applying global data augmentation to generate different views of the skeleton sequence for contrastive learning. However, due to the rich…

2023

Transfer Knowledge From Head to Tail: Uncertainty Calibration Under Long-Tailed Distribution

CVPR 2023poster

How to estimate the uncertainty of a given model is a crucial problem. Current calibration techniques treat different classes equally and thus implicitly assume that the distribution of training data is balanced, but ignore the fact that real-world data often follows a long-tailed distribution. In t…

2022

Interventional Contrastive Learning with Meta Semantic Regularizer

ICML 2022spotlight

Contrastive learning (CL)-based self-supervised learning models learn visual representations in a pairwise manner. Although the prevailing CL model has achieved great progress, in this paper, we uncover an ever-overlooked phenomenon: When the CL model is trained with full images, the performance tes…

Cited by 34SourcePDFScholar
2022

MetAug: Contrastive Learning via Meta Feature Augmentation

ICML 2022spotlight

What matters for contrastive learning? We argue that contrastive learning heavily relies on informative features, or “hard” (positive or negative) features. Early works include more informative features by applying complex data augmentations and large batch size or memory bank, and recent works desi…

Cited by 41SourcePDFScholar
2022

MetaMask: Revisiting Dimensional Confounder for Self-Supervised Learning

NeurIPS 2022accept

As a successful approach to self-supervised learning, contrastive learning aims to learn invariant information shared among distortions of the input sample. While contrastive learning has yielded continuous advancements in sampling strategy and architecture design, it still remains two persistent de…

Cited by 16SourcePDFScholar
2022

Optimal Partial Transport Based Sentence Selection for Long-form Document Matching

COLING 2022main

One typical approach to long-form document matching is first conducting alignment between cross-document sentence pairs, and then aggregating all of the sentence-level matching signals. However, this approach could be problematic because the alignment between documents is partial — despite two docum…

2022

SemMAE: Semantic-Guided Masking for Learning Masked Autoencoders

NeurIPS 2022accept

Recently, significant progress has been made in masked image modeling to catch up to masked language modeling. However, unlike words in NLP, the lack of semantic decomposition of images still makes masked autoencoding (MAE) different between vision and language. In this paper, we explore a potential…

2022

Temporal Alignment Prediction for Supervised Representation Learning and Few-Shot Sequence Classification

ICLR 2022poster

Explainable distances for sequence data depend on temporal alignment to tackle sequences with different lengths and local variances. Most sequence alignment methods infer the optimal alignment by solving an optimization problem under pre-defined feasible alignment constraints, which not only is time…

2020

Online Joint Multi-Metric Adaptation From Frequent Sharing-Subset Mining for Person Re-Identification

CVPR 2020poster

Person Re-IDentification (P-RID), as an instance-level recognition problem, still remains challenging in computer vision community. Many P-RID works aim to learn faithful and discriminative features/metrics from offline training data and directly use them for the unseen online testing data. However,…

Cited by 61PDFScholar
2018

Easy Identification From Better Constraints: Multi-Shot Person Re-Identification From Reference Constraints

CVPR 2018poster

Multi-shot person re-identification (MsP-RID) utilizes multiple images from the same person to facilitate identification. Considering the fact that motion information may not be discriminative nor reliable enough for MsP-RID, this paper is focused on handling the large variations in the visual appea…

Cited by 21SourcePDFScholar