← Search

Yanjun Li

15 accepted papers

2026

Apo2Mol: 3D Molecule Generation via Dynamic Pocket-Aware Diffusion Models

AAAI 2026technical

Deep generative models are rapidly advancing structure-based drug design, offering substantial promise for generating small molecule ligands that bind to specific protein targets. However, most current approaches assume a rigid protein binding pocket, neglecting the intrinsic flexibility of proteins

Cited by 3SourcePDFScholar
2026

EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question Answering

AAAI 2026technical

Recent advances in Multimodal Large Language Models (MLLMs) have significantly pushed the frontier of egocentric video question answering (EgocentricQA). However, existing benchmarks and studies are mainly limited to common daily activities such as cooking and cleaning. In contrast, real-world deplo

Cited by 0SourcePDFScholar
2026

GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and a Comprehensive Multimodal Dataset Towards General Medical AI

AAAI 2026technical

Despite significant advancements in general AI, its effectiveness in the medical domain is limited by the lack of specialized medical knowledge. To address this, we formulate GMAI-VL-5.5M, a multimodal medical dataset created by converting hundreds of specialized medical datasets with various annot

Cited by 0SourcePDFScholar
2026

Spectral Heat Flow for Conservative Token Condensation in Vision-Language Models

ICML 2026poster

Vision-Language Models (VLMs) are costly at inference time because they must process long sequences of visual tokens. Existing token pruning methods often degrade under high compression by blindly discarding information, breaking spatial structure or collapsing diversity. We propose SpecFlow, a trai…

Cited by 0SourceScholar
2025

DecoyDB: A Dataset for Graph Contrastive Learning in Protein-Ligand Binding Affinity Prediction

NeurIPS 2025poster

Predicting the binding affinity of protein-ligand complexes plays a vital role in drug discovery. Unfortunately, progress has been hindered by the lack of large-scale and high-quality binding affinity labels. The widely used PDBbind dataset has fewer than 20K labeled complexes. Self-supervised learn…

Cited by 0SourceScholar
2025

SlideChat: A Large Vision-Language Assistant for Whole-Slide Pathology Image Understanding

CVPR 2025poster

Despite the progress made by multimodal large language models (MLLMs) in computational pathology, they remain limited by a predominant focus on patch-level analysis, missing essential contextual information at the whole-slide level. The lack of large-scale instruction datasets and the gigapixel scal…

2024

Capturing Closely Interacted Two-Person Motions with Reaction Priors

CVPR 2024poster

In this paper we focus on capturing closely interacted two-person motions from monocular videos an important yet understudied topic. Unlike less-interacted motions closely interacted motions contain frequently occurring inter-human occlusions which pose significant challenges to existing capturing a…

Cited by 1SourcePDFScholar
2024

GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI

NeurIPS 2024poster

Large Vision-Language Models (LVLMs) are capable of handling diverse data types such as imaging, text, and physiological signals, and can be applied in various fields. In the medical field, LVLMs have a high potential to offer substantial assistance for diagnosis and treatment. Before that, it is cr…

2023

A Bioinspired Synthetic Nervous System Controller for Pick-and-Place Manipulation

ICRA 2023poster

The Synthetic Nervous System (SNS) is a biologically inspired neural network (NN). Due to its capability of capturing complex mechanisms underlying neural computation, an SNS model is a candidate for building compact and interpretable NN controllers for robots. Previous work on SNSs has focused on a…

Cited by 6SourceScholar
2022

Study on Time-of-Flight Estimation in Ultrasonic Well Logging Tool: Model-Driven Transfer Learning

ICASSP 2022accepted

Time-of-flight (ToF) of ultrasonic waves is essential for petroleum well logging to draw borehole-wall images. This paper proposed a method that boosted accuracy of ToFs estimation in a complex geological environment. Unlike other classical methods, the proposed one adopts a one-dimensional convolut…

Cited by 0SourceScholar
2018

Comfort-Centered Design of a Lightweight and Backdrivable Knee Exoskeleton

RA-L 2018

This letter presents design principles for comfort-centered wearable robots and their application in a lightweight and backdrivable knee exoskeleton. The mitigation of discomfort is treated as mechanical design and control issues and three solutions are proposed in this letter: 1) a new wearable str

Cited by 116SourceScholar
2017

Joint Adaptive Sparsity and Low-Rankness on the Fly: An Online Tensor Reconstruction Scheme for Video Denoising

ICCV 2017poster

Recent works on adaptive sparse and low-rank signal modeling have demonstrated their usefulness, especially in image/video processing applications. While a patch-based sparse model imposes local structure, low-rankness of the grouped patches exploits non-local correlation. Applying either approach a…

Cited by 56PDFScholar
2017

When sparsity meets low-rankness: Transform learning with non-local low-rank constraint for image restoration

ICASSP 2017accepted

Recent works on adaptive sparse signal modeling have demonstrated their usefulness in various image/video processing applications. As the popular synthesis dictionary learning methods involve NP-hard sparse coding and expensive learning steps, transform learning has recently received more interest f…

Cited by 0SourceScholar