← Search

Wenxuan Xie

12 accepted papers

2026

Drugging the Undruggable: Benchmarking and Modeling Fragment-Based Screening

ICLR 2026poster

A significant portion of disease-relevant proteins remain undruggable due to shallow, flexible, or otherwise ill-defined binding pockets that hinder conventional molecule screening. Fragment-based drug discovery (FBDD) offers a promising alternative, as small, low-complexity fragments can flexibly e…

Cited by 0SourceScholar
2025

CIDD: Collaborative Intelligence for Structure-Based Drug Design Empowered by LLMs

NeurIPS 2025poster

Structure-guided molecular generation is pivotal in early-stage drug discovery, enabling the design of compounds tailored to specific protein targets. However, despite recent advances in 3D generative modeling, particularly in improving docking scores, these methods often produce rare and intrinsica…

Cited by 0SourceScholar
2025

Motion-Guided Dual-Camera Tracker for Endoscope Tracking and Motion Analysis in a Mechanical Gastric Simulator

ICRA 2025

Flexible endoscope motion tracking and analysis in mechanical simulators have proven useful for endoscopy training. Common motion tracking methods based on electromagnetic tracker are however limited by their high cost and material susceptibility. In this work, the motion-guided dual-camera vision t

Cited by 11SourcecodeScholar
2025

Self-Sufficient 5-DoF Discrete Global Localization for Magnetically-Actuated Endoscope in Bronchoscopy

ICRA 2025

Existing sensor-based global localization methods limit the miniaturization potential of magnetically-actuated endoscopes (MAE) while localization based on external medical imaging demands accurate registration and imposes a variety of modality-specific challenges during continuous image acquisition

Cited by 0SourceScholar
2025

Towards Practical Real-Time Neural Video Compression

CVPR 2025poster

We introduce a practical real-time neural video codec (NVC) designed to deliver high compression ratio, low latency and broad versatility. In practice, the coding speed of NVCs depends on 1) computational costs, and 2) non-computational operational costs, such as memory I/O and the number of functio…

2024

Slot-VLM: Object-Event Slots for Video-Language Modeling

NeurIPS 2024poster

Video-Language Models (VLMs), powered by the advancements in Large Language Models (LLMs), are charting new frontiers in video understanding. A pivotal challenge is the development of an effective method to encapsulate video content into a set of representative tokens to align with LLMs. In this wor…

Cited by 0SourcePDFScholar
2024

Text Grouping Adapter: Adapting Pre-trained Text Detector for Layout Analysis

CVPR 2024poster

Significant progress has been made in scene text detection models since the rise of deep learning but scene text layout analysis which aims to group detected text instances as paragraphs has not kept pace. Previous works either treated text detection and grouping using separate models or train a mod…

Cited by 1SourcePDFScholar
2023

Unifying Layout Generation With a Decoupled Diffusion Model

CVPR 2023poster

Layout generation aims to synthesize realistic graphic scenes consisting of elements with different attributes including category, size, position, and between-element relation. It is a crucial task for reducing the burden on heavy-duty graphic design works for formatted scenes, e.g., publications, d…

Cited by 45SourcePDFScholar
2022

Sparse MLP for Image Recognition: Is Self-Attention Really Necessary?

AAAI 2022technical

Transformers have sprung up in the field of computer vision. In this work, we explore whether the core self-attention module in Transformer is the key to achieving excellent performance in image recognition. To this end, we build an attention-free network called sMLPNet based on the existing MLP-bas…

2021

Unsupervised Visual Representation Learning by Tracking Patches in Video

CVPR 2021poster

Inspired by the fact that human eyes continue to develop tracking ability in early and middle childhood, we propose to use tracking as a proxy task for a computer vision system to learn the visual representations. Modelled on the Catch game played by the children, we design a Catch-the-Patch (CtP) g…

Cited by 31PDFcodeScholar
2020

Joint Time-Frequency and Time Domain Learning for Speech Enhancement

IJCAI 2020poster

For single-channel speech enhancement, both time-domain and time-frequency-domain methods have their respective pros and cons. In this paper, we present a cross-domain framework named TFT-Net, which takes time-frequency spectrogram as input and produces time-domain waveform as output. Such a framewo…

Cited by 0SourcePDFScholar