← Search

Weijian Li

14 accepted papers

2026

StarEmbed: Benchmarking Time Series Foundation Models on Astronomical Observations of Variable Stars

ICML 2026poster

Time series foundation models (TSFMs) are increasingly adopted as general-purpose time series learners. Although their training corpora are vast, they exclude peta-scale astronomical time series that exhibit unique challenges (e.g., irregular sampling, multiple variates, and heteroskedasticity) and …

Cited by 0SourceScholar
2024

BiSHop: Bi-Directional Cellular Learning for Tabular Data with Generalized Sparse Modern Hopfield Model

ICML 2024poster

We introduce the **Bi**-Directional **S**parse **Hop**field Network (**BiSHop**), a novel end-to-end framework for tabular learning. BiSHop handles the two major challenges of deep tabular learning: non-rotationally invariant data structure and feature sparsity in tabular data. Our key motivation co…

2024

DNABERT-2: Efficient Foundation Model and Benchmark For Multi-Species Genomes

ICLR 2024poster

Decoding the linguistic intricacies of the genome is a crucial problem in biology, and pre-trained foundational models such as DNABERT and Nucleotide Transformer have made significant strides in this area. Existing works have largely hinged on k-mer, fixed-length permutations of A, T, C, and G, as t…

2024

Multivariate Time Series Forecasting By Graph Attention Networks With Theoretical Guarantees

AISTATS 2024poster

Multivariate time series forecasting (MTSF) aims to predict future values of multiple variables based on past values of multivariate time series, and has been applied in fields including traffic flow prediction, stock price forecasting, and anomaly detection. Capturing the inter-dependencies among m…

Cited by 4SourcePDFScholar
2024

Outlier-Efficient Hopfield Layers for Large Transformer-Based Models

ICML 2024poster

We introduce an Outlier-Efficient Modern Hopfield Model (termed `OutEffHop`) and use it to address the outlier inefficiency problem of training gigantic transformer-based models. Our main contribution is a novel associative memory model facilitating _outlier-efficient_ associative memory retrievals.…

2024

STanHop: Sparse Tandem Hopfield Model for Memory-Enhanced Time Series Prediction

ICLR 2024poster

We present **STanHop-Net** (**S**parse **Tan**dem **Hop**field **Net**work) for multivariate time series prediction with memory-enhanced capabilities. At the heart of our approach is **STanHop**, a novel Hopfield-based neural network block, which sparsely learns and stores both temporal and cross-se…

Cited by 47SourcePDFScholar
2023

DocTr: Document Transformer for Structured Information Extraction in Documents

ICCV 2023poster

We present a new formulation for structured information extraction (SIE) from visually rich documents. We address the limitations of existing IOB tagging and graph-based formulations, which are either overly reliant on the correct ordering of input text or struggle with decoding a complex graph. Ins…

Cited by 23PDFScholar
2023

Unsupervised Self-Driving Attention Prediction via Uncertainty Mining and Knowledge Embedding

ICCV 2023poster

Predicting attention regions of interest is an important yet challenging task for self-driving systems. Existing methodologies rely on large-scale labeled traffic datasets that are labor-intensive to obtain. Besides, the huge domain gap between natural scenes and traffic scenes in current datasets a…

Cited by 6PDFcodeScholar
2022

Communication-Efficient Topologies for Decentralized Learning with $O(1)$ Consensus Rate

NeurIPS 2022accept

Decentralized optimization is an emerging paradigm in distributed learning in which agents achieve network-wide solutions by peer-to-peer communication without the central server. Since communication tends to be slower than computation, when each agent communicates with only a few neighboring agent…

2021

Learning Bias-Invariant Representation by Cross-Sample Mutual Information Minimization

ICCV 2021poster

Deep learning algorithms mine knowledge from the training data and thus would likely inherit the dataset's bias information. As a result, the obtained model would generalize poorly and even mislead the decision process in real-life applications. We propose to remove the bias information misused by t…

Cited by 50PDFScholar
2020

Anatomy-Aware Siamese Network: Exploiting Semantic Asymmetry for Accurate Pelvic Fracture Detection in X-ray Images

ECCV 2020poster

Trauma PXR are essential for instantaneous pelvic bone fracture detection. However, small, pathologically critical fractures can be missed, even by experienced clinicians, under the very limited diagnosis times allowed in urgent care. As a result, fracture CAD has very high demands to save time and…

Cited by 43SourcePDFScholar
2020

Structured Landmark Detection via Topology-Adapting Deep Graph Learning

ECCV 2020poster

Image landmark detection aims to automatically identify the locations of predefined fiducial points. Despite recent success in this field, higher-ordered structural modeling to capture implicit or explicit relationships among anatomical landmarks has not been adequately exploited. In this work, we p…

Cited by 121SourcePDFScholar
2019

Attentive Relational Networks for Mapping Images to Scene Graphs

CVPR 2019poster

Scene graph generation refers to the task of automatically mapping an image into a semantic structural graph, which requires correctly labeling each extracted object and their interaction relationships. Despite the recent success in object detection using deep learning techniques, inferring complex…

Cited by 199PDFScholar