← Search

xiang lin

5 accepted papers

2026

RMAdapter: Reconstruction-based Multi-Modal Adapter for Vision-Language Models

AAAI 2026technical

Pre-trained Vision-Language Models (VLMs), e.g. CLIP, have become essential tools in multimodal transfer learning. However, fine-tuning VLMs in few-shot scenarios poses significant challenges in balancing task-specific adaptation and generalization in the obtained model. Meanwhile, current researc

Cited by 0SourcePDFScholar
2026

SPATIA: Multimodal Generation and Prediction of Spatial Cell Phenotypes

ICML 2026poster

Understanding how cellular morphology, gene expression, and spatial context jointly shape tissue function is a central challenge in biology. Image-based spatial transcriptomics technologies now provide high-resolution measurements of cell images and gene expression profiles, but existing methods typ…

Cited by 0SourceScholar
2022

Chart-to-Text: A Large-Scale Benchmark for Chart Summarization

ACL 2022long

Charts are commonly used for exploring data and communicating insights. Generating natural language summaries from charts can be very helpful for people in inferring key insights that would otherwise require a lot of cognitive and perceptual efforts. We present Chart-to-text, a large-scale benchmark…

Cited by 151SourcePDFScholar
2022

Rethinking Self-Supervision Objectives for Generalizable Coherence Modeling

ACL 2022long

Given the claims of improved text generation quality across various pre-trained neural models, we consider the coherence evaluation of machine generated text to be one of the principal applications of coherence models that needs to be investigated. Prior work in neural coherence modeling has primari…

2021

Straight to the Gradient: Learning to Use Novel Tokens for Neural Text Generation

ICML 2021oral

Advanced large-scale neural language models have led to significant success in many language generation tasks. However, the most commonly used training objective, Maximum Likelihood Estimation (MLE), has been shown problematic, where the trained model prefers using dull and repetitive phrases. In th…