← Search

Xingyi Cheng

8 accepted papers

2024

MSAGPT: Neural Prompting Protein Structure Prediction via MSA Generative Pre-Training

NeurIPS 2024poster

Multiple Sequence Alignment (MSA) plays a pivotal role in unveiling the evolutionary trajectories of protein families. The accuracy of protein structure predictions is often compromised for protein sequences that lack sufficient homologous information to construct high-quality MSA. Although various…

2024

Training Compute-Optimal Protein Language Models

NeurIPS 2024spotlight

We explore optimally training protein language models, an area of significant interest in biological research where guidance on best practices is limited. Most models are trained with extensive compute resources until performance gains plateau, focusing primarily on increasing model sizes rather tha…

2023

Injecting Multimodal Information into Rigid Protein Docking via Bi-level Optimization

NeurIPS 2023poster

The structure of protein-protein complexes is critical for understanding binding dynamics, biological mechanisms, and intervention strategies. Rigid protein docking, a fundamental problem in this field, aims to predict the 3D structure of complexes from their unbound states without conformational ch…

Cited by 6SourcePDFScholar
2023

Revisiting Out-of-distribution Robustness in NLP: Benchmarks, Analysis, and LLMs Evaluations

NeurIPS 2023poster

This paper reexamines the research on out-of-distribution (OOD) robustness in the field of NLP. We find that the distribution shift settings in previous studies commonly lack adequate challenges, hindering the accurate evaluation of OOD robustness. To address these issues, we propose a benchmark con…

2023

Won’t Get Fooled Again: Answering Questions with False Premises

ACL 2023long

Pre-trained language models (PLMs) have shown unprecedented potential in various fields, especially as the backbones for question-answering (QA) systems. However, they tend to be easily deceived by tricky questions such as “How many eyes does the sun have?”. Such frailties of PLMs often allude to th…

2023

xTrimoGene: An Efficient and Scalable Representation Learner for Single-Cell RNA-Seq Data

NeurIPS 2023poster

Advances in high-throughput sequencing technology have led to significant progress in measuring gene expressions at the single-cell level. The amount of publicly available single-cell RNA-seq (scRNA-seq) data is already surpassing 50M records for humans with each record measuring 20,000 genes. This…

Cited by 28SourcePDFScholar
2020

Towards Fast and Accurate Neural Chinese Word Segmentation with Multi-Criteria Learning

COLING 2020main

The ambiguous annotation criteria lead to divergence of Chinese Word Segmentation (CWS) datasets in various granularities. Multi-criteria Chinese word segmentation aims to capture various annotation criteria among datasets and leverage their common underlying knowledge. In this paper, we propose a d…