← Search

Jian Pei

31 accepted papers

2026

Cite Pretrain: Retrieval-Free Knowledge Attribution for Large Language Models

ICLR 2026poster

Trustworthy language models should provide both correct and verifiable answers. However, citations generated directly by standalone LLMs are often unreliable due to hallucinations. As a result, current systems insert citations by querying an external retriever at inference time, introducing latency,…

Cited by 0SourceScholar
2026

TrustGen: A Platform of Dynamic Benchmarking on the Trustworthiness of Generative Foundation Models

ICLR 2026poster

Generative foundation models (GenFMs), such as large language models and text-to-image systems, have demonstrated remarkable capabilities in various downstream applications. As they are increasingly deployed in high-stakes applications, assessing their trustworthiness has become both a critical nece…

Cited by 0SourceScholar
2026

pTNAS: Progressive Neural Architecture Search for Tabular Data

ICML 2026poster

Recent advances have shifted the paradigm of tabular learning toward tabular foundation models, yet their accuracy relies on a heavy inference cost that scales poorly with context size. Deep neural networks remain a highly competitive and more efficient modeling paradigm when equipped with well-desi…

Cited by 0SourceScholar
2024

Fair and Efficient Contribution Valuation for Vertical Federated Learning

ICLR 2024poster

Federated learning is an emerging technology for training machine learning models across decentralized data sources without sharing data. Vertical federated learning, also known as feature-based federated learning, applies to scenarios where data sources have the same sample IDs but different featur…

Cited by 49SourcePDFScholar
2024

Position: TrustLLM: Trustworthiness in Large Language Models

ICML 2024poster

Large language models (LLMs) have gained considerable attention for their excellent natural language processing capabilities. Nonetheless, these LLMs present many challenges, particularly in the realm of trustworthiness. This paper introduces TrustLLM, a comprehensive study of trustworthiness in LLM…

Cited by 95SourcePDFScholar
2024

ReCaLL: Membership Inference via Relative Conditional Log-Likelihoods

EMNLP 2024main

The rapid scaling of large language models (LLMs) has raised concerns about the transparency and fair use of the data used in their pretraining. Detecting such content is challenging due to the scale of the data and limited exposure of each instance during training. We propose ReCaLL (Relative Condi…

Cited by 11SourcePDFScholar
2023

A Graph Fusion Approach for Cross-Lingual Machine Reading Comprehension

AAAI 2023technical

Although great progress has been made for Machine Reading Comprehension (MRC) in English, scaling out to a large number of languages remains a huge challenge due to the lack of large amounts of annotated training data in non-English languages. To address this challenge, some recent efforts of cross-…

2023

Alleviating Over-smoothing for Unsupervised Sentence Representation

ACL 2023long

Currently, learning better unsupervised sentence representations is the pursuit of many natural language processing communities. Lots of approaches based on pre-trained language models (PLMs) and contrastive learning have achieved promising results on this task. Experimentally, we observe that the o…

2023

Bias Reduced Semidefinite Relaxation Method for Multistatic Localization in the Absence of Transmitter Position And Its Synchronization

ICASSP 2023accepted

This paper addresses the challenging problem of multistatic localization of a stationary object with a set of synchronized receivers, when the transmitter position is unknown and the synchronization with the transmitter is unavailable. Using the time delay measurements from the direct and indirect p…

Cited by 0SourceScholar
2023

Identify Event Causality with Knowledge and Analogy

AAAI 2023technical

Event causality identification (ECI) aims to identify the causal relationship between events, which plays a crucial role in deep text understanding. Due to the diversity of real-world causality events and difficulty in obtaining sufficient training data, existing ECI approaches have poor generalizab…

2023

LazyGNN: Large-Scale Graph Neural Networks via Lazy Propagation

ICML 2023poster

Recent works have demonstrated the benefits of capturing long-distance dependency in graphs by deeper graph neural networks (GNNs). But deeper GNNs suffer from the long-lasting scalability challenge due to the neighborhood explosion problem in large-scale graphs. In this work, we propose to capture…

2023

Structural Contrastive Pretraining for Cross-Lingual Comprehension

ACL 2023findings

To present, multilingual language models trained using various pre-training tasks like mask language modeling (MLM) have yielded encouraging results on a wide range of downstream tasks. Despite the promising performances, structural knowledge in cross-lingual corpus is less explored in current works…

2022

Bridging the Gap between Language Models and Cross-Lingual Sequence Labeling

NAACL 2022long

Large-scale cross-lingual pre-trained language models (xPLMs) have shown effective in cross-lingual sequence labeling tasks (xSL), such as machine reading comprehension (xMRC) by transferring knowledge from a high-resource language to low-resource languages. Despite the great success, we draw an emp…

2022

Cosine Model Watermarking against Ensemble Distillation

AAAI 2022technical

Many model watermarking methods have been developed to prevent valuable deployed commercial models from being stealthily stolen by model distillations. However, watermarks produced by most existing model watermarking methods can be easily evaded by ensemble distillation, because averaging the outpu…

Cited by 26SourcePDFScholar
2022

From Good to Best: Two-Stage Training for Cross-Lingual Machine Reading Comprehension

AAAI 2022technical

Cross-lingual Machine Reading Comprehension (xMRC) is a challenging task due to the lack of training data in low-resource languages. Recent approaches use training data only in a resource-rich language (such as English) to fine-tune large-scale cross-lingual pre-trained language models, which transf…

2022

Label-aware Multi-level Contrastive Learning for Cross-lingual Spoken Language Understanding

EMNLP 2022main

Despite the great success of spoken language understanding (SLU) in high-resource languages, it remains challenging in low-resource languages mainly due to the lack of labeled training data. The recent multilingual code-switching approach achieves better alignments of model representations across la…

2022

Lexicon-Enhanced Self-Supervised Training for Multilingual Dense Retrieval

EMNLP 2022finding

Recent multilingual pre-trained models have shown better performance in various multilingual tasks. However, these models perform poorly on multilingual retrieval tasks due to lacking multilingual training data. In this paper, we propose to mine and generate self-supervised training data based on a…

2022

Revisiting Graph Contrastive Learning from the Perspective of Graph Spectrum

NeurIPS 2022accept

Graph Contrastive Learning (GCL), learning the node representations by augmenting graphs, has attracted considerable attentions. Despite the proliferation of various graph augmentation strategies, there are still some fundamental questions unclear: what information is essentially learned by GCL? Are…

2021

Finding Representative Interpretations on Convolutional Neural Networks

ICCV 2021poster

Interpreting the decision logic behind effective deep convolutional neural networks (CNN) on images complements the success of deep learning models. However, the existing methods can only interpret some specific decision logic on individual or a small number of images. To facilitate human understand…

Cited by 11PDFScholar
2021

Knowledge-Enhanced Hierarchical Graph Transformer Network for Multi-Behavior Recommendation

AAAI 2021technical

Accurate user and item embedding learning is crucial for modern recommender systems. However, most existing recommendation techniques have thus far focused on modeling users' preferences over singular type of user-item interactions. Many practical recommendation scenarios involve multi-typed user in…

2021

Learning from Multiple Noisy Augmented Data Sets for Better Cross-Lingual Spoken Language Understanding

EMNLP 2021main

Lack of training data presents a grand challenge to scaling out spoken language understanding (SLU) to low-resource languages. Although various data augmentation approaches have been proposed to synthesize training data in low-resource target languages, the augmented data sets are often noisy, and t…

2021

Personalized Cross-Silo Federated Learning on Non-IID Data

AAAI 2021technical

Non-IID data present a tough challenge for federated learning. In this paper, we explore a novel idea of facilitating pairwise collaborations between clients with similar data. We propose FedAMP, a new method employing federated attentive message passing to facilitate similar clients to collaborate…

Cited by 744SourcePDFScholar
2021

Reasoning over Entity-Action-Location Graph for Procedural Text Understanding

ACL 2021long

Procedural text understanding aims at tracking the states (e.g., create, move, destroy) and locations of the entities mentioned in a given paragraph. To effectively track the states and locations, it is essential to capture the rich semantic relations between entities, actions, and locations in the…

2021

Reinforced Multi-Teacher Selection for Knowledge Distillation

AAAI 2021technical

In natural language processing (NLP) tasks, slow inference speed and huge footprints in GPU usage remain the bottleneck of applying pre-trained deep models in production. As a popular method for model compression, knowledge distillation transfers knowledge from one or multiple large (teacher) models…

2021

Robust Counterfactual Explanations on Graph Neural Networks

NeurIPS 2021poster

Massive deployment of Graph Neural Networks (GNNs) in high-stake applications generates a strong demand for explanations that are robust to noise and align well with human intuition. Most existing methods generate explanations by identifying a subgraph of an input graph that has a strong correlation…

Cited by 139SourcePDFScholar
2020

A Graph Representation of Semi-structured Data for Web Question Answering

COLING 2020main

The abundant semi-structured data on the Web, such as HTML-based tables and lists, provide commercial search engines a rich information source for question answering (QA). Different from plain text passages in Web documents, Web tables and lists have inherent structures, which carry semantic correla…

Cited by 15SourcePDFScholar
2020

Cross-lingual Machine Reading Comprehension with Language Branch Knowledge Distillation

COLING 2020main

Cross-lingual Machine Reading Comprehension (CLMRC) remains a challenging problem due to the lack of large-scale annotated datasets in low-source languages, such as Arabic, Hindi, and Vietnamese. Many previous approaches use translation data by translating from a rich-source language, such as Englis…

Cited by 20SourcePDFScholar
2020

Discrete Model Compression With Resource Constraint for Deep Neural Networks

CVPR 2020poster

In this paper, we target to address the problem of compression and acceleration of Convolutional Neural Networks (CNNs). Specifically, we propose a novel structural pruning method to obtain a compact CNN with strong discriminative power. To find such networks, we propose an efficient discrete optimi…

Cited by 100PDFScholar
2020

Sinkhorn Regression

IJCAI 2020poster

This paper introduces a novel Robust Regression (RR) model, named Sinkhorn regression, which imposes Sinkhorn distances on both loss function and regularization. Traditional RR methods target at searching for an element-wise loss function (e.g., Lp-norm) to characterize the errors such that ou…

Cited by 0SourcePDFScholar