← Search

Qingcheng Zeng

16 accepted papers

2025

CARES: Comprehensive Evaluation of Safety and Adversarial Robustness in Medical LLMs

NeurIPS 2025poster

Large language models (LLMs) are increasingly deployed in medical contexts, raising critical concerns about safety, alignment, and susceptibility to adversarial manipulation. While prior benchmarks assess model refusal capabilities for harmful prompts, they often lack clinical specificity, graded ha…

Cited by 0SourceScholar
2025

Distributed Invariant Kalman Filter for Object-Level Multi-Robot Pose SLAM

ICRA 2025

Cooperative localization and target tracking are essential for multi-robot systems to implement high-level tasks. To this end, we propose a distributed invariant Kalman filter (KF) based on covariance intersection (CI) for effective multi-robot pose estimation. The paper utilizes the object-level me

Cited by 0SourcecodeScholar
2025

Exploring Concept Depth: How Large Language Models Acquire Knowledge and Concept at Different Layers?

COLING 2025main

Large language models (LLMs) have shown remarkable performances across a wide range of tasks. However, the mechanisms by which these models encode tasks of varying complexities remain poorly understood. In this paper, we explore the hypothesis that LLMs process concepts of varying complexities in di…

2025

Good Intentions Beyond ACL: Who Does NLP for Social Good, and Where?

EMNLP 2025

The social impact of Natural Language Processing (NLP) is increasingly important, with a rising community focus on initiatives related to NLP for Social Good (NLP4SG). Indeed, in recent years, almost 20% of all papers in the ACL Anthology address topics related to social good as defined by the UN Su

2025

Leveraging Human Production-Interpretation Asymmetries to Test LLM Cognitive Plausibility

ACL 2025short

Whether large language models (LLMs) process language similarly to humans has been the subject of much theoretical and practical debate. We examine this question through the lens of the production-interpretation distinction found in human sentence processing and evaluate the extent to which instruct…

Cited by 0SourcePDFScholar
2025

MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation

EMNLP 2025

Existing large language model (LLM) evaluation benchmarks primarily focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-lingual reasoning abilities. This dual limitation makes it challenging to assess LLMs’ performance in the multilingual setting

Cited by 0SourcePDFScholar
2025

SCORE: Saturated Consensus Relocalization in Semantic Line Maps

IROS 2025

We present SCORE, a visual relocalization system that achieves unprecedented map compactness through semantically labeled 3D line maps. SCORE requires only 0.01%-0.1% of the storage needed by structure-based or learning-based baselines, while maintaining practical accuracy and comparable runtime. Th

Cited by 0SourcecodeScholar
2025

Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models

EMNLP 2025

Uncertainty quantification is essential for assessing the reliability and trustworthiness of modern AI systems. Among existing approaches, verbalized uncertainty, where models express their confidence through natural language, has emerged as a lightweight and interpretable solution in large language

Cited by 0SourcePDFScholar
2025

ThinkBench: Dynamic Out-of-Distribution Evaluation for Robust LLM Reasoning

NeurIPS 2025poster

Evaluating large language models (LLMs) poses significant challenges, particularly due to issues of data contamination and the leakage of correct answers. To address these challenges, we introduce ThinkBench, a novel evaluation framework designed to robustly evaluate the reasoning capability of LLMs…

Cited by 0SourcecodeScholar
2024

Evaluating Large Language Models on Wikipedia-Style Survey Generation

ACL 2024findings

Educational materials such as survey articles in specialized fields like computer science traditionally require tremendous expert inputs and are therefore expensive to create and update. Recently, Large Language Models (LLMs) have achieved significant success across various general tasks. However, t…

2023

GreenPLM: Cross-Lingual Transfer of Monolingual Pre-Trained Language Models at Almost No Cost

IJCAI 2023poster

Large pre-trained models have revolutionized natural language processing (NLP) research and applications, but high training costs and limited data resources have prevented their benefits from being shared equally amongst speakers of all the world's languages. To address issues of cross-linguistic ac…

2023

Large Language Models Are Partially Primed in Pronoun Interpretation

ACL 2023findings

While a large body of literature suggests that large language models (LLMs) acquire rich linguistic representations, little is known about whether they adapt to linguistic biases in a human-like way. The present study probes this question by asking whether LLMs display human-like referential biases…

2023

Masked Spectrogram Prediction for Self-Supervised Audio Pre-Training

ICASSP 2023accepted

Transformer-based models attain excellent results and generalize well when trained on sufficient amounts of data. However, constrained by the limited data available in the audio domain, most transformer-based models for audio tasks are finetuned from pre-trained models in other domains (e.g. image),…

Cited by 0SourceScholar
2022

A Survey in Automatic Irony Processing: Linguistic, Cognitive, and Multi-X Perspectives

COLING 2022main

Irony is a ubiquitous figurative language in daily communication. Previously, many researchers have approached irony from linguistic, cognitive science, and computational aspects. Recently, some progress have been witnessed in automatic irony processing due to the rapid development in deep neural mo…

Cited by 13SourcePDFScholar