← Search

Lechen Zhang

8 accepted papers

2026

SPRIG: Improving Large Language Model Performance by System Prompt Optimization

ICLR 2026poster

Large Language Models (LLMs) have shown impressive capabilities in many scenarios, but their performance depends, in part, on the choice of prompt. Past research has focused on optimizing prompts specific to a task. However, much less attention has been given to optimizing the general instructions i…

Cited by 0SourcecodeScholar
2025

AutoURDF: Unsupervised Robot Modeling from Point Cloud Frames Using Cluster Registration

CVPR 2025poster

Robot description models are essential for simulation and control, yet their creation often requires significant manual effort. To streamline this modeling process, we introduce AutoURDF, an unsupervised approach for constructing description files for unseen robots from point cloud frames. Our metho…

2025

Causally Modeling the Linguistic and Social Factors that Predict Email Response

NAACL 2025long

Email is a vital conduit for human communication across businesses, organizations, and broader societal contexts. In this study, we aim to model the intents, expectations, and responsiveness in email exchanges. To this end, we release SIZZLER, a new dataset containing 1800 emails annotated with nuan…

Cited by 0SourcePDFScholar
2025

FactBench: A Dynamic Benchmark for In-the-Wild Language Model Factuality Evaluation

ACL 2025long

The rapid adoption of language models (LMs) across diverse applications has raised concerns about their factuality, i.e., their consistency with real-world facts. We introduce VERIFY, an evidence-based evaluation pipeline that measures LMs’ factuality in real-world user interactions. VERIFY consider…

Cited by 0SourcePDFScholar
2025

MoD-SLAM: Monocular Dense Mapping for Unbounded 3D Scene Reconstruction

RA-L 2025

Monocular SLAM has received a lot of attention due to its simple RGB inputs and the lifting of complex sensor constraints. However, existing monocular SLAM systems lack accurate depth estimation, which limits the accuracy of tracking and mapping performance. To address this limitation, we propose Mo

Cited by 33SourceScholar
2025

Toward Global AI Inclusivity: A Large-Scale Multilingual Terminology Dataset (GIST)

ACL 2025finding

The field of machine translation has achieved significant advancements, yet domain-specific terminology translation, particularly in AI, remains challenging. This work introduces GIST, a large-scale multilingual AI terminology dataset containing 5K terms extracted from top AI conference papers spann…

2025

VeriFact: Enhancing Long-Form Factuality Evaluation with Refined Fact Extraction and Reference Facts

EMNLP 2025

Large language models (LLMs) excel at generating long-form responses, but evaluating their factuality remains challenging due to complex inter-sentence dependencies within the generated facts. Prior solutions predominantly follow a decompose-decontextualize-verify pipeline but often fail to capture

Cited by 0SourcePDFScholar
2024

You don’t need a personality test to know these models are unreliable: Assessing the Reliability of Large Language Models on Psychometric Instruments

NAACL 2024long

The versatility of Large Language Models (LLMs) on natural language understanding tasks has made them popular for research in social sciences. To properly understand the properties and innate personas of LLMs, researchers have performed studies that involve using prompts in the form of questions tha…