← Search

Xu Xu

4 accepted papers

2026

Local Success Does Not Compose: Benchmarking Large Language Models for Compositional Formal Verification

ICLR 2026poster

Despite rapid advances in code generation, current Large Language Models (LLMs) still lack an essential capability for reliable and verifiable code generation: compositional reasoning across multi-function programs. To explore this potential and important gap, we introduce DafnyCOMP, a benchmark des…

Cited by 0SourceScholar
2026

VeriEquivBench: An Equivalence Score for Ground-Truth-Free Evaluation of Formally Verifiable Code

ICLR 2026poster

Formal verification is the next frontier for ensuring the correctness of code generated by Large Language Models (LLMs). While methods that co-generate code and formal specifications in formal languages, like Dafny, can, in principle, prove alignment with user intent, progress is bottlenecked by sp…

Cited by 0SourcecodeScholar
2024

Breaking Language Barriers: Cross-Lingual Continual Pre-Training at Scale

EMNLP 2024main

In recent years, Large Language Models (LLMs) have made significant strides towards Artificial General Intelligence. However, training these models from scratch requires substantial computational resources and vast amounts of text data. In this paper, we explores an alternative approach to construct…

Cited by 4SourcePDFScholar
2024

Optimal Measurement Poses Using LSSA for Robot Kinematics-Flexibility Calibration

RA-L 2024

The absolute positioning accuracy of robots is a primary factor limiting their applications. In order to enhance the efficiency and precision of robot calibration, this study introduces a method that utilizes the Levy flight and Sparrow Search Algorithm (LSSA) to optimize the measurement pose of the

Cited by 10SourceScholar