← Search

Azim Ospanov

6 accepted papers

2026

HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs

ICML 2026poster

Informal mathematics has been central to modern large language model (LLM) reasoning, offering flexibility and efficient construction of arguments. However, purely informal reasoning is prone to logical gaps and subtle errors that are difficult to detect and correct. In contrast, formal theorem prov…

Cited by 0SourceScholar
2025

APOLLO: Automated LLM and Lean Collaboration for Advanced Formal Reasoning

NeurIPS 2025poster

Formal reasoning and automated theorem proving constitute a challenging subfield of machine learning, in which machines are tasked with proving mathematical theorems using formal languages like Lean. A formal verification system can check whether a formal proof is correct or not almost instantaneous…

Cited by 0SourcecodeScholar
2025

Do Vendi Scores Converge with Finite Samples? Truncated Vendi Score for Finite-Sample Convergence Guarantees

UAI 2025

Evaluating the diversity of generative models without reference data poses methodological challenges. The reference-free Vendi and RKE scores address this by quantifying the diversity of generated data using matrix-based entropy measures. Among these two, the Vendi score is typically computed via th

2025

Scendi Score: Prompt-Aware Diversity Evaluation via Schur Complement of CLIP Embeddings

ICCV 2025accepted

The use of CLIP embeddings to assess the fidelity of samples produced by text-to-image generative models has been extensively explored in the literature. While the widely adopted CLIPScore, derived from the cosine similarity of text and image embeddings, effectively measures the alignment of a gener…

2025

miniF2F-Lean Revisited: Reviewing Limitations and Charting a Path Forward

NeurIPS 2025poster

We perform a thorough analysis of the formal and informal statements in the miniF2F benchmark from the perspective of an AI system that is tasked to participate in a math Olympiad consisting of the problems in miniF2F. In such setting, the model has to read and comprehend the problems in natural lan…

Cited by 0SourcecodeScholar
2024

Towards a Scalable Reference-Free Evaluation of Generative Models

NeurIPS 2024poster

While standard evaluation scores for generative models are mostly reference-based, a reference-dependent assessment of generative models could be generally difficult due to the unavailability of applicable reference datasets. Recently, the reference-free entropy scores, VENDI and RKE, have been prop…