← Search

Daksh Dobhal

2 accepted papers

2025

Autonomous Evaluation of LLMs for Truth Maintenance and Reasoning Tasks

ICLR 2025poster

This paper presents AutoEval, a novel benchmark for scaling Large Language Model (LLM) assessment in formal tasks with clear notions of correctness, such as truth maintenance in translation and logical reasoning. AutoEval is the first benchmarking paradigm that offers several key advantages necessar…

Cited by 0SourcePDFScholar
2025

Using Explainable AI and Hierarchical Planning for Outreach with Robots

AAAI 2025technical

Understanding how robots plan and execute tasks is crucial in today's world, where they are becoming more prevalent in our daily lives. However, teaching non-experts, such as K-12 students, the complexities of robot planning can be challenging. This work presents an open-source platform, JEDAI.Ed, t…