← Search

Ankit Shah

20 accepted papers

2026

Inference Time Optimization with Confidence Dynamics

ICML 2026poster

Inference time optimization techniques, such as repeated sampling, have significantly advanced the reasoning capabilities of Large Language Models (LLMs). However, the critical role of model uncertainty remains largely underexplored in these optimization strategies. In this paper, we investigate the…

Cited by 0SourceScholar
2026

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

ICLR 2026poster

We introduce MCP-Bench, a benchmark for evaluating large language models (LLMs) on realistic, multi-step tasks that demand tool use, cross-tool coordination, precise parameter control, and planning/reasoning for solving tasks. Built on the Model Context Protocol (MCP), MCP-Bench connects LLMs to 28…

Cited by 0SourcecodeScholar
2026

ProRefine: Inference-Time Prompt Refinement with Textual Feedback (Student Abstract)

AAAI 2026technical

Agentic workflows, where multiple AI agents collaborate to accomplish complex tasks like reasoning or planning, play a substantial role in many cutting-edge commercial applications. These workflows depend critically on the prompts used to provide the roles models play in such workflows. Poorly desig

Cited by 0SourcePDFScholar
2025

Improving Data Efficiency via Curating LLM-Driven Rating Systems

ICLR 2025poster

Instruction tuning is critical for adapting large language models (LLMs) to downstream tasks, and recent studies have demonstrated that small amounts of human-curated data can outperform larger datasets, challenging traditional data scaling laws. While LLM-based data quality rating systems offer a c…

Cited by 3SourcePDFScholar
2025

LLM Unlearning via Loss Adjustment with Only Forget Data

ICLR 2025poster

Unlearning in Large Language Models (LLMs) is essential for ensuring ethical and responsible AI use, especially in addressing privacy leak, bias, safety, and evolving regulations. Existing approaches to LLM unlearning often rely on retain data or a reference LLM, yet they struggle to adequately bala…

Cited by 2SourcePDFScholar
2024

Imprecise Label Learning: A Unified Framework for Learning with Various Imprecise Label Configurations

NeurIPS 2024poster

Learning with reduced labeling standards, such as noisy label, partial label, and supplementary unlabeled data, which we generically refer to as imprecise label, is a commonplace challenge in machine learning tasks. Previous methods tend to propose specific designs for every emerging imprecise label…

2024

Lang2LTL-2: Grounding Spatiotemporal Navigation Commands Using Large Language and Vision-Language Models

IROS 2024poster

Grounding spatiotemporal navigation commands to structured task specifications enables autonomous robots to understand a broad range of natural language and solve long-horizon tasks with safety guarantees. Prior works mostly focus on grounding spatial or temporally extended language for robots. We p…

Cited by 6SourceScholar
2024

Plug in the Safety Chip: Enforcing Constraints for LLM-driven Robot Agents

ICRA 2024poster

Recent advancements in large language models (LLMs) have enabled a new research domain, LLM agents, for solving robotics and planning tasks by leveraging the world knowledge and general reasoning abilities of LLMs obtained during pretraining. However, while considerable effort has been made to teach…

Cited by 52SourceScholar
2024

Skill Transfer for Temporal Task Specification

ICRA 2024poster

Deploying robots in real-world environments, such as households and manufacturing lines, requires generalization across novel task specifications without violating safety constraints. Linear temporal logic (LTL) is a widely used task specification language with a compositional grammar that naturally…

Cited by 19SourceScholar
2024

Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks

ICLR 2024spotlight

Pre-training on large-scale datasets and then fine-tuning on downstream tasks have become a standard practice in deep learning. However, pre-training data often contain label noise that may adversely affect the generalization of the model. This paper aims to understand the nature of noise in pre-tra…

2023

An Approach to Ontological Learning from Weak Labels

ICASSP 2023accepted

Ontologies encompass a formal representation of knowledge through the definition of concepts or properties of a domain, and the relationships between those concepts. In this work, we seek to investigate whether using this ontological information will improve learning from weakly labeled data, which…

Cited by 0SourceScholar
2023

Grounding Complex Natural Language Commands for Temporal Tasks in Unseen Environments

CoRL 2023poster

Grounding navigational commands to linear temporal logic (LTL) leverages its unambiguous semantics for reasoning about long-horizon tasks and verifying the satisfaction of temporal constraints. Existing approaches require training data from the specific environment and landmarks that will be used in…

Cited by 46SourceScholar
2022

Temporal Logic Imitation: Learning Plan-Satisficing Motion Policies from Demonstrations

CoRL 2022oral

Learning from demonstration (LfD) has successfully solved tasks featuring a long time horizon. However, when the problem complexity also includes human-in-the-loop perturbations, state-of-the-art approaches do not guarantee the successful reproduction of a task. In this work, we identify the roots o…

Cited by 26SourceScholar
2021

Bayes-TrEx: a Bayesian Sampling Approach to Model Transparency by Example

AAAI 2021technical

Post-hoc explanation methods are gaining popularity for interpreting, understanding, and debugging neural networks. Most analyses using such methods explain decisions in response to inputs drawn from the test set. However, the test set may have few examples that trigger some model behaviors, such a…

2021

Provably Safe and Efficient Motion Planning with Uncertain Human Dynamics

RSS 2021poster

Ensuring human safety without unnecessarily impacting task efficiency during human-robot interactive manipulation tasks is a critical challenge. In this work; we formally define human physical safety as collision avoidance or safe impact in the event of a collision. We developed a motion planner tha…

Cited by 29SourcePDFScholar
2018

Bayesian Inference of Temporal Task Specifications from Demonstrations

NeurIPS 2018poster

When observing task demonstrations, human apprentices are able to identify whether a given task is executed correctly long before they gain expertise in actually performing that task. Prior research into learning from demonstrations (LfD) has failed to capture this notion of the acceptability of an…

Cited by 106SourcePDFScholar
2018

Content-Based Representations of Audio Using Siamese Neural Networks

ICASSP 2018accepted

In this paper, we focus on the problem of content-based retrieval for audio, which aims to retrieve all semantically similar audio recordings for a given audio clip query. This problem is similar to the problem of query by example of audio, which aims to retrieve media samples from a database, which…

Cited by 0SourceScholar
2018

Framework for Evaluation of Sound Event Detection in Web Videos

ICASSP 2018accepted

The largest source of sound events is web videos. Most videos lack sound event labels at segment level, however, a significant number of them do respond to text queries, from a match found using metadata by search engines. In this paper we explore the extent to which a search query can be used as th…

Cited by 0SourceScholar