← Search

Akshay Nambi

8 accepted papers

2026

Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Use

ICML 2026poster

Agentic language models operate in a fundamentally different safety regime than chat models: they must plan, call tools, and execute long-horizon actions where a single misstep, such as accessing files or entering credentials, can cause irreversible harm. Existing alignment methods, largely optimize…

Cited by 0SourceScholar
2025

Bridging the Language Gap: Dynamic Learning Strategies for Improving Multilingual Performance in LLMs

COLING 2025main

Large language models (LLMs) have revolutionized various domains but still struggle with non-Latin scripts and low-resource languages. This paper addresses the critical challenge of improving multilingual performance without extensive fine-tuning. We introduce a novel dynamic learning approach that…

Cited by 0SourcePDFScholar
2025

Exposing the Achilles’ Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical Reasoning

ACL 2025long

Large Language Models (LLMs) have significantly impacted the field of Math Word Problems (MWPs), transforming how these problems are approached and solved, particularly in educational contexts. However, existing evaluations often focus on final accuracy, neglecting the critical aspect of reasoning c…

Cited by 0SourcePDFScholar
2025

Multimodal Needle in a Haystack: Benchmarking Long-Context Capability of Multimodal Large Language Models

NAACL 2025long

Multimodal Large Language Models (MLLMs) have shown significant promise in various applications, leading to broad interest from researchers and practitioners alike. However, a comprehensive evaluation of their long-context capabilities remains underexplored. To address these gaps, we introduce the M…

2025

PromptWizard: Optimizing Prompts via Task-Aware, Feedback-Driven Self-Evolution

ACL 2025finding

Large language models (LLMs) have transformed AI across diverse domains, with prompting being central to their success in guiding model outputs. However, manual prompt engineering is both labor-intensive and domain-specific, necessitating the need for automated solutions. We introduce PromptWizard,…

Cited by 0SourcePDFScholar
2024

TorchSpatial: A Location Encoding Framework and Benchmark for Spatial Representation Learning

NeurIPS 2024poster

Spatial representation learning (SRL) aims at learning general-purpose neural network representations from various types of spatial data (e.g., points, polylines, polygons, networks, images, etc.) in their native formats. Learning good spatial representations is a fundamental problem for various dow…

2023

Chanakya: Learning Runtime Decisions for Adaptive Real-Time Perception

NeurIPS 2023poster

Real-time perception requires planned resource utilization. Computational planning in real-time perception is governed by two considerations -- accuracy and latency. There exist run-time decisions (e.g. choice of input resolution) that induce tradeoffs affecting performance on a given hardware, aris…

2023

MEGA: Multilingual Evaluation of Generative AI

EMNLP 2023long main

Generative AI models have shown impressive performance on many Natural Language Processing tasks such as language understanding, reasoning, and language generation. An important question being asked by the AI community today is about the capabilities and limits of these models, and it is clear that…

Cited by 0SourceScholar