← Search

Sameep Mehta

14 accepted papers

2026

Automated Creation and Enrichment Framework for Improved Invocation of Enterprise APIs as Tools

AAAI 2026technical

Recent advancements in Large Language Models (LLMs) has lead to the development of agents capable of complex reasoning and interaction with external tools. In enterprise contexts, the effective use of such tools that are often enabled by application programming interfaces (APIs) is hindered by poor

Cited by 0SourcePDFScholar
2026

DFAgent: From Natural Language Data Interactions to Reusable Agent-Ready Tools

AAAI 2026technical

We present DataFoundry Agent (DFAgent), a system that forges reusable, agent-ready tools from interactive data exploration, quality, and remediation tasks. Users engage with data through natural-language prompts for operations that include inspection, transformation, and visualization. These interac

Cited by 0SourcePDFScholar
2026

From Natural Language to Executable ETL Flows: The IBM DataStage Assistant

AAAI 2026technical

Modern ETL (Extract, Transform, Load) tools offer graphical, no-code interfaces for workflow creation but still require users to manually identify transformation functions and configure their properties, which is time-consuming and demands prior expertise. We present the research and engineering fou

Cited by 0SourcePDFScholar
2026

ToolSmith: A Multi-Agent Framework for Enterprise Tool Creation

AAAI 2026technical

Although LLMs can generate tools for generic domains and tasks, they struggle with enterprise-related domains that involve proprietary APIs and data schemas. We present ToolSmith, a framework for autonomously generating and validating agent-compatible tools. Given an API specification and a Tool Spe

Cited by 0SourcePDFScholar
2025

CodeGenWrangler: Data Wrangling task automation using Code-Generating Models

NAACL 2025industry

Assuring the data quality of tabular datasets is essential for the efficiency of the diverse tabular downstream tasks (like summarization and fact-checking). Data-wrangling tasks effectively address the challenges associated with structured data processing to improve the quality of tabular data. Tra…

Cited by 0SourcePDFScholar
2025

Question-guided Insights Generation for Automated Exploratory Data Analysis

AAAI 2025technical

Exploratory Data Analysis (EDA) derives meaningful insights from extensive and complex datasets. This process typically involves a series of analytical operations to identify the patterns within the data. However, the effectiveness of EDA is often limited by the user's domain knowledge and proficien…

Cited by 0SourcePDFScholar
2025

Schema and Natural Language Aware In-Context Learning for Improved GraphQL Query Generation

NAACL 2025industry

GraphQL offers a flexible alternative to REST APIs, allowing precise data retrieval across multiple sources in a single query. However, generating complex GraphQL queries remains a significant challenge. Large Language Models (LLMs), while powerful, often produce suboptimal queries due to limited ex…

Cited by 0SourcePDFScholar
2024

GraphQL Query Generation: A Large Training and Benchmarking Dataset

EMNLP 2024industry

GraphQL is a powerful query language for APIs that allows clients to fetch precise data efficiently and flexibly, querying multiple resources with a single request. However, crafting complex GraphQL query operations can be challenging. Large Language Models (LLMs) offer an alternative by generating…

2024

LLM-powered GraphQL Generator for Data Retrieval

IJCAI 2024poster

GraphQL offers an efficient, powerful, and flexible alternative to REST APIs. However, application developers writing GraphQL clients need both technical and domain-specific expertise to reap its benefits, and avoid over-fetching or under-fetching data. Automated GraphQL generation has so far proven…

2024

LLMGuard: Guarding against Unsafe LLM Behavior

AAAI 2024technical

Although the rise of Large Language Models (LLMs) in enterprise settings brings new opportunities and capabilities, it also brings challenges, such as the risk of generating inappropriate, biased, or misleading content that violates regulations and can have legal concerns. To alleviate this, we pres…

Cited by 11SourcePDFScholar
2024

Sequential API Function Calling Using GraphQL Schema

EMNLP 2024main

Function calling using Large Language Models (LLMs) is an active research area that aims to empower LLMs with the ability to execute APIs to perform real-world tasks. However, sequential function calling using LLMs with interdependence between functions is still under-explored. To this end, we intro…

Cited by 1SourcePDFScholar
2023

CFL: Causally Fair Language Models Through Token-level Attribute Controlled Generation

ACL 2023findings

We propose a method to control the attributes of Language Models (LMs) for the text generation task using Causal Average Treatment Effect (ATE) scores and counterfactual augmentation. We explore this method, in the context of LM detoxification, and propose the Causally Fair Language (CFL) architectu…

Cited by 5SourcePDFScholar
2019

Learning Convolutional Neural Networks with Deep Part Embeddings

ICASSP 2019accepted

We propose a novel concept of Deep Part Embeddings (DPEs), which can be used to learn new Convolutional Neural Networks (CNNs) for different classes. We define DPE as a neuron of a trained CNN along with its network of filter activations that is interpretable as a part of a class that the neuron con…

Cited by 0SourceScholar
2019

Radial Loss for Learning Fine-grained Video Similarity Metric

ICASSP 2019accepted

In this paper, we propose the Radial Loss which utilizes category and sub-category labels to learn an order-preserving fine-grained video similarity metric. We propose an end-to-end quadlet-based Convolutional Neural Network (CNN) combined with Long Short-term Memory (LSTM) Unit to model video simil…

Cited by 0SourceScholar