← Search

Rajarshi Roy

14 accepted papers

2026

About Time: Model-Free Reinforcement Learning with Timed Reward Machines

IJCAI 2026

Reward specification plays a central role in reinforcement learning (RL), guiding the agent’s behavior. To express non-Markovian rewards, formalisms such as reward machines have been introduced to capture dependencies on histories. However, traditional reward machines lack the ability to model preci

Cited by 0Scholar
2026

DETONATE – A Benchmark for Text-to-Image Alignment and Kernelized Direct Preference Optimization

AAAI 2026technical

Alignment is crucial for text-to-image (T2I) models to ensure that the generated images faithfully capture user intent while maintaining safety and fairness. Direct Preference Optimization (DPO) has emerged as a key alignment technique for large language models (LLMs), and its influence is now exten

Cited by 0SourcePDFScholar
2025

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization

ACL 2025finding

The rapid advancement of large language models (LLMs) has revolutionized numerous applications, but presents significant challenges in aligning these models with diverse human values, ethical standards, and specific user preferences. Direct Preference Optimization (DPO) has become a cornerstone for…

2025

Learning Probabilistic Temporal Logic Specifications for Stochastic Systems

IJCAI 2025

There has been substantial progress in the inference of formal behavioural specifications from sample trajectories, for example using Linear Temporal Logic (LTL). However, these techniques cannot handle specifications that correctly characterise systems with stochastic behaviour, which occur commonl

2025

NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models

ICLR 2025spotlight

Decoder-only large language model (LLM)-based embedding models are beginning to outperform BERT or T5-based embedding models in general-purpose text embedding tasks, including dense vector-based retrieval. In this work, we introduce the NV-Embed model, incorporating architectural designs, training p…

Cited by 158SourcePDFScholar
2025

Nemotron-CORTEXA: Enhancing LLM Agents for Software Engineering Tasks via Improved Localization and Solution Diversity

ICML 2025poster

Large Language Models (LLMs) have demonstrated significant potential in code generation by following natural language instructions. Unfortunately, crucial real-world software engineering tasks, such as debugging or repository-level feature implementation, involve processing extensive contexts beyon…

Cited by 0SourcePDFScholar
2024

ChatQA: Surpassing GPT-4 on Conversational QA and RAG

NeurIPS 2024poster

In this work, we introduce ChatQA, a suite of models that outperform GPT-4 on retrieval-augmented generation (RAG) and conversational question answering (QA). To enhance generation, we propose a two-stage instruction tuning method that significantly boosts the performance of RAG. For effective ret…

Cited by 35SourcePDFScholar
2023

Learning Interpretable Temporal Properties from Positive Examples Only

AAAI 2023technical

We consider the problem of explaining the temporal behavior of black-box systems using human-interpretable models. Following recent research trends, we rely on the fundamental yet interpretable models of deterministic finite automata (DFAs) and linear temporal logic (LTL_f) formulas. In contrast to…

2020

Learning Interpretable Models in the Property Specification Language

IJCAI 2020poster

We address the problem of learning human-interpretable descriptions of a complex system from a finite set of positive and negative examples of its behavior. In contrast to most of the recent work in this area, which focuses on descriptions expressed in Linear Temporal Logic (LTL), we develop a learn…

2016

Complementary model update: A method for simultaneous registration and stiffness mapping in flexible environments

ICRA 2016

Registering a surgical tool to an a priori model of the environment is an important first step in computer-aided surgery. In this paper we present an approach for simultaneous registration and stiffness mapping using blind exploration of flexible environments. During contact-based exploration of fle

Cited by 32SourceScholar
2016

Concurrent nonparametric estimation of organ geometry and tissue stiffness using continuous adaptive palpation

ICRA 2016

Surgeons often manually palpate tissue or organs in order to find tumors or other anatomical structures. Information about organ geometry and tissue stiffness gained from palpation can also be extremely useful in robotic surgery for diagnosis, surgical guidance, and registration to other preoperativ

Cited by 33SourceScholar
2016

Using Bayesian optimization to guide probing of a flexible environment for simultaneous registration and stiffness mapping

ICRA 2016

One of the goals of computer-aided surgery is to register intraoperative data to preoperative model of the anatomy, and hence add complementary information that can facilitate the task of surgical navigation. In this context, mechanical palpation can reveal critical anatomical features such as arter

Cited by 46SourceScholar