← Search

Lalithkumar Seenivasan

4 accepted papers

2025

ConceptAgent: LLM-Driven Precondition Grounding and Tree Search for Robust Task Planning and Execution

ICRA 2025

Robotic planning and execution in open-world environments is a complex problem due to the vast state spaces and high variability of task embodiment. Recent advances in perception algorithms, combined with Large Language Models (LLMs) for planning, offer promising solutions to these challenges, as th

Cited by 7SourceScholar
2025

Online Reasoning Video Segmentation with Just-in-Time Digital Twins

ICCV 2025poster

Reasoning segmentation (RS) aims to identify and segment objects of interest based on implicit text queries. As such, RS is a catalyst for embodied AI agents, enabling them to interpret high-level commands without requiring explicit step-by-step guidance. However, current RS approaches rely heavily…

2023

Surgical-VQLA:Transformer with Gated Vision-Language Embedding for Visual Question Localized-Answering in Robotic Surgery

ICRA 2023poster

Despite the availability of computer-aided simulators and recorded videos of surgical procedures, junior residents still heavily rely on experts to answer their queries. However, expert surgeons are often overloaded with clinical and academic workloads and limit their time in answering. For this pur…

Cited by 34SourcecodeScholar
2022

Rethinking Feature Extraction: Gradient-Based Localized Feature Extraction for End-To-End Surgical Downstream Tasks

RA-L 2022

Several approaches have been introduced to understand surgical scenes through downstream tasks like captioning and surgical scene graph generation. However, most of them heavily rely on an independent object detector and region-based feature extractor. Encompassing computationally expensive detectio

Cited by 4SourceScholar