← Search

Kenneth Marino

14 accepted papers

2025

BriefMe: A Legal NLP Benchmark for Assisting with Legal Briefs

ACL 2025finding

A core part of legal work that has been underexplored in Legal NLP is the writing and editing of legal briefs. This requires not only a thorough understanding of the law of a jurisdiction, from judgments to statutes, but also the ability to make new arguments to try to expand the law in a new direct…

Cited by 0SourcePDFScholar
2025

Efficient Exploration and Discriminative World Model Learning with an Object-Centric Abstraction

ICLR 2025poster

In the face of difficult exploration problems in reinforcement learning, we study whether giving an agent an object-centric mapping (describing a set of items and their attributes) allow for more efficient learning. We found this problem is best solved hierarchically by modelling items at a higher l…

Cited by 0SourcePDFScholar
2025

ReCogLab: a framework testing relational reasoning & cognitive hypotheses on LLMs

ICLR 2025poster

A fundamental part of human cognition is the ability to not only recall previous memories, but also reason across them to draw conclusions. In cognitive science and psychology, this is termed relational reasoning and a number of effects and biases have been observed in human cognition. Designing exp…

2024

VLM Agents Generate Their Own Memories: Distilling Experience into Embodied Programs of Thought

NeurIPS 2024spotlight

Large-scale generative language and vision-language models (LLMs and VLMs) excel in few-shot in-context learning for decision making and instruction following. However, they require high-quality exemplar demonstrations to be included in their context window. In this work, we ask: Can LLMs and VLMs g…

Cited by 5SourcePDFScholar
2023

Distilling Internet-Scale Vision-Language Models into Embodied Agents

ICML 2023poster

Instruction-following agents must ground language into their observation and action spaces. Learning to ground language is challenging, typically requiring domain-specific engineering or large quantities of human interaction data. To address this challenge, we propose using pretrained vision-languag…

Cited by 29SourcePDFScholar
2022

A-OKVQA: A Benchmark for Visual Question Answering Using World Knowledge

ECCV 2022poster

"The Visual Question Answering (VQA) task aspires to provide a meaningful testbed for the development of AI models that can jointly reason over visual and natural language inputs. Despite a proliferation of VQA datasets, this goal is hindered by a set of common limitations. These include a reliance…

2022

Learning to Navigate Wikipedia by Taking Random Walks

NeurIPS 2022accept

A fundamental ability of an intelligent web-based agent is seeking out and acquiring new information. Internet search engines reliably find the correct vicinity but the top results may be a few links away from the desired target. A complementary approach is navigation via hyperlinks, employing a pol…

Cited by 5SourcePDFScholar
2021

Ask Your Humans: Using Human Instructions to Improve Generalization in Reinforcement Learning

ICLR 2021poster

Complex, multi-task problems have proven to be difficult to solve efficiently in a sparse-reward reinforcement learning setting. In order to be sample efficient, multi-task learning requires reuse and sharing of low-level policies. To facilitate the automatic decomposition of hierarchical tasks, we…

2021

KRISP: Integrating Implicit and Symbolic Knowledge for Open-Domain Knowledge-Based VQA

CVPR 2021poster

One of the most challenging question types in VQA is when answering the question requires outside knowledge not present in the image. In this work we study open-domain knowledge, the setting when the knowledge required to answer a question is not given/annotated, neither at training nor test time. W…

Cited by 247PDFScholar
2020

Same Object, Different Grasps: Data and Semantic Knowledge for Task-Oriented Grasping

CoRL 2020

Despite the enormous progress and generalization in robotic grasping in recent years, existing methods have yet to scale and generalize task-oriented grasping to the same extent. This is largely due to the scale of the datasets both in terms of the number of objects and tasks studied. We address the

2019

Hierarchical RL Using an Ensemble of Proprioceptive Periodic Policies

ICLR 2019poster

In this paper we introduce a simple, robust approach to hierarchically training an agent in the setting of sparse reward tasks. The agent is split into a low-level and a high-level policy. The low-level policy only accesses internal, proprioceptive dimensions of the state observation. The low-level…

Cited by 20SourcePDFScholar
2019

OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge

CVPR 2019poster

Visual Question Answering (VQA) in its ideal form lets us study reasoning in the joint space of vision and language and serves as a proxy for the AI task of scene understanding. However, most VQA benchmarks to date are focused on questions such as simple counting, visual attributes, and object detec…

Cited by 1199PDFScholar