← Search

Jonathan Francis

22 accepted papers

2026

DYMO-Hair: Generalizable Volumetric Dynamics Modeling for Robot Hair Manipulation

ICRA 2026poster

Hair care is an essential daily activity, yet it remains inaccessible to individuals with limited mobility and challenging for autonomous robot systems due to the fine-grained physical structure and complex dynamics of hair. In this work, we present DYMO-Hair, a model-based robot hair care system. W…

2026

GRAPPA: Generalizing and Adapting Robot Policies via Online Agentic Guidance

RA-L 2026

Robot learning approaches such as behavior cloning and reinforcement learning have shown great promise in synthesizing robot skills from human demonstrations in specific environments. However, these approaches often struggle to generalize to unseen real-world settings because they rely on task-speci

Cited by 0SourceScholar
2026

LightTact: A Visual-Tactile Fingertip Sensor for Deformation-Independent Contact Sensing

RSS 2026poster

Contact often occurs without macroscopic surface deformation, such as during interaction with liquids, semi-liquids, or ultra-soft materials. However, most existing tactile sensors rely on deformation to infer contact, making such light-contact interactions difficult to perceive robustly. To address…

Cited by 0SourceScholar
2026

RIO: Flexible Real-time Robot I/O for Cross-Embodiment Robot Learning

RSS 2026poster

Despite recent efforts to collect multi-task or multiembodiment datasets, to design efficient recipes for training Vision-Language-Action models (VLAs), and to showcase these models on selected robot platforms, generalist robot capabilities and cross-embodiment transfer remain largely elusive ideals…

Cited by 0SourceScholar
2026

SEAL: Towards Safe Autonomous Driving Via Skill-Enabled Adversary Learning for Closed-Loop Scenario Generation

ICRA 2026poster

Verification and validation of autonomous driving (AD) systems and components is of increasing importance, as such technology increases in real-world prevalence. Safety-critical scenario generation is a key approach to robustify AD policies through closed-loop training. However, existing approaches …

2026

STRIVE: Structured Representation Integrating VLM Reasoning for Efficient Object Navigation

ICRA 2026poster

Vision-Language Models (VLMs) have been increasingly integrated into object navigation tasks for their rich prior knowledge and strong reasoning abilities. However, applying VLMs to navigation presents two key challenges: effectively parsing and structuring complex environment information and determ…

2025

CaDRE: Controllable and Diverse Generation of Safety-Critical Driving Scenarios Using Real-World Trajectories

ICRA 2025

Simulation is an indispensable tool in the development and testing of autonomous vehicles (AVs), offering an efficient and safe alternative to road testing. An outstanding challenge with simulation-based testing is the generation of safety-critical scenarios, which are essential to ensure that AVs c

Cited by 11SourceScholar
2025

GraphEQA: Using 3D Semantic Scene Graphs for Real-time Embodied Question Answering

CoRL 2025poster

In Embodied Question Answering (EQA), agents must explore and develop a semantic understanding of an unseen environment in order to answer a situated question with confidence. This remains a challenging problem in robotics, due to the difficulties in obtaining useful semantic representations, updati…

Cited by 0SourcecodeScholar
2025

Human2LocoMan: Learning Versatile Quadrupedal Manipulation with Human Pretraining

RSS 2025poster

Quadrupedal robots have demonstrated impressive locomotion capabilities in complex environments, but equipping them with autonomous versatile manipulation skills in a scalable way remains a significant challenge. In this work, we introduce a system that integrates data collection and imitation learn…

Cited by 0PDFcodeScholar
2025

KineSoft: Learning Proprioceptive Manipulation Policies with Soft Robot Hands

CoRL 2025oral

Underactuated soft robot hands offer inherent safety and adaptability advantages over rigid systems, but developing dexterous manipulation skills remains challenging. While imitation learning shows promise for complex manipulation tasks, traditional approaches struggle with soft systems due to demon…

Cited by 0SourceScholar
2025

MOSAIC: Generating Consistent, Privacy-Preserving Scenes from Multiple Depth Views in Multi-Room Environments

ICCV 2025poster

We introduce a diffusion-based approach for generating privacy-preserving digital twins of multi-room indoor environments from depth images only. Central to our approach is a novel Multi-view Overlapped Scene Alignment with Implicit Consistency (MOSAIC) model that explicitly considers cross-view dep…

Cited by 0SourcePDFScholar
2025

SEAL: Towards Safe Autonomous Driving via Skill-Enabled Adversary Learning for Closed-Loop Scenario Generation

RA-L 2025

Verification and validation of autonomous driving (AD) systems and components is of increasing importance, as such technology increases in real-world prevalence. Safety-critical scenario generation is a key approach to robustify AD policies through closed-loop training. However, existing approaches

Cited by 10SourceScholar
2025

STRAP: Robot Sub-Trajectory Retrieval for Augmented Policy Learning

ICLR 2025poster

Robot learning is witnessing a significant increase in the size, diversity, and complexity of pre-collected datasets, mirroring trends in domains such as natural language processing and computer vision. Many robot learning methods treat such datasets as multi-task expert data and learn a multi-task,…

Cited by 1SourcePDFScholar
2024

MOSAIC: Learning Unified Multi-Sensory Object Property Representations for Robot Learning via Interactive Perception

ICRA 2024poster

A holistic understanding of object properties across diverse sensory modalities (e.g., visual, audio, and haptic) is essential for tasks ranging from object categorization to complex manipulation. Drawing inspiration from cognitive science studies that emphasize the significance of multi-sensory int…

Cited by 2SourcecodeScholar
2023

Core Challenges in Embodied Vision-Language Planning (Extended Abstract)

IJCAI 2023poster

Recent advances in the areas of Multimodal Machine Learning and Artificial Intelligence (AI) have led to the development of challenging tasks at the intersection of Computer Vision, Natural Language Processing, and Robotics. Whereas many approaches and previous survey pursuits have characterised one…

Cited by 0SourcePDFScholar
2023

Transferring Implicit Knowledge of Non-Visual Object Properties Across Heterogeneous Robot Morphologies

ICRA 2023poster

Humans leverage multiple sensor modalities when interacting with objects and discovering their intrinsic properties. Using the visual modality alone is insufficient for deriving intuition behind object properties (e.g., which of two boxes is heavier), making it essential to consider non-visual modal…

Cited by 17SourcecodeScholar
2023

What Went Wrong? Closing the Sim-to-Real Gap via Differentiable Causal Discovery

CoRL 2023poster

Training control policies in simulation is more appealing than on real robots directly, as it allows for exploring diverse states in an efficient manner. Yet, robot simulators inevitably exhibit disparities from the real-world \rebut{dynamics}, yielding inaccuracies that manifest as the dynamical si…

Cited by 33SourceScholar
2022

Coalescing Global and Local Information for Procedural Text Understanding

COLING 2022main

Procedural text understanding is a challenging language reasoning task that requires models to track entity states across the development of a narrative. We identify three core aspects required for modeling this task, namely the local and global view of the inputs, as well as the global view of outp…

2021

Exploring Strategies for Generalizable Commonsense Reasoning with Pre-trained Models

EMNLP 2021main

Commonsense reasoning benchmarks have been largely solved by fine-tuning language models. The downside is that fine-tuning may cause models to overfit to task-specific data and thereby forget their knowledge gained during pre-training. Recent works only propose lightweight model updates as models ma…

2021

Knowledge-driven Data Construction for Zero-shot Evaluation in Commonsense Question Answering

AAAI 2021technical

Recent developments in pre-trained neural language modeling have led to leaps in accuracy on common-sense question-answering benchmarks. However, there is increasing concern that models overfit to specific tasks, without learning to utilize external knowledge or perform general semantic reasoning.…

2021

Learn-To-Race: A Multimodal Control Environment for Autonomous Racing

ICCV 2021poster

Existing research on autonomous driving primarily focuses on urban driving, which is insufficient for characterising the complex driving behaviour underlying high-speed racing. At the same time, existing racing simulation frameworks struggle in capturing realism, with respect to visual rendering, ve…

Cited by 40PDFcodeScholar
2020

Diverse and Admissible Trajectory Prediction through Multimodal Context Understanding

ECCV 2020poster

Multi-agent trajectory forecasting in autonomous driving requires an agent to accurately anticipate the behaviors of the surrounding vehicles and pedestrians, for safe and reliable decision-making. Due to partial observability in these dynamical scenes, directly obtaining the posterior distribution…