← Search

Brian M Sadler

25 accepted papers

2026

Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel Approach

ICLR 2026poster

Hierarchical reinforcement learning (HRL) enables agents to solve complex, long-horizon tasks by decomposing them into manageable sub-tasks. However, HRL methods face two fundamental challenges: (i) non-stationarity caused by the evolving lower-level policy during training, which destabilizes higher…

Cited by 0SourceScholar
2025

Confidence-Controlled Exploration: Efficient Sparse-Reward Policy Learning for Robot Navigation

IROS 2025

Reinforcement learning (RL) is a promising approach for robotic navigation, allowing robots to learn through trial and error. However, real-world robotic tasks often suffer from sparse rewards, leading to inefficient exploration and suboptimal policies due to sample inefficiency of RL. In this work,

Cited by 4SourceScholar
2025

On the Vulnerability of LLM/VLM-Controlled Robotics

IROS 2025

In this work, we highlight vulnerabilities in robotic systems integrating large language models (LLMs) and vision-language models (VLMs) due to input modality sensitivities. While LLM/VLM-controlled robots show impressive performance across various tasks, their reliability under slight input variati

Cited by 13SourcecodeScholar
2024

Multi-Antenna ISAC Receiver with n-Tuple Blind Deconvolution

ICASSP 2024accepted

Recent developments in spectrum-sharing technologies include integrated sensing and communications (ISAC) systems to save resources, cost, and power. In this paper, we consider a co-existence topology with n-tuple radar and communications transmitters, wherein neither the transmitted signal nor the…

Cited by 0SourceScholar
2024

PIPER: Primitive-Informed Preference-based Hierarchical Reinforcement Learning via Hindsight Relabeling

ICML 2024poster

In this work, we introduce PIPER: Primitive-Informed Preference-based Hierarchical reinforcement learning via Hindsight Relabeling, a novel approach that leverages preference-based learning to learn a reward model, and subsequently uses this reward model to relabel higher-level replay buffers. Since…

2024

Sampling-based Safe Reinforcement Learning for Nonlinear Dynamical Systems

AISTATS 2024poster

We develop provably safe and convergent reinforcement learning (RL) algorithms for control of nonlinear dynamical systems, bridging the gap between the hard safety guarantees of control theory and the convergence guarantees of RL theory. Recent advances at the intersection of control and RL follow a…

2024

Towards Global Optimality for Practical Average Reward Reinforcement Learning without Mixing Time Oracles

ICML 2024poster

In the context of average-reward reinforcement learning, the requirement for oracle knowledge of the mixing time, a measure of the duration a Markov chain under a fixed policy needs to achieve its stationary distribution, poses a significant challenge for the global convergence of policy gradient me…

Cited by 2SourcePDFScholar
2023

Beyond Exponentially Fast Mixing in Average-Reward Reinforcement Learning via Multi-Level Monte Carlo Actor-Critic

ICML 2023poster

Many existing reinforcement learning (RL) methods employ stochastic gradient iteration on the back end, whose stability hinges upon a hypothesis that the data-generating process mixes exponentially fast with a rate parameter that appears in the step-size selection. Unfortunately, this assumption is…

Cited by 14SourcePDFScholar
2023

Identifying Coordination in a Cognitive Radar Network - A Multi-Objective Inverse Reinforcement Learning Approach

ICASSP 2023accepted

Consider a target being tracked by a cognitive radar network. If the target can intercept some radar network emissions, how can it detect coordination among the radars? By 'coordination' we mean that the radar emissions satisfy Pareto optimality with respect to multiobjective optimization over each…

Cited by 0SourceScholar
2023

LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language Models

ICCV 2023poster

This study focuses on using large language models (LLMs) as a planner for embodied agents that can follow natural language instructions to complete complex tasks in a visually-perceived environment. The high data cost and poor sample efficiency of existing methods hinders the development of versatil…

Cited by 577PDFScholar
2023

Unique Bispectrum Inversion for Signals with Finite Spectral/Temporal Support

ICASSP 2023accepted

Retrieving a signal from its triple correlation spectrum, also called bispectrum, arises in a wide range of signal processing problems. Conventional methods do not provide an accurate inversion of bispectrum to the underlying signal. In this paper, we present an approach that uniquely recovers signa…

Cited by 1SourceScholar
2022

Joint Radar-Communications Processing from A Dual-Blind Deconvolution Perspective

ICASSP 2022accepted

We consider a general spectral coexistence scenario, wherein the channels and transmit signals of both radar and communications systems are unknown at the receiver. In this dual-blind deconvolution (DBD) problem, a common receiver admits the multi-carrier wireless communications signal that is overl…

Cited by 0SourceScholar
2022

On the Hidden Biases of Policy Mirror Ascent in Continuous Action Spaces

ICML 2022spotlight

We focus on parameterized policy search for reinforcement learning over continuous action spaces. Typically, one assumes the score function associated with a policy is bounded, which {fails to hold even for Gaussian policies. } To properly address this issue, one must introduce an exploration tolera…

Cited by 20SourcePDFScholar
2022

One Step at a Time: Long-Horizon Vision-and-Language Navigation With Milestones

CVPR 2022poster

We study the problem of developing autonomous agents that can follow human instructions to infer and perform a sequence of actions to complete the underlying task. Significant progress has been made in recent years, especially for tasks with short horizons. However, when it comes to long-horizon tas…

Cited by 33PDFcodeScholar
2022

Robotic Parasitic Array Control for Increased RSS in Non-Line-of-Sight

RA-L 2022

Safety, security, and rescue missions typically occur in environments where unreliable wireless communication impedes cooperative human and robot missions. A low-VHF band antenna array formed through the coordination of ground robots has the potential to transmit and receive signals more reliably at

Cited by 2SourceScholar
2021

Banraw: Band-Limited Radar Waveform Design Via Phase Retrieval

ICASSP 2021accepted

This paper presents a uniqueness result which states that a band- limited signal can be recovered from at least 3B measurements where B is the bandwidth from the radar ambiguity function (AF). This function is a two-dimensional mapping of the propagation delay and Doppler frequency. This formal mode…

Cited by 9SourceScholar
2021

Performance Analysis of Spatial and Frequency Domain Index-Modulated Reconfigurable Intelligent Metasurfaces

ICASSP 2021accepted

Higher spectral and energy efficiencies are the envisioned defining characteristics of next-generation high data-rate sixth-generation (6G) wireless networks. One of the enabling technologies to meet these requirements is index modulation (IM), which transmits information through permutations of ind…

Cited by 0SourceScholar
2021

Physical-Layer Security via Distributed Beamforming in the Presence of Adversaries with Unknown Locations

ICASSP 2021accepted

We study the problem of securely communicating a sequence of information bits with a client in the presence of multiple adversaries at unknown locations in the environment. We assume that the client and the adversaries are located in the far-field region, and all possible directions for each adversa…

Cited by 0SourceScholar
2021

VGAI: End-to-End Learning of Vision-Based Decentralized Controllers for Robot Swarms

ICASSP 2021accepted

Decentralized coordination of a robot swarm requires addressing the tension between local perceptions and actions, and the accomplishment of a global objective. In this work, we propose to learn decentralized controllers based solely on raw visual inputs. For the first time, this integrates the lear…

Cited by 0SourceScholar
2020

Game Theoretic Formation Design for Probabilistic Barrier Coverage

IROS 2020poster

We study strategies to deploy defenders/sensors to detect intruders that approach a targeted region. This scenario is formulated as a barrier coverage, which aims to minimize the number of unseen paths. The problem becomes challenging when the number of defenders is insufficient for a full coverage,…

Cited by 7SourceScholar
2018

Unequal Error Protection Querying Policies for the Noisy 20 Questions Problem

ICASSP 2018accepted

We propose a non-adaptive unequal error protection (UEP) querying policy based on superposition coding for the noisy 20 questions problem. In this problem, a player wishes to successively refine an estimate of the value of a continuous random variable by posing binary queries and receiving noisy res…

Cited by 0SourceScholar