← Search

Robert McCarthy

7 accepted papers

2026

Early Signs of Steganographic Capabilities in Frontier LLMs

ICLR 2026poster

Monitoring Large Language Model (LLM) outputs is crucial for mitigating risks from misuse and misalignment. However, LLMs could evade monitoring through steganography: Encoding hidden information within seemingly benign generations. In this paper, we evaluate the steganography capabilities in fronti…

Cited by 0SourcecodeScholar
2026

Reasoning Models Struggle to Control their Chains of Thought

ICML 2026poster

Instruction following in LLMs captures models' ability to change their visible behaviors as requested by users. Instead, we study models' ability to control their chain-of-thought (CoT). This capability -- CoT controllability -- is undesirable because it could allow models to suppress signs of misbe…

Cited by 0SourceScholar
2026

Split Personality Training: Revealing Latent Knowledge Through Alternate Personalities

ICML 2026poster

Detecting misalignment in large language models is challenging because models may learn to conceal misbehavior during training. Standard auditing techniques fall short: black-box methods often cannot distinguish misaligned outputs from benign ones, and mechanistic interpretability does not scale wit…

Cited by 0SourceScholar
2026

Towards Generalist Robot Learning from Internet Video: A Survey (Abstract Reprint)

AAAI 2026technical

Scaling deep learning to massive and diverse internet data has driven remarkable breakthroughs in domains such as video generation and natural language processing. Robot learning, however, has thus far failed to replicate this success and remains constrained by a scarcity of available data. Learning

Cited by 0SourcePDFScholar
2025

Large language models can learn and generalize steganographic chain-of-thought under process supervision

NeurIPS 2025poster

Chain-of-thought (CoT) reasoning not only enhances large language model performance but also provides critical insights into decision-making processes, marking it as a useful tool for monitoring model intent and planning. By proactively preventing models from acting on CoT indicating misaligned or h…

Cited by 0SourceScholar
2024

Identifying Expert Behavior in Offline Training Datasets Improves Behavioral Cloning of Robotic Manipulation Policies

RA-L 2024

This letter presents our solution for the Real Robot Challenge III <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> , aiming to address dexterous robotic manipulation tasks through learning from offline data. In this competition, participants wer

Cited by 13SourcecodeScholar
2023

Improving Behavioural Cloning with Positive Unlabeled Learning

CoRL 2023poster

Learning control policies offline from pre-recorded datasets is a promising avenue for solving challenging real-world problems. However, available datasets are typically of mixed quality, with a limited number of the trajectories that we would consider as positive examples; i.e., high-quality demons…

Cited by 8SourceScholar