← Search

Harry Coppock

5 accepted papers

2026

Quantifying Biases in LLM-as-a-Judge Evaluations

ICML 2026poster

The evaluation of large language models (LLMs) is increasingly performed by other LLMs, a setup commonly known as "LLM-as-a-judge", or autograders. While autograders offer a scalable alternative to human evaluation, they are not free from biases (e.g., favouring longer outputs or generations from th…

Cited by 0SourceScholar
2026

Quantifying Frontier LLM Capabilities for Container Sandbox Escape

ICML 2026oral

Large Language Models (LLMs) increasingly act as autonomous agents with tool use, ability to execute code, file I/O, and network access. These capabilities create novel security risks. To mitigate these risks, agents are often deployed and evaluated in isolated environments commonly referred to as s…

Cited by 0SourceScholar
2025

Establishing Best Practices in Building Rigorous Agentic Benchmarks

NeurIPS 2025poster

Benchmarks are essential for quantitatively tracking progress in AI. As AI agents become increasingly capable, researchers and practitioners have introduced agentic benchmarks to evaluate agents on complex, real-world tasks. These benchmarks typically measure agent capabilities by evaluating task ou…

Cited by 0SourceScholar
2024

Synthia's Melody: A Benchmark Framework for Unsupervised Domain Adaptation in Audio

ICASSP 2024accepted

Despite significant advancements in deep learning for vision and natural language, unsupervised domain adaptation in audio remains relatively unexplored. We, in part, attribute this to the lack of an appropriate benchmark dataset. To address this gap, we present Synthia’s melody, a novel audio data…

Cited by 0SourceScholar
2023

Audio Barlow Twins: Self-Supervised Audio Representation Learning

ICASSP 2023accepted

The Barlow Twins self-supervised learning objective requires neither negative samples or asymmetric learning updates, achieving results on a par with the current state-of-the-art within Computer Vision. As such, we present Audio Barlow Twins, a novel self-supervised audio representation learning app…

Cited by 0SourceScholar