← Search

Robert D. Mullins

7 accepted papers

2026

SURGE: Surprise-Guided Token Reduction for Efficient Video Understanding with VLMs

ICLR 2026poster

Videos contain rich information but also high redundancy, as consecutive frames often share similar backgrounds and predictable motions. Current video-language models (VLMs) are unable to exploit this redundancy and therefore perform a significant amount of superfluous computation, processing thousa…

Cited by 0SourceScholar
2025

Hardware and Software Platform Inference

ICML 2025poster

It is now a common business practice to buy access to large language model (LLM) inference rather than self-host, because of significant upfront hardware infrastructure and energy costs. However, as a buyer, there is no mechanism to verify the authenticity of the advertised service including the ser…

Cited by 0SourcePDFScholar
2025

Inverse Constitutional AI: Compressing Preferences into Principles

ICLR 2025poster

Feedback data is widely used for fine-tuning and evaluating state-of-the-art AI models. Pairwise text preferences, where human or AI annotators select the “better” of two options, are particularly common. Such preferences are used to train (reward) models or to rank models with aggregate statistics.…

2025

PhySwin: An Efficient and Physically-Informed Foundation Model for Multispectral Earth Observation

NeurIPS 2025poster

Recent progress on Remote Sensing Foundation Models (RSFMs) aims toward universal representations for Earth observation imagery. However, current efforts often scale up in size significantly without addressing efficiency constraints critical for real-world applications (e.g., onboard processing, rap…

Cited by 0SourceScholar
2024

Beyond Slow Signs in High-fidelity Model Extraction

NeurIPS 2024poster

Deep neural networks, costly to train and rich in intellectual property value, are increasingly threatened by model extraction attacks that compromise their confiden- tiality. Previous attacks have succeeded in reverse-engineering model parameters up to a precision of float64 for models trained on r…

2023

Dynamic Stashing Quantization for Efficient Transformer Training

EMNLP 2023short findings

Large Language Models (LLMs) have demonstrated impressive performance on a range of Natural Language Processing (NLP) tasks. Unfortunately, the immense amount of computations and memory accesses required for LLM training makes them prohibitively expensive in terms of hardware cost, and thus challeng…

Cited by 0SourceScholar
2022

Rapid Model Architecture Adaption for Meta-Learning

NeurIPS 2022accept

Network Architecture Search (NAS) methods have recently gathered much attention. They design networks with better performance and use a much shorter search time compared to traditional manual tuning. Despite their efficiency in model deployments, most NAS algorithms target a single task on a fixed h…

Cited by 6SourcePDFScholar