← Search

Mark Hamilton

9 accepted papers

2026

MathNet: A Global Multimodal Benchmark for Mathematical Reasoning and Retrieval

ICLR 2026poster

Mathematical problem solving remains a challenging test of reasoning for large language and multimodal models, yet existing benchmarks are limited in size, language coverage, and task diversity. We introduce *MathNet*, a large-scale, high-quality, multilingual, and multimodal dataset of Olympiad-lev…

Cited by 0SourcecodeScholar
2025

I-Con: A Unifying Framework for Representation Learning

ICLR 2025poster

As the field of representation learning grows, there has been a proliferation of different loss functions to solve different classes of problems. We introduce a single information-theoretic equation that generalizes a large collection of mod- ern loss functions in machine learning. In particular, we…

Cited by 0SourcePDFScholar
2024

COCO-Periph: Bridging the Gap Between Human and Machine Perception in the Periphery

ICLR 2024poster

Evaluating deep neural networks (DNNs) as models of human perception has given rich insights into both human visual processing and representational properties of DNNs. We extend this work by analyzing how well DNNs perform compared to humans when constrained by peripheral vision -- which limits huma…

Cited by 3SourcePDFScholar
2024

FeatUp: A Model-Agnostic Framework for Features at Any Resolution

ICLR 2024poster

Deep features are a cornerstone of computer vision research, capturing image semantics and enabling the community to solve downstream tasks even in the zero- or few-shot regime. However, these features often lack the spatial resolution to directly perform dense prediction tasks like segmentation and…

2024

Separating the "Chirp" from the "Chat": Self-supervised Visual Grounding of Sound and Language

CVPR 2024poster

We present DenseAV a novel dual encoder grounding architecture that learns high-resolution semantically meaningful and audio-visual aligned features solely through watching videos. We show that DenseAV can discover the "meaning" of words and the "location" of sounds without explicit localization sup…

2023

Exploring perceptual straightness in learned visual representations

ICLR 2023poster

Humans have been shown to use a ''straightened'' encoding to represent the natural visual world as it evolves in time (Henaff et al. 2019). In the context of discrete video sequences, ''straightened'' means that changes between frames follow a more linear path in representation space at progressivel…

Cited by 5SourcePDFScholar
2022

Axiomatic Explanations for Visual Search, Retrieval, and Similarity Learning

ICLR 2022poster

Visual search, recommendation, and contrastive similarity learning power technologies that impact billions of users worldwide. Modern model architectures can be complex and difficult to interpret, and there are several competing techniques one can use to explain a search engine's behavior. We show t…

Cited by 9SourcePDFScholar
2022

Unsupervised Semantic Segmentation by Distilling Feature Correspondences

ICLR 2022poster

Unsupervised semantic segmentation aims to discover and localize semantically meaningful categories within image corpora without any form of annotation. To solve this task, algorithms must produce features for every pixel that are both semantically meaningful and compact enough to form distinct clus…