← Search

Patrick Emami

5 accepted papers

2026

Mitigating the Modality Gap in Vision–Language Models with Fractal Spectral Geometry

ICML 2026poster

Vision–language models such as CLIP embed images and text into a shared space, but still suffer from a modality gap, where image and text features cluster separately and nearest neighbors are dominated by same-modality rather than true cross-modal matches. Existing works alleviate the modality gap b…

Cited by 0SourceScholar
2025

SysCaps: Language Interfaces for Simulation Surrogates of Complex Systems

ICLR 2025poster

Surrogate models are used to predict the behavior of complex energy systems that are too expensive to simulate with traditional numerical methods. Our work introduces the use of language descriptions, which we call "system captions" or SysCaps, to interface with such surrogates. We argue that inte…

Cited by 7SourcePDFScholar
2023

BuildingsBench: A Large-Scale Dataset of 900K Buildings and Benchmark for Short-Term Load Forecasting

NeurIPS 2023poster

Short-term forecasting of residential and commercial building energy consumption is widely used in power systems and continues to grow in importance. Data-driven short-term load forecasting (STLF), although promising, has suffered from a lack of open, large-scale datasets with high building diversit…

2022

Self-Supervised Robust Scene Flow Estimation via the Alignment of Probability Density Functions

AAAI 2022technical

In this paper, we present a new self-supervised scene flow estimation approach for a pair of consecutive point clouds. The key idea of our approach is to represent discrete point clouds as continuous probability density functions using Gaussian mixture models. Scene flow estimation is therefore conv…

Cited by 11SourcePDFScholar
2021

Efficient Iterative Amortized Inference for Learning Symmetric and Disentangled Multi-Object Representations

ICML 2021spotlight

Unsupervised multi-object representation learning depends on inductive biases to guide the discovery of object-centric representations that generalize. However, we observe that methods for learning these representations are either impractical due to long training times and large memory consumption o…