← Search

Alvaro Arroyo

8 accepted papers

2026

Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin

ICLR 2026poster

Attention sinks and compression valleys have attracted significant attention as two puzzling phenomena in large language models, but have been studied in isolation. In this work, we present a surprising connection between attention sinks and compression valleys, tracing both to the formation of mass…

Cited by 0SourceScholar
2026

gLSTM: Mitigating Over-Squashing by Increasing Storage Capacity

ICLR 2026poster

Graph Neural Networks (GNNs) leverage the graph structure to transmit information between nodes, typically through the message-passing mechanism. While these models have found a wide variety of applications, they are known to suffer from over-squashing, where information from a large receptive field…

Cited by 0SourcecodeScholar
2025

On Vanishing Gradients, Over-Smoothing, and Over-Squashing in GNNs: Bridging Recurrent and Graph Learning

NeurIPS 2025poster

Graph Neural Networks (GNNs) are models that leverage the graph structure to transmit information between nodes, typically through the message-passing operation. While widely successful, this approach is well-known to suffer from representational collapse as the number of layers increases and insens…

Cited by 0SourceScholar
2025

Return of ChebNet: Understanding and Improving an Overlooked GNN on Long Range Tasks

NeurIPS 2025spotlight

ChebNet, one of the earliest spectral GNNs, has largely been overshadowed by Message Passing Neural Networks (MPNNs), which gained popularity for their simplicity and effectiveness in capturing local graph structure. Despite their success, MPNNs are limited in their ability to capture long-range dep…

Cited by 0SourceScholar
2024

Rough Transformers: Lightweight and Continuous Time Series Modelling through Signature Patching

NeurIPS 2024poster

Time-series data in real-world settings typically exhibit long-range dependencies and are observed at non-uniform intervals. In these settings, traditional sequence-based recurrent models struggle. To overcome this, researchers often replace recurrent models with Neural ODE-based architectures to ac…

2023

Neural Latent Geometry Search: Product Manifold Inference via Gromov-Hausdorff-Informed Bayesian Optimization

NeurIPS 2023poster

Recent research indicates that the performance of machine learning models can be improved by aligning the geometry of the latent space with the underlying data structure. Rather than relying solely on Euclidean space, researchers have proposed using hyperbolic and spherical spaces with constant curv…

Cited by 10SourcePDFScholar
2022

Dynamic Portfolio Cuts: A Spectral Approach to Graph-Theoretic Diversification

ICASSP 2022accepted

Stock market returns are typically analyzed using standard regression models yet they reside on irregular domains, a natural scenario for graph signal processing. This motivates us to consider a market graph as an intuitive way to represent the relationships between financial assets. Traditional met…

Cited by 0SourceScholar
2021

Nonstationary Portfolios: Diversification in the Spectral Domain

ICASSP 2021accepted

Classical portfolio optimization methods typically determine an optimal capital allocation through the implicit, yet critical, assumption of statistical time-invariance. Such models are inadequate for real-world markets as they employ standard time-averaging based estimators which suffer significant…

Cited by 0SourceScholar