← Search

Sho Sonoda

11 accepted papers

2026

Why Agentic Theorem Prover Works: A Statistical Provability Theory of Mathematical Reasoning Models

ICML 2026poster

Agentic theorem provers---pipelines that couple a mathematical reasoning model with library retrieval, decomposition/search, and a proof assistant verifier---have recently achieved striking empirical success, yet it remains unclear which components drive performance and why such systems work at all …

Cited by 0SourceScholar
2026

Why High-rank Neural Networks Generalize?: An Algebraic Framework with RKHSs

ICLR 2026poster

We derive a new Rademacher complexity bound for deep neural networks using Koopman operators, group representations, and reproducing kernel Hilbert spaces (RKHSs). The proposed bound describes why the models with high-rank weight matrices generalize well. Although there are existing bounds that atte…

Cited by 0SourceScholar
2025

Deep Ridgelet Transform and Unified Universality Theorem for Deep and Shallow Joint-Group-Equivariant Machines

ICML 2025poster

We present a constructive universal approximation theorem for learning machines equipped with joint-group-equivariant feature maps, called the joint-equivariant machines, based on the group representation theory. ``Constructive'' here indicates that the distribution of parameters is given in a close…

Cited by 0SourcePDFScholar
2024

Koopman-based generalization bound: New aspect for full-rank weights

ICLR 2024poster

We propose a new bound for generalization of neural networks using Koopman operators. Whereas most of existing works focus on low-rank weight matrices, we focus on full-rank weight matrices. Our bound is tighter than existing norm-based bounds when the condition numbers of weight matrices are small.…

Cited by 3SourcePDFScholar
2023

How Powerful are Shallow Neural Networks with Bandlimited Random Weights?

ICML 2023poster

We investigate the expressive power of depth-2 bandlimited random neural networks. A random net is a neural network where the hidden layer parameters are frozen with random assignment, and only the output layer parameters are trained by loss minimization. Using random weights for a hidden layer is a…

Cited by 10SourcePDFScholar
2023

Quantum Ridgelet Transform: Winning Lottery Ticket of Neural Networks with Quantum Computation

ICML 2023poster

A significant challenge in the field of quantum machine learning (QML) is to establish applications of quantum computation to accelerate common tasks in machine learning such as those for neural networks. Ridgelet transform has been a fundamental mathematical tool in the theoretical studies of neura…

Cited by 6SourcePDFScholar
2022

Fully-Connected Network on Noncompact Symmetric Space and Ridgelet Transform based on Helgason-Fourier Analysis

ICML 2022spotlight

Neural network on Riemannian symmetric space such as hyperbolic space and the manifold of symmetric positive definite (SPD) matrices is an emerging subject of research in geometric deep learning. Based on the well-established framework of the Helgason-Fourier transform on the noncompact symmetric sp…

Cited by 19SourcePDFScholar
2022

Universality of Group Convolutional Neural Networks Based on Ridgelet Analysis on Groups

NeurIPS 2022accept

We show the universality of depth-2 group convolutional neural networks (GCNNs) in a unified and constructive manner based on the ridgelet theory. Despite widespread use in applications, the approximation property of (G)CNNs has not been well investigated. The universality of (G)CNNs has been shown…

Cited by 11SourcePDFScholar
2021

Differentiable Multiple Shooting Layers

NeurIPS 2021poster

We detail a novel class of implicit neural models. Leveraging time-parallel methods for differential equations, Multiple Shooting Layers (MSLs) seek solutions of initial value problems via parallelizable root-finding algorithms. MSLs broadly serve as drop-in replacements for neural ordinary differe…

Cited by 23SourcePDFScholar
2021

Ridge Regression with Over-parametrized Two-Layer Networks Converge to Ridgelet Spectrum

AISTATS 2021poster

Characterization of local minima draws much attention in theoretical studies of deep learning. In this study, we investigate the distribution of parameters in an over-parametrized finite neural network trained by ridge regularized empirical square risk minimization (RERM). We develop a new theory of…

Cited by 17SourcePDFScholar
2020

Learning with Optimized Random Features: Exponential Speedup by Quantum Machine Learning without Sparsity and Low-Rank Assumptions

NeurIPS 2020poster

Kernel methods augmented with random features give scalable algorithms for learning from big data. But it has been computationally hard to sample random features according to a probability distribution that is optimized for the data, so as to minimize the required number of features for achieving th…

Cited by 22SourcePDFScholar