← Search

Siddhartha Mishra

26 accepted papers

2026

Imposing Boundary Conditions on Neural Operators via Learned Function Extensions

ICML 2026poster

Neural operators have emerged as powerful surrogates for the solution of partial differential equations (PDEs), yet their ability to handle general, highly variable boundary conditions (BCs) remains limited. Existing approaches often fail when the solution operator exhibits strong sensitivity to bou…

Cited by 0SourceScholar
2026

Learning, Solving and Optimizing PDEs with TensorGalerkin: an efficient high-performance Galerkin assembly algorithm

ICML 2026poster

We present a unified algorithmic framework for the numerical solution, constrained optimization, and physics-informed learning of PDEs with a variational structure. Our framework is based on a Galerkin discretization of the underlying variational forms, and its high efficiency stems from a novel hig…

Cited by 0SourceScholar
2026

Neuro-Symbolic AI for Analytical Solutions of Differential Equations

ICML 2026poster

Analytical solutions to differential equations offer exact, interpretable insight but are rarely available because discovering them requires expert intuition or exhaustive search in combinatorial spaces. We introduce SIGS, a neuro-symbolic framework that automates this process. SIGS uses a formal gr…

Cited by 0SourceScholar
2026

Towards a Certificate of Trust: Task-Aware OOD Detection for Scientific AI

ICLR 2026poster

Data-driven models are increasingly adopted in critical scientific fields like weather forecasting and fluid dynamics. These methods can fail on out-of-distribution (OOD) data, but detecting such failures in regression tasks is an open challenge. We propose a new OOD detection method based on estima…

Cited by 0SourceScholar
2025

Geometry Aware Operator Transformer as an efficient and accurate neural surrogate for PDEs on arbitrary domains

NeurIPS 2025poster

The very challenging task of learning solution operators of PDEs on arbitrary domains accurately and efficiently is of vital importance to engineering and industrial simulations. Despite the existence of many operator learning algorithms to approximate such PDEs, we find that accurate models are not…

Cited by 0SourcecodeScholar
2025

HyPINO: Multi-Physics Neural Operators via HyperPINNs and the Method of Manufactured Solutions

NeurIPS 2025spotlight

We present HyPINO, a multi-physics neural operator designed for zero-shot generalization across a broad class of parametric PDEs without requiring task-specific fine-tuning. Our approach combines a Swin Transformer-based hypernetwork with mixed supervision: (i) labeled data from analytical solutions…

Cited by 0SourcecodeScholar
2025

RIGNO: A Graph-based Framework For Robust And Accurate Operator Learning For PDEs On Arbitrary Domains

NeurIPS 2025poster

Learning the solution operators of PDEs on arbitrary domains is challenging due to the diversity of possible domain shapes, in addition to the often intricate underlying physics. We propose an end-to-end graph neural network (GNN) based neural operator to learn PDE solution operators from data on po…

Cited by 0SourcecodeScholar
2024

An operator preconditioning perspective on training in physics-informed machine learning

ICLR 2024poster

In this paper, we investigate the behavior of gradient descent algorithms in physics-informed machine learning methods like PINNs, which minimize residuals connected to partial differential equations (PDEs). Our key result is that the difficulty in training these models is closely related to the con…

Cited by 29SourcePDFScholar
2024

Beyond Regular Grids: Fourier-Based Neural Operators on Arbitrary Domains

ICML 2024poster

The computational efficiency of many neural operators, widely used for learning solutions of PDEs, relies on the fast Fourier transform (FFT) for performing spectral computations. As the FFT is limited to equispaced (rectangular) grids, this limits the efficiency of such neural operators when applie…

Cited by 6SourcePDFScholar
2024

FUSE: Fast Unified Simulation and Estimation for PDEs

NeurIPS 2024poster

The joint prediction of continuous fields and statistical estimation of the underlying discrete parameters is a common problem for many physical systems, governed by PDEs. Hitherto, it has been separately addressed by employing operator learning surrogates for field prediction while using simulation…

Cited by 2SourcePDFScholar
2024

Poseidon: Efficient Foundation Models for PDEs

NeurIPS 2024poster

We introduce Poseidon, a foundation model for learning the solution operators of PDEs. It is based on a multiscale operator transformer, with time-conditioned layer norms that enable continuous-in-time evaluations. A novel training strategy leveraging the semi-group property of time-dependent PDEs t…

2024

SmallToLarge (S2L): Scalable Data Selection for Fine-tuning Large Language Models by Summarizing Training Trajectories of Small Models

NeurIPS 2024poster

Despite the effectiveness of data selection for pretraining and instruction fine-tuning large language models (LLMs), improving data efficiency in supervised fine-tuning (SFT) for specialized domains poses significant challenges due to the complexity of fine-tuning data. To bridge this gap, we intro…

2023

Convolutional Neural Operators for robust and accurate learning of PDEs

NeurIPS 2023poster

Although very successfully used in conventional machine learning, convolution based neural network architectures -- believed to be inconsistent in function space -- have been largely ignored in the context of learning solution operators of PDEs. Here, we present novel adaptations for convolutional n…

2023

Gradient Gating for Deep Multi-Rate Learning on Graphs

ICLR 2023poster

We present Gradient Gating (G$^2$), a novel framework for improving the performance of Graph Neural Networks (GNNs). Our framework is based on gating the output of GNN layers with a mechanism for multi-rate flow of message passing information across nodes of the underlying graph. Local gradients are…

2023

Neural Inverse Operators for Solving PDE Inverse Problems

ICML 2023poster

A large class of inverse problems for PDEs are only well-defined as mappings from operators to functions. Existing operator learning frameworks map functions to functions and need to be modified to learn inverse maps from data. We propose a novel architecture termed Neural Inverse Operators (NIOs) t…

Cited by 48SourcePDFScholar
2023

Nonlinear Reconstruction for Operator Learning of PDEs with Discontinuities

ICLR 2023top-25%

Discontinuous solutions arise in a large class of hyperbolic and advection-dominated PDEs. This paper investigates, both theoretically and empirically, the operator learning of PDEs with discontinuous solutions. We rigorously prove, in terms of lower approximation bounds, that methods which entail a…

Cited by 37SourcePDFScholar
2023

Representation Equivalent Neural Operators: a Framework for Alias-free Operator Learning

NeurIPS 2023poster

Recently, operator learning, or learning mappings between infinite-dimensional function spaces, has garnered significant attention, notably in relation to learning partial differential equations from data. Conceptually clear when outlined on paper, neural operators necessitate discretization in the…

Cited by 52SourcePDFScholar
2022

An Evaluative Measure of Clustering Methods Incorporating Hyperparameter Sensitivity

AAAI 2022technical

Clustering algorithms are often evaluated using metrics which compare with ground-truth cluster assignments, such as Rand index and NMI. Algorithm performance may vary widely for different hyperparameters, however, and thus model selection based on optimal performance for these metrics is discordant…

2022

Generic bounds on the approximation error for physics-informed (and) operator learning

NeurIPS 2022accept

We propose a very general framework for deriving rigorous bounds on the approximation error for physics-informed neural networks (PINNs) and operator learning architectures such as DeepONets and FNOs as well as for physics-informed operator learning. These bounds guarantee that PINNs and (physics-in…

Cited by 80SourcePDFScholar
2022

Graph-Coupled Oscillator Networks

ICML 2022spotlight

We propose Graph-Coupled Oscillator Networks (GraphCON), a novel framework for deep learning on graphs. It is based on discretizations of a second-order system of ordinary differential equations (ODEs), which model a network of nonlinear controlled and damped oscillators, coupled via the adjacency s…

2022

Long Expressive Memory for Sequence Modeling

ICLR 2022spotlight

We propose a novel method called Long Expressive Memory (LEM) for learning long-term sequential dependencies. LEM is gradient-based, it can efficiently process sequential tasks with very long-term dependencies, and it is sufficiently expressive to be able to learn complicated input-output maps. To d…

2022

Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks

EMNLP 2022main

How well can NLP models generalize to a variety of unseen tasks when provided with task instructions? To address this question, we first introduce Super-NaturalInstructions, a benchmark of 1,616 diverse NLP tasks and their expert-written instructions. Our collection covers 76 distinct task types, in…

2022

Word2Box: Capturing Set-Theoretic Semantics of Words using Box Embeddings

ACL 2022long

Learning representations of words in a continuous space is perhaps the most fundamental task in NLP, however words interact in ways much richer than vector dot product similarity can provide. Many relationships between words can be expressed set-theoretically, for example, adjective-noun compounds (…

2021

Coupled Oscillatory Recurrent Neural Network (coRNN): An accurate and (gradient) stable architecture for learning long time dependencies

ICLR 2021oral

Circuits of biological neurons, such as in the functional parts of the brain can be modeled as networks of coupled oscillators. Inspired by the ability of these systems to express a rich set of outputs while keeping (gradients of) state variables bounded, we propose a novel architecture for recurren…

2021

UnICORNN: A recurrent model for learning very long time dependencies

ICML 2021spotlight

The design of recurrent neural networks (RNNs) to accurately process sequential inputs with long-time dependencies is very challenging on account of the exploding and vanishing gradient problem. To overcome this, we propose a novel RNN architecture which is based on a structure preserving discretiza…