publications
2026
- Boundary Variance Inflation Causes Acquisition Bias in Gaussian ProcessesMaria Bånkestad, Sanna Jarl, and Jens SjölundarXiv preprint arXiv:2606.07561, 2026
Gaussian processes with stationary kernels on bounded domains exhibit inflated posterior variance near the boundary. Despite being a long-recognized artifact in geostatistics and a source of over-exploration in Bayesian optimization, the causes and effects of boundary-induced acquisition bias are underexplored. We trace the root cause to a simple geometric mechanism: the truncation of the kernel correlation neighborhood at the domain boundary creates an observation-independent distortion that worsens with dimensionality. We show how this distortion manifests across three acquisition classes: variance maximization concentrates selections at the corners, whereas negative integrated posterior variance and expected predictive information gain move selections inward to axis-aligned interior shells. These patterns arise without reference to any objective function, meaning that acquisition behavior can be dominated by kernel geometry rather than the desired task-specific uncertainty. To quantify this, we introduce a function-free selection-profile diagnostic for arbitrary acquisitions, kernels, and bounded-domain geometries.
@article{bankestad2026boundary, title = {Boundary Variance Inflation Causes Acquisition Bias in Gaussian Processes}, author = {B{\aa}nkestad, Maria and Jarl, Sanna and Sj{\"o}lund, Jens}, journal = {arXiv preprint arXiv:2606.07561}, year = {2026}, } -
A differentiable machine learning small-angle X-ray scattering analysis framework for structure elucidation of lipid nanoparticlesMaria Bånkestad, Sandra Barman, Magnus Röding, and 11 more authorsarXiv preprint arXiv:2606.05200, 2026Lipid nanoparticles (LNPs) are efficient delivery systems for negatively charged nucleic acids. Their multi-component architecture yields a core-shell structure. Small-angle X-ray scattering (SAXS) is an important characterization technique for LNPs, but recovering internal structure and size distribution from SAXS is an inverse problem with non-unique solutions. Realistic models are often too expensive for systematic exploration. We introduce a machine-learning-accelerated, differentiable framework for SAXS analysis of heterogeneous, polydisperse LNPs. The forward model combines a core-shell particle with a Gaussian random-field interior, a neural surrogate for the monodisperse SAXS map, and a differentiable layer integrating over particle-size distributions. The surrogate reduces prediction cost by four orders of magnitude, while differentiability enables large-scale multi-start fitting and ensemble identifiability analysis. Applied to synthetic and experimental MC3 LNP data, the framework shows that near-identical SAXS fits can arise from distinct parameter modes, with the experimental fits dominated by a trade-off between size-distribution and interior-structure parameters.
@article{bankestad2026saxs, title = {A differentiable machine learning small-angle X-ray scattering analysis framework for structure elucidation of lipid nanoparticles}, author = {B{\aa}nkestad, Maria and Barman, Sandra and R{\"o}ding, Magnus and Kaunisto, Erik and Meklesh, Viktoriia and Gallud, Audrey and Mendez, Marco and Yanez Arteta, Marianna and Norberg, Stefan and Terry, Ann and Chakraborty, Smita and Yu, Shun and R{\"o}nnols, Jerk and Pashami, Sepideh}, journal = {arXiv preprint arXiv:2606.05200}, year = {2026}, } -
Observation-dependent Bayesian active learning via input-warped Gaussian processesSanna Jarl, Maria Bånkestad, Jonathan J. S. Scragg, and 1 more authorarXiv preprint arXiv:2602.01898, 2026Bayesian active learning relies on the precise quantification of predictive uncertainty to explore unknown function landscapes. While Gaussian process surrogates are the standard for such tasks, an underappreciated fact is that their posterior variance depends on the observed outputs only through the hyperparameters, rendering exploration largely insensitive to the actual measurements. We propose to inject observation-dependent feedback by warping the input space with a learned, monotone reparameterization. This mechanism allows the design policy to expand or compress regions of the input space in response to observed variability, thereby shaping the behavior of variance-based acquisition functions. We demonstrate that while such warps can be trained via marginal likelihood, a novel self-supervised objective yields substantially better performance. Our approach improves sample efficiency across a range of active learning benchmarks, particularly in regimes where non-stationarity challenges traditional methods.
@article{jarl2026observation, title = {Observation-dependent Bayesian active learning via input-warped Gaussian processes}, author = {Jarl, Sanna and B{\aa}nkestad, Maria and Scragg, Jonathan J. S. and Sj{\"o}lund, Jens}, journal = {arXiv preprint arXiv:2602.01898}, year = {2026}, }
2025
- EurIPS-W
DiffWake: A General Differentiable Wind Farm Solver in JAXMaria Bånkestad, Leon Sütfeld, Aleksis Pirinen, and 1 more authorIn Workshop on Differentiable Systems and Scientific Machine Learning (EurIPS) , 2025We present DiffWake, a general differentiable wind farm solver implemented in JAX, including the first differentiable implementation of the cumulative-curl wake model. End-to-end differentiability enables fast gradient-based wind farm layout optimization and probabilistic calibration of turbulence intensity from data. On layout optimization, L-BFGS converges roughly fifty times faster than a FLORIS/SciPy baseline, and the calibration approach improves turbulence-intensity prediction on operational SCADA data.
@inproceedings{bankestad2025diffwake, title = {DiffWake: A General Differentiable Wind Farm Solver in JAX}, author = {B{\aa}nkestad, Maria and S{\"u}tfeld, Leon and Pirinen, Aleksis and Abedi, Hamidreza}, booktitle = {Workshop on Differentiable Systems and Scientific Machine Learning (EurIPS)}, year = {2025}, } - ThesisStructured models for scientific machine learning: From graphs to kernelsMaria BånkestadUppsala University , 2025
@phdthesis{bankestad2025thesis, title = {Structured models for scientific machine learning: From graphs to kernels}, author = {B{\aa}nkestad, Maria}, school = {Uppsala University}, year = {2025}, } - LoG
Ising on the Graph: Task-Specific Graph Subsampling via the Ising ModelMaria Bånkestad, Jennifer R. Andersson, Sebastian Mair, and 1 more authorIn Proceedings of the Third Learning on Graphs Conference , 2025Reducing a graph while preserving its overall structure is an important problem with many applications. Typically, the reduction approaches either remove edges (sparsification) or merge nodes (coarsening) in an unsupervised way with no specific downstream task in mind. In this paper, we present an approach for subsampling graph structures using an Ising model defined on either the nodes or edges and learning the external magnetic field of the Ising model using a graph neural network. Our approach is task-specific as it can learn how to reduce a graph for a specific downstream task in an end-to-end fashion. The utilized loss function of the task does not even have to be differentiable. We showcase the versatility of our approach on distinct applications, including image segmentation, explainability for graph classification, 3D shape sparsification, and sparse approximate matrix inverse determination.
@inproceedings{bankestad2025ising, title = {Ising on the Graph: Task-Specific Graph Subsampling via the Ising Model}, author = {B{\aa}nkestad, Maria and Andersson, Jennifer R. and Mair, Sebastian and Sj{\"o}lund, Jens}, booktitle = {Proceedings of the Third Learning on Graphs Conference}, series = {Proceedings of Machine Learning Research}, volume = {269}, publisher = {PMLR}, year = {2025}, }
2024
-
Flexible SE(2) graph neural networks with applications to PDE surrogatesMaria Bånkestad, Olof Mogren, and Aleksis PirinenarXiv preprint arXiv:2405.20287, 2024This paper presents a novel approach for constructing graph neural networks equivariant to 2D rotations and translations and leveraging them as PDE surrogates on non-gridded domains. We show that aligning the representations with the principal axis allows us to sidestep many constraints while preserving SE(2) equivariance. By applying our model as a surrogate for fluid flow simulations and conducting thorough benchmarks against non-equivariant models, we demonstrate significant gains in terms of both data efficiency and accuracy.
@article{bankestad2024flexible, title = {Flexible SE(2) graph neural networks with applications to PDE surrogates}, author = {B{\aa}nkestad, Maria and Mogren, Olof and Pirinen, Aleksis}, journal = {arXiv preprint arXiv:2405.20287}, year = {2024}, } - RSC Adv.
Carbohydrate NMR chemical shift prediction by GeqShift employing E(3) equivariant graph neural networksMaria Bånkestad, Keven M. Dorst, Göran Widmalm, and 1 more authorRSC Advances, 2024Carbohydrates, vital components of biological systems, are well-known for their structural diversity. Nuclear Magnetic Resonance (NMR) spectroscopy plays a crucial role in understanding their intricate molecular arrangements and is essential in assessing and verifying the molecular structure of organic molecules. An important part of this process is to predict the NMR chemical shift from the molecular structure. This work introduces a novel approach that leverages E(3) equivariant graph neural networks to predict carbohydrate NMR spectra. Notably, our model achieves a substantial reduction in mean absolute error, up to threefold, compared to traditional models that rely solely on two-dimensional molecular structure. Even with limited data, the model excels, highlighting its robustness and generalization capabilities.
@article{bankestad2024carbohydrate, title = {Carbohydrate NMR chemical shift prediction by GeqShift employing E(3) equivariant graph neural networks}, author = {B{\aa}nkestad, Maria and Dorst, Keven M. and Widmalm, G{\"o}ran and R{\"o}nnols, Jerk}, journal = {RSC Advances}, volume = {14}, pages = {26585--26595}, year = {2024}, doi = {10.1039/D4RA03428G}, }
2023
- TMLR
Variational Elliptical ProcessesMaria Bånkestad, Jens Sjölund, Jalil Taghia, and 1 more authorTransactions on Machine Learning Research, 2023We present elliptical processes—a family of non-parametric probabilistic models that subsumes Gaussian processes and Student’s t processes. This generalization includes a range of new heavy-tailed behaviors while retaining computational tractability. Elliptical processes are based on a representation of elliptical distributions as a continuous mixture of Gaussian distributions. We parameterize this mixture distribution as a spline normalizing flow, which we train using variational inference. The proposed form of the variational posterior enables a sparse variational elliptical process applicable to large-scale problems. We highlight advantages compared to Gaussian processes through regression and classification experiments. Elliptical processes can supersede Gaussian processes in several settings, including cases where the likelihood is non-Gaussian or when accurate tail modeling is essential.
@article{bankestad2023variational, title = {Variational Elliptical Processes}, author = {B{\aa}nkestad, Maria and Sj{\"o}lund, Jens and Taghia, Jalil and Sch{\"o}n, Thomas B.}, journal = {Transactions on Machine Learning Research}, issn = {2835-8856}, year = {2023}, url = {https://openreview.net/forum?id=djN3TaqbdA}, }
2022
-
Graph-based neural acceleration for nonnegative matrix factorizationJens Sjölund, and Maria BånkestadarXiv preprint arXiv:2202.00264, 2022We describe a graph-based neural acceleration technique for nonnegative matrix factorization that builds upon a connection between matrices and bipartite graphs that is well-known in certain fields, e.g., sparse linear algebra, but has not yet been exploited to design graph neural networks for matrix computations. We first consider low-rank factorization more broadly and propose a graph representation of the problem suited for graph neural networks. Then, we focus on the task of nonnegative matrix factorization and propose a graph neural network that interleaves bipartite self-attention layers with updates based on the alternating direction method of multipliers. Our empirical evaluation on synthetic and two real-world datasets shows that we attain substantial acceleration, even though we only train in an unsupervised fashion on smaller synthetic instances.
@article{sjolund2022graph, title = {Graph-based neural acceleration for nonnegative matrix factorization}, author = {Sj{\"o}lund, Jens and B{\aa}nkestad, Maria}, journal = {arXiv preprint arXiv:2202.00264}, year = {2022}, } - ICML-WPre-training Transformers for Molecular Property Prediction Using Reaction PredictionJohan Broberg, Maria Bånkestad, and Erik YlipääIn ICML 2022 2nd AI for Science Workshop , 2022
Molecular property prediction is essential in chemistry, especially for drug discovery applications. However, available molecular property data is often limited, encouraging the transfer of information from related data. Transfer learning has had a tremendous impact in fields like computer vision and natural language processing signaling for its potential in molecular property prediction. We present a pre-training procedure for molecular representation learning using reaction data and use it to pre-train a SMILES transformer. We fine-tune and evaluate the pre-trained model on 12 molecular property prediction tasks from MoleculeNet within physical chemistry, biophysics, and physiology and show a statistically significant positive effect on 5 of the 12 tasks compared to a non-pre-trained baseline model.
@inproceedings{broberg2022pretraining, title = {Pre-training Transformers for Molecular Property Prediction Using Reaction Prediction}, author = {Broberg, Johan and B{\aa}nkestad, Maria and Ylip{\"a}{\"a}, Erik}, booktitle = {ICML 2022 2nd AI for Science Workshop}, year = {2022}, }
2019
- Constructing the Matrix Multilayer Perceptron and its Application to the VAEJalil Taghia, Maria Bånkestad, Fredrik Lindsten, and 1 more authorarXiv preprint arXiv:1902.01182, 2019
Like most learning algorithms, the multilayer perceptrons (MLP) is designed to learn a vector of parameters from data. However, in certain scenarios we are interested in learning structured parameters (predictions) in the form of symmetric positive definite matrices. Here, we introduce a variant of the MLP, referred to as the matrix MLP, that is specialized at learning symmetric positive definite matrices. We demonstrate its usefulness by applying the matrix MLP to the task of learning the covariance matrix of the Gaussian posterior in the variational autoencoder.
@article{taghia2019constructing, title = {Constructing the Matrix Multilayer Perceptron and its Application to the VAE}, author = {Taghia, Jalil and B{\aa}nkestad, Maria and Lindsten, Fredrik and Sch{\"o}n, Thomas B.}, journal = {arXiv preprint arXiv:1902.01182}, year = {2019}, }
2018
- NeuroImageBayesian uncertainty quantification in linear models for diffusion MRIJens Sjölund, Anders Eklund, Evren Özarslan, and 3 more authorsNeuroImage, 2018
Diffusion MRI (dMRI) is a valuable tool in the assessment of tissue microstructure. By fitting a model to the dMRI signal it is possible to derive various quantitative features. Several of the most popular dMRI signal models are expansions in an appropriately chosen basis, where the coefficients are determined using some variation of least-squares. However, such approaches lack any notion of uncertainty, which could be valuable in e.g. group analyses. In this work, we use a probabilistic interpretation of linear least-squares methods to recast popular dMRI models as Bayesian ones.
@article{sjolund2018bayesian, title = {Bayesian uncertainty quantification in linear models for diffusion MRI}, author = {Sj{\"o}lund, Jens and Eklund, Anders and {\"O}zarslan, Evren and Herberthson, Magnus and B{\aa}nkestad, Maria and Knutsson, Hans}, journal = {NeuroImage}, volume = {175}, pages = {272--285}, year = {2018}, doi = {10.1016/j.neuroimage.2018.03.059}, }