Group picture from Mathematics of Transformers workshop 2025 in DESY, Hamburg

Workshop on the Mathematics of Transformers (Aug 2026)

This workshop will focus on the transformer architecture and its underlying (self-)attention mechanisms that gained substantial interest in recent years. Despite their empirical success and groundbreaking advances in natural language processing, computer vision, and scientific computing, the mathematical understanding of transformers is still in its infancy, with many fundamental questions only starting to be posed and addressed. 

After the first edition of the Workshop on the Mathematics of Transformers having taken place at DESY, Hamburg in September 2025, we are excited to host the second edition of this workshop at the Mathematical Institute in Oxford.

We aim to bring together researchers with backgrounds in multi-agent dynamics, optimal transport, and PDEs, to initiate discussions on a variety of aspects connected to the theoretical principles governing transformers. By fostering discussions, we seek to advance this young and rapidly evolving research field, uncovering new mathematical perspectives on transformer models.

 

Location and Date

The workshop will be hosted at the Mathematical Institute of the University of Oxford (Andrew Wiles Building) in lecture hall L5 on August 21, 2026.

 

Registration

Registration for the workshop is now closed.

 

Confirmed Speakers

Albert Alcalde, University of Erlangen–Nuremberg
Edoardo Calvello, California Institute of Technology
Mitia Duerinckx, Université Libre de Bruxelles
William Gibson, University of Oxford
Samira Kabri, Helmholtz Imaging, DESY
Thomas Jacob Maranzatto, University of Maryland
Jan Peszek, University of Warsaw
Domènec Ruiz-Balet, University of Barcelona

 

Schedule

The preliminary schedule of the workshop is as follows.

Friday, August 21stSchedule
9:00–9:30Registration
Get-to-know-each-other
9:30–9:45Welcome & Introduction
9:45–10:20                                                                                                                                                                                                                                                                                                                                                                                                                                                                   

Talk 1: Anisotropic interaction in self-attention layers
by Samira Kabri

Abstract: In recent years, analysing transformer architectures from the perspective of interacting particle systems as proposed by Sander et al. (2022) has gained strong interest. The work by Geshkovski et al. (2024) connects the mean-field dynamics induced by self-attention layers to a time-continuous evolution of probability measures on a high dimensional sphere. In particular, it analyses the long-time clustering behaviour under isotropic interaction, which corresponds to a fixed choice of the so-called key, query and value parameters. In this talk, we relax the assumptions on the parameters such that the interaction between particles becomes anisotropic. We discuss how anisotropic interaction influences the structure of the self-attention layer and the stationary states of the dynamics.

10:20–10:55

Talk 2: Diverse Dynamical Behaviors in Deep Linear Transformers
by Thomas Jacob Maranzatto and Jan Peszek

Abstract: In this talk, we discuss recent work on the dynamics arising in linearized transformers when the token embedding space is two-dimensional. In contrast to much of the recent literature, we do not rely on the gradient flow framework. Instead, we identify a lower-dimensional invariant manifold and derive a reduced state-space description that allows us to analyze token dynamics beyond consensus and clustering. We discuss this reduction and its connection to the Kuramoto model, and present examples of token convergence, clustering, cycling, and bifurcations across different parameter regimes.
This talk is based on the paper "On the Diverse Dynamical Behaviors Arising in Deep Linear Transformers".

10:55–11:20Group Photo
Coffee break
(included)
11:20–11:55

Talk 3: Hardmax attention dynamics
by Domènec Ruiz-Balet

Abstract: In this talk we will discuss some theoretical aspects of hardmax attention models from a PDE viewpoint.

11:55–12:30Talk 4: Quantifying Low-Temperature Concentration in Mean-Field Transformers
by Albert Alcalde
Abstract: We study the interaction between the temperature of softmax self-attention and the inference depth in a continuous-time mean-field transformer model. Under symmetry and spectral gap assumptions on the attention weights, we show that, at low temperature, the token distribution rapidly approaches a concentrated state determined jointly by the initial distribution and the query, key, and value matrices. This state is metastable: at inverse temperature $\beta$, it persists for inference times of order $\log \beta$, revealing a quantitative trade-off between these two hyperparameters. Numerical experiments support this initial concentration timescale and suggest that, at longer times and finite temperature, the dynamics transition to a distinct regime associated with a dominant eigenspace of the value matrix.
12:30–14:00Lunch break at Café Pi (Mathematical Institute) (1.5h)
(self-paid)
14:00–14:35                                                                                                                                                                                                                                                                                   

Talk 5: Transformers for Operator Learning on Probability Measures: From Theory to Scientific Workflows
by Edoardo Calvello

Abstract: Attention mechanisms, the core of transformer architectures that power modern large language models, have been widely successful at modeling nonlocal correlations in data. Motivated by recent formulations of attention as a measure-to-measure mapping, in this talk we show how the mathematical properties of the attention mechanism enable the design of neural architectures approximating nonlinear operators on probability measures. We leverage this perspective to design novel methods for nonlinear data assimilation, surpassing the performance of ensemble Kalman method baselines. Starting from a universal approximation theorem we obtain for the conditioning operator on densities, we discuss key considerations and open questions towards the development of scalable transformer-based Bayesian inference workflows.

14:35–15:10

Talk 6: Uniform Scaling Limits in AdamW trained Transformers
by William Gibson

Abstract: We study the large-depth limit of transformers trained with AdamW, by modelling the hidden-state dynamics as an interacting particle system (IPS) coupled through the attention mechanism.
Under appropriate scaling of the attention heads, we prove that the joint dynamics of the hidden states and backpropagated variables converge in $L^2$, uniformly over the initial condition, to the solution of a forward–backward system of ODEs at rate $\mathcal O(L^{-1}+L^{-1/3}H^{-1/2})$. Here, $L$ and $H$ denote the depth and number of heads of the transformer, respectively. The limiting system of ODEs can be identified with a McKean–Vlasov ODE (MVODE) when the attention heads do not incorporate causal masking. By using the flow maps associated with this MVODE and applying concentration of measure techniques, we obtain bounds on the difference between the discrete and continuous models that are uniform over compact sets of initial conditions. As this is achieved without resorting to a covering argument, the constants in our bounds are independent of the number of tokens. Furthermore, under a suitable adaptation to AdamW, the bounds become independent of the token embedding dimension.

15:10–15:40Coffee break
(included)
15:40–17:30Discussion Session
Format: TBA
(integral part of workshop)
17:30–18:30Social Event
Oxford walking tour
to Queen's College
18:30–21:30                Dinner at Queen's College
18:30 – Pre-drinks
19:00 – Call for dinner
(self-paid but subsidized, organized,
dress code: business casual,
registration via oxforduniversitystores link)

 

Travel and Accommodation

The closest international airport is London Heathrow Airport (LHR). An alternative is London Gatwick Airport (LGW).
The Oxford airline bus frequently commutes between Oxford and LHR/LGW.
For accommodation, it may be worth checking college accommodation.

 

Organizers

José A. Carrillo (University of Oxford)
Daniel Kelly (University of Oxford)
Konstantin Riedl (University of Oxford)
Tim Roith (Technical University of Munich)

 

We gratefully acknowledge support from the following agencies in funding the workshop:

G-Research (KR's G-Research January 2026 grant).

The Advanced Grant Nonlocal-CPD: "Nonlocal PDEs for Complex Particle Dynamics: Phase Transitions, Patterns and Synchronization" of the European Research Council Executive Agency (ERC) (JAC's research group).

Last updated on 29 Jul 2026, 12:57pm. Please contact us with feedback and comments about this page.