Workshop on the Mathematics of Transformers (Aug 2026)
This workshop will focus on the transformer architecture and its underlying (self-)attention mechanisms that gained substantial interest in recent years. Despite their empirical success and groundbreaking advances in natural language processing, computer vision, and scientific computing, the mathematical understanding of transformers is still in its infancy, with many fundamental questions only starting to be posed and addressed.
After the first edition of the Workshop on the Mathematics of Transformers having taken place at DESY, Hamburg in September 2025, we are excited to host the second edition of this workshop at the Mathematical Institute in Oxford.
We aim to bring together researchers with backgrounds in multi-agent dynamics, optimal transport, and PDEs, to initiate discussions on a variety of aspects connected to the theoretical principles governing transformers. By fostering discussions, we seek to advance this young and rapidly evolving research field, uncovering new mathematical perspectives on transformer models.
Location and Date
The workshop will be hosted at the Mathematical Institute of the University of Oxford (Andrew Wiles Building) in lecture hall L5 on August 21, 2026.
Registration
Registration for the workshop is now closed.
Confirmed Speakers
Albert Alcalde, University of Erlangen–Nuremberg
Edoardo Calvello, California Institute of TechnologyMitia Duerinckx, Université Libre de Bruxelles
William Gibson, University of Oxford
Samira Kabri, Helmholtz Imaging, DESY
Thomas Jacob Maranzatto, University of Maryland
Jan Peszek, University of Warsaw
Domènec Ruiz-Balet, University of Barcelona
Schedule
The preliminary schedule of the workshop is as follows.
| Friday, August 21st | Schedule |
|---|---|
| 9:00–9:30 | Registration Get-to-know-each-other |
| 9:30–9:45 | Welcome & Introduction |
| 9:45–10:20 | Talk 1: Anisotropic interaction in self-attention layers Abstract: In recent years, analysing transformer architectures from the perspective of interacting particle systems as proposed by Sander et al. (2022) has gained strong interest. The work by Geshkovski et al. (2024) connects the mean-field dynamics induced by self-attention layers to a time-continuous evolution of probability measures on a high dimensional sphere. In particular, it analyses the long-time clustering behaviour under isotropic interaction, which corresponds to a fixed choice of the so-called key, query and value parameters. In this talk, we relax the assumptions on the parameters such that the interaction between particles becomes anisotropic. We discuss how anisotropic interaction influences the structure of the self-attention layer and the stationary states of the dynamics. |
| 10:20–10:55 | Talk 2: Diverse Dynamical Behaviors in Deep Linear Transformers Abstract: In this talk, we discuss recent work on the dynamics arising in linearized transformers when the token embedding space is two-dimensional. In contrast to much of the recent literature, we do not rely on the gradient flow framework. Instead, we identify a lower-dimensional invariant manifold and derive a reduced state-space description that allows us to analyze token dynamics beyond consensus and clustering. We discuss this reduction and its connection to the Kuramoto model, and present examples of token convergence, clustering, cycling, and bifurcations across different parameter regimes. |
| 10:55–11:20 | Group Photo Coffee break (included) |
| 11:20–11:55 | Talk 3: Hardmax attention dynamics Abstract: In this talk we will discuss some theoretical aspects of hardmax attention models from a PDE viewpoint. |
| 11:55–12:30 | Talk 4: Quantifying Low-Temperature Concentration in Mean-Field Transformers by Albert Alcalde Abstract: We study the interaction between the temperature of softmax self-attention and the inference depth in a continuous-time mean-field transformer model. Under symmetry and spectral gap assumptions on the attention weights, we show that, at low temperature, the token distribution rapidly approaches a concentrated state determined jointly by the initial distribution and the query, key, and value matrices. This state is metastable: at inverse temperature $\beta$, it persists for inference times of order $\log \beta$, revealing a quantitative trade-off between these two hyperparameters. Numerical experiments support this initial concentration timescale and suggest that, at longer times and finite temperature, the dynamics transition to a distinct regime associated with a dominant eigenspace of the value matrix. |
| 12:30–14:00 | Lunch break at Café Pi (Mathematical Institute) (1.5h) (self-paid) |
| 14:00–14:35 | Talk 5: Transformers for Operator Learning on Probability Measures: From Theory to Scientific Workflows Abstract: Attention mechanisms, the core of transformer architectures that power modern large language models, have been widely successful at modeling nonlocal correlations in data. Motivated by recent formulations of attention as a measure-to-measure mapping, in this talk we show how the mathematical properties of the attention mechanism enable the design of neural architectures approximating nonlinear operators on probability measures. We leverage this perspective to design novel methods for nonlinear data assimilation, surpassing the performance of ensemble Kalman method baselines. Starting from a universal approximation theorem we obtain for the conditioning operator on densities, we discuss key considerations and open questions towards the development of scalable transformer-based Bayesian inference workflows. |
| 14:35–15:10 | Talk 6: Uniform Scaling Limits in AdamW trained Transformers Abstract: We study the large-depth limit of transformers trained with AdamW, by modelling the hidden-state dynamics as an interacting particle system (IPS) coupled through the attention mechanism. |
| 15:10–15:40 | Coffee break (included) |
| 15:40–17:30 | Discussion Session Format: TBA (integral part of workshop) |
| 17:30–18:30 | Social Event Oxford walking tour to Queen's College |
| 18:30–21:30 | Dinner at Queen's College 18:30 – Pre-drinks 19:00 – Call for dinner (self-paid but subsidized, organized, dress code: business casual, registration via oxforduniversitystores link) |
Travel and Accommodation
The closest international airport is London Heathrow Airport (LHR). An alternative is London Gatwick Airport (LGW).
The Oxford airline bus frequently commutes between Oxford and LHR/LGW.
For accommodation, it may be worth checking college accommodation.
Organizers
José A. Carrillo (University of Oxford)
Daniel Kelly (University of Oxford)
Konstantin Riedl (University of Oxford)
Tim Roith (Technical University of Munich)
We gratefully acknowledge support from the following agencies in funding the workshop:
G-Research (KR's G-Research January 2026 grant).
The Advanced Grant Nonlocal-CPD: "Nonlocal PDEs for Complex Particle Dynamics: Phase Transitions, Patterns and Synchronization" of the European Research Council Executive Agency (ERC) (JAC's research group).