Workshop Group Picture

In the last few years, large language models (LLMs) have undergone significant development: in their capabilities, in the variety of their applications, and in their prevalence in our world. They have evolved from relatively fluent text generators to proficient reasoning models with which most people interact on a daily basis.

While this progress has, in part, been driven by scale, data, and compute as well as careful engineering, an additional contributor to this development has been an improved theoretical understanding of the transformer architecture on which these models are based. In particular, the self-attention mechanism distinguishes the transformer from machine learning models which came before it. The mathematical structure of this component has proven considerably richer than might be expected from its relatively simple formulation. There is, therefore, an ongoing programme which seeks to provide insight into the underlying behaviour of LLMs by establishing a firm theoretical understanding of the transformer.

On August 21, 2026, Oxford Mathematics hosted the second edition of the “Workshop on the Mathematics of Transformers”, following a first edition at DESY in Hamburg in September 2025. The workshop aimed to bring together researchers in multi-agent dynamics, optimal transport, and partial differential equations, with the common objective to improve the theoretical understanding of the transformer and the underlying self-attention mechanism.

Close to 40 participants, coming from different institutions, countries and disciplines, gathered to discuss open questions and results that we hope will shape a fundamental understanding of LLMs.

Programme

The workshop featured six invited talks given by,

  • Samira Kabri (Helmholtz Imaging, DESY) on Anisotropic interaction in self-attention layers,
  • Thomas Jacob Maranzatto (University of Maryland) and Jan Peszek (University of Warsaw) on Diverse dynamical behaviors in deep linear transformers,
  • Domènec Ruiz-Balet (University of Barcelona) on Hardmax attention dynamics,
  • Albert Alcalde (University of Erlangen–Nuremberg) on Quantifying low-temperature concentration in mean-field transformers,
  • Edoardo Calvello (California Institute of Technology) on Transformers for operator learning on probability measures: from theory to scientific workflows,
  • William Gibson (University of Oxford) on Uniform scaling limits in AdamW-trained transformers.

Most of the talks were concerned with the dynamics of self-attention at inference. The starting point in this line of research is a setting where tokens converge to consensus, a small number of clusters, or concentrate on lower-dimensional spaces. Several of the results presented investigated this picture in different regimes. Allowing the interaction between tokens to be anisotropic changes the stationary states of the dynamics. Working outside the gradient flow framework, on a lower-dimensional invariant manifold, makes it possible to observe cycling and bifurcations as well as clustering, and connects the dynamics to the Kuramoto model. The hardmax limit of the attention mechanism leads to a different class of PDE problems. A further result showed that at low temperature the token distribution concentrates quickly into a metastable state, and that the metastability of this state depends on a hyperparameter.

Two talks went beyond this setting. One used the description of attention as a map between probability measures as a design principle rather than as a tool for analysis, and applied it to nonlinear data assimilation, where it improved on ensemble Kalman baselines. The second considered the large-depth limit of transformers trained with AdamW, and established convergence of the hidden state dynamics to a forward-backward system of ODEs, with bounds that do not depend on the number of tokens.

Outlook

The response to this workshop, and to its predecessor, has been overwhelmingly positive. What came out of the day was a sharper sense of the open problems: the discussions brought them into clearer focus and left participants with a better sense of which questions are worth pursuing and which tools are likely to reach them. Several exchanges pointed towards new collaborations across institutions. Given the extent of the engagement and curiosity for the subject, we were very happy to find that a third edition of the workshop was a late-night topic of discussion.

Thanks

The workshop was jointly organised by José A. Carrillo, Daniel Kelly, Konstantin Riedl (University of Oxford) and Tim Roith (Technical University of Munich), and kindly supported by G-Research and the Advanced Grant Nonlocal-CPD: “Nonlocal PDEs for Complex Particle Dynamics: Phase Transitions, Patterns and Synchronization” of the European Research Council Executive Agency (ERC).

Posted on 4 Sep 2026, 10:58am. Please contact us with feedback and comments about this page.