Itay Lavie, Kirsten Fischer, Andrey Lekov, Frederic Van Maele, Zohar Ringel, and Moritz Helias
In submission, 2026
Attention is the key mechanism underlying in-context learning in transformers, and attention patterns have been observed empirically to emerge abruptly during training. We present a Bayesian theory of feature learning in attention; we then focus on how the copy subcircuit in the first layer of an induction head is learned by analyzing a single-layer softmax attention network trained on a copy task. We derive a closed-form posterior over the attention matrix and reduce it to a low-dimensional order parameter space. This reduction reveals a phase transition in the amount of training data, which we verify using both Bayesian sampling and standard training with Adam. We contrast our results with linear attention and find that softmax attention exhibits a first-order phase transition while in linear attention an initial second-order phase transition is followed by a smooth, continuous evolution toward the structured attention pattern. Our work provides a first-principles theoretical account of the abrupt emergence of the copy subcircuit, reminiscent of the one observed in training large language models.
@misc{lavie2026phase,title={Phase Transitions in Attention: A Bayesian Theory of Copy Head Emergence},author={Lavie, Itay and Fischer, Kirsten and Lekov, Andrey and Van Maele, Frederic and Ringel, Zohar and Helias, Moritz},year={2026},publisher={arXiv},}
@misc{lavie2026statistical,title={Statistical Properties of Training \& Generalization},author={Lavie, Itay and Levi, Noam and Kahn, Yonatan},year={2026},publisher={arXiv},}
Kernel ridge regression (KRR) and Gaussian processes (GPs) are fundamental tools in statistics and machine learning, with recent applications to highly over-parameterized deep neural networks. The ability of these tools to learn a target function is directly related to the eigenvalues of their kernel sampled on the input data distribution. Targets that have support on higher eigenvalues are more learnable. However, solving such eigenvalue problems on real-world data remains a challenge. Here, we consider cross-dataset learnability and show that one may use eigenvalues and eigenfunctions associated with highly idealized data measures to reveal spectral bias on complex datasets and bound learnability on real-world data. This allows us to leverage various symmetries that realistic kernels manifest to unravel their spectral bias.
@misc{lavie2024demystifying,title={Demystifying Spectral Bias on Real-World Data},author={Lavie, Itay and Ringel, Zohar},year={2024},publisher={arXiv},}
In Proceedings of the 41st International Conference on Machine Learning, 2024
We study inductive bias in Transformers in the infinitely over-parameterized Gaussian process limit and argue transformers tend to be biased towards more permutation symmetric functions in sequence space. We show that the representation theory of the symmetric group can be used to give quantitative analytical predictions when the dataset is symmetric to permutations between tokens. We present a simplified transformer block and solve the model at the limit, including accurate predictions for the learning curves and network outputs. We show that in common setups, one can derive tight bounds in the form of a scaling law for the learnability as a function of the context length. Finally, we argue WikiText dataset, does indeed possess a degree of permutation symmetry.
@inproceedings{lavie2024towards,title={Towards Understanding Inductive Bias in Transformers: A View From Infinity},author={Lavie, Itay and Gur-Ari, Guy and Ringel, Zohar},booktitle={Proceedings of the 41st International Conference on Machine Learning},year={2024},pages={26043--26069},publisher={PMLR},}
In ICLR 2024 Workshop on Mathematical and Empirical Understanding of Foundation Models, 2024
@inproceedings{lavie2024spectral,title={Transformers' Spectral Bias and The Symmetric Group},author={Lavie, Itay and Gur-Ari, Guy and Ringel, Zohar},booktitle={ICLR 2024 Workshop on Mathematical and Empirical Understanding of Foundation Models},year={2024},}
Ronny Frumkin, Eric Kuflik, Itay Lavie, and Tal Silverwater
Physical Review Letters. Authors listed alphabetically, per the high-energy phenomenology convention, 2023
We study the general properties of the freeze-out of a thermal relic. We give analytic estimates of the relic abundance for an arbitrary freeze-out process, showing when instantaneous freeze-out is appropriate and how it can be corrected when freeze-out is slow. This is used to generalize the relationship between the dark matter mass and coupling that matches the observed abundance. The result encompasses well-studied particular examples, such as weakly interacting massive particles (WIMPs), strongly interacting massive particles, coannihilation, coscattering, inverse decays, and forbidden channels, and generalizes beyond them. In turn, this gives an approximate perturbative unitarity bound on the dark matter mass for an arbitrary thermal freeze-out process. We show that going beyond the maximal masses allowed for freeze-out via dark matter self-annihilations predicts that there are nearly degenerate states with the dark matter and that the dark matter is generically metastable.
@article{frumkin2023roadmap,title={Roadmap to Thermal Dark Matter beyond the Weakly Interacting Dark Matter Unitarity Bound},author={Frumkin, Ronny and Kuflik, Eric and Lavie, Itay and Silverwater, Tal},journal={Physical Review Letters},volume={130},number={17},pages={171001},year={2023},publisher={American Physical Society},doi={10.1103/PhysRevLett.130.171001},}