## Workshop

Researchers working on related topics joined us for a series of talks, followed by a selection of posters tackling various parts of the decentralized training stack. The workshop was organized in collaboration with Professor Namhoon Lee and POSTECH.

## Protocol Learning

Training frontier foundation models today demands massive, co-located clusters of high-end GPUs - accessible only to a handful of the most well-resourced organizations. Protocol Learning removes this co-location requirement, enabling multi-participant training of foundation models across open, permissionless networks of globally distributed compute, where no single participant has, or can ever obtain, a full copy of the model.

This requires solving hard open problems in low-bandwidth model parallelism, asynchronous distributed optimization, supporting heterogeneous hardware, fault-tolerant training systems, Byzantine robustness, and trustless verification. The workshop convened the researchers advancing these building blocks to define the challenges ahead and chart a research roadmap for training the next generation of community-owned frontier models with self-sustaining economics.

## Talks

[Konstantin Mishchenko, Research Scientist, Meta FAIR - Distributed Optimization Stories](https://www.youtube.com/watch?v=i4v0rO5yILQ)  
 [Anastasiia Koloskova, Assistant Professor, University of Zurich - FedMuon: Federated Learning with Bias-corrected LMO-based Optimization](https://www.youtube.com/watch?v=3ysoli1ROzc)  
 [Aaron Defazio, Research Scientist, Meta - Scaling Schedule-Free and Learning Rate Free for Continual, Anytime, LLM Training](https://www.youtube.com/watch?v=dJ5dZb-YZUI)  
 [Kun Yuan, Assistant Professor, Peking University - Efficient and Effective Communication Topologies for Decentralized Deep Learning](https://www.youtube.com/watch?v=X3jVoN1BbGA)  
 [Hadi Mohaghegh Dolatabadi, Research Scientist, Pluralis Research - Agora: a tale of compute, models, and signals](https://www.youtube.com/watch?v=wP7Nyg3sOvk)  
 [Samuel Horvath, Assistant Professor, MBZUAI - Protocols for Distributed Adaptation and Efficient Training of Foundation Models](https://www.youtube.com/watch?v=JtUlEJ5ihMM)  
 [Riccardo Patana, Head of Strategy, Product and Safety, Pluralis Research - Collective Sovereignty: The Third Path in AI](https://www.youtube.com/watch?v=XML8xqNycxY)

## Lightning Talks

[Chamin Hewa Koneputugodage, Research Scientist, Pluralis Research - Factored Gossip DiLoCo](https://www.youtube.com/watch?v=2IeFLg8fkYk)  
 [Paul Janson, Mila Quebec AI Institute / Concordia University - Stabilizing Native Low-Rank LLM Pretraining](https://www.youtube.com/watch?v=WpXOSTV2-NY)  
 [Rustem Islamov, PhD Student, University of Basel - Safe-EF: Error Feedback for Nonsmooth Constrained Optimization](https://www.youtube.com/watch?v=4OS-IX0YSaY)  
 [Hyunji Jung, MS student, POSTECH - Mitigating Staleness in Asynchronous Pipeline Parallelism via Basis Rotation](https://www.youtube.com/watch?v=rKezWQ_HfsY)  
 [Dongyeop Lee, PhD Student, POSTECH - FRESCO: A Novel Consistency Control for Asynchronous Pipeline Parallel Training](https://www.youtube.com/watch?v=_lRT8EE33WI)  
 [Tongtian Zhu, PhD student, Zhejiang University - On the Surprising Effectiveness of Single Global Merging in Decentralized Learning](https://www.youtube.com/watch?v=6heVcFvSUeU)  
 [Egor Shulgin, PhD Student, KAUST - Deriving Hyperparameter Scaling Laws via Optimization Theory](https://www.youtube.com/watch?v=kMJ2iTTf5Ks)

## Photos

## Poster Sessions

Sungbin Shin, Hyunji Jung  
[Mitigating Staleness in Asynchronous Pipeline Parallelism via Basis Rotation](https://arxiv.org/abs/2602.03515)  
Zhiwei Bai  
[Adaptive Preconditioners Trigger Loss Spikes in Adam](https://arxiv.org/abs/2506.04805)  
Jin Lee  
[SPARe: Stacked Parallelism with Adaptive Reordering for Fault-Tolerant LLM Pretraining Systems with 100k+ GPUs](https://arxiv.org/abs/2603.00357)  
Andrej Jovanović  
[LoRDO: Distributed Low-Rank Optimization with Infrequent Communication](https://arxiv.org/abs/2602.04396)  
Egor Shulgin  
[General Analysis of LMO-based Optimizers: Beyond Bounded Variance](https://icml.cc/virtual/2026/poster/61267)  
Xingyu Qu  
[Can Muon Fine-tune Adam-Pretrained Models?](https://arxiv.org/abs/2605.10468)  
Benjamin Thérien  
[MuLoCo: Muon is a Practical Inner Optimizer for DiLoCo](https://arxiv.org/abs/2505.23725)  
Paul Janson  
[Stabilizing Native Low-Rank LLM Pretraining](https://www.researchgate.net/publication/400812518_Stabilizing_Native_Low-Rank_LLM_Pretraining)  
Zhuoli Ouyang  
[RMNP: Row-Momentum Normalized Preconditioning for Scalable Matrix-Based Optimization](https://arxiv.org/abs/2603.20527)  
Jeffrey T. H. Wong (Imperial College London)  
[A3: an Analytical Low-Rank Approximation Framework for Attention](https://arxiv.org/abs/2505.12942)  
Philip Zmushko (ISTA), Egor Petrov (Yandex Research)  
[One-Step Gradient Delay is Not a Barrier for Large-Scale Asynchronous Pipeline Parallel LLM Pretraining](https://arxiv.org/abs/2606.30634)  
Han Shi (Huawei)  
[POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation](https://arxiv.org/pdf/2603.05500)  
Dongyeop Lee  
FRESCO: A Novel Consistency Control for Asynchronous Pipeline Parallel Training
