ICML Workshop: Protocol Learning | Pluralis Research

Workshop

Researchers working on related topics joined us for a series of talks, followed by a selection of posters tackling various parts of the decentralized training stack. The workshop was organized in collaboration with Professor Namhoon Lee and POSTECH.

Protocol Learning

Training frontier foundation models today demands massive, co-located clusters of high-end GPUs - accessible only to a handful of the most well-resourced organizations. Protocol Learning removes this co-location requirement, enabling multi-participant training of foundation models across open, permissionless networks of globally distributed compute, where no single participant has, or can ever obtain, a full copy of the model.

This requires solving hard open problems in low-bandwidth model parallelism, asynchronous distributed optimization, supporting heterogeneous hardware, fault-tolerant training systems, Byzantine robustness, and trustless verification. The workshop convened the researchers advancing these building blocks to define the challenges ahead and chart a research roadmap for training the next generation of community-owned frontier models with self-sustaining economics.

Talks

Konstantin Mishchenko, Research Scientist, Meta FAIR - Distributed Optimization Stories
Anastasiia Koloskova, Assistant Professor, University of Zurich - FedMuon: Federated Learning with Bias-corrected LMO-based Optimization
Aaron Defazio, Research Scientist, Meta - Scaling Schedule-Free and Learning Rate Free for Continual, Anytime, LLM Training
Kun Yuan, Assistant Professor, Peking University - Efficient and Effective Communication Topologies for Decentralized Deep Learning
Hadi Mohaghegh Dolatabadi, Research Scientist, Pluralis Research - Agora: a tale of compute, models, and signals
Samuel Horvath, Assistant Professor, MBZUAI - Protocols for Distributed Adaptation and Efficient Training of Foundation Models
Riccardo Patana, Head of Strategy, Product and Safety, Pluralis Research - Collective Sovereignty: The Third Path in AI

Lightning Talks

Chamin Hewa Koneputugodage, Research Scientist, Pluralis Research - Factored Gossip DiLoCo
Paul Janson, Mila Quebec AI Institute / Concordia University - Stabilizing Native Low-Rank LLM Pretraining
Rustem Islamov, PhD Student, University of Basel - Safe-EF: Error Feedback for Nonsmooth Constrained Optimization
Hyunji Jung, MS student, POSTECH - Mitigating Staleness in Asynchronous Pipeline Parallelism via Basis Rotation
Dongyeop Lee, PhD Student, POSTECH - FRESCO: A Novel Consistency Control for Asynchronous Pipeline Parallel Training
Tongtian Zhu, PhD student, Zhejiang University - On the Surprising Effectiveness of Single Global Merging in Decentralized Learning
Egor Shulgin, PhD Student, KAUST - Deriving Hyperparameter Scaling Laws via Optimization Theory

Photos

Poster Sessions

Sungbin Shin, Hyunji Jung
Mitigating Staleness in Asynchronous Pipeline Parallelism via Basis Rotation
Zhiwei Bai
Adaptive Preconditioners Trigger Loss Spikes in Adam
Jin Lee
SPARe: Stacked Parallelism with Adaptive Reordering for Fault-Tolerant LLM Pretraining Systems with 100k+ GPUs
Andrej Jovanović
LoRDO: Distributed Low-Rank Optimization with Infrequent Communication
Egor Shulgin
General Analysis of LMO-based Optimizers: Beyond Bounded Variance
Xingyu Qu
Can Muon Fine-tune Adam-Pretrained Models?
Benjamin Thérien
MuLoCo: Muon is a Practical Inner Optimizer for DiLoCo
Paul Janson
Stabilizing Native Low-Rank LLM Pretraining
Zhuoli Ouyang
RMNP: Row-Momentum Normalized Preconditioning for Scalable Matrix-Based Optimization
Jeffrey T. H. Wong (Imperial College London)
A3: an Analytical Low-Rank Approximation Framework for Attention
Philip Zmushko (ISTA), Egor Petrov (Yandex Research)
One-Step Gradient Delay is Not a Barrier for Large-Scale Asynchronous Pipeline Parallel LLM Pretraining
Han Shi (Huawei)
POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation
Dongyeop Lee
FRESCO: A Novel Consistency Control for Asynchronous Pipeline Parallel Training