Node0 Dashboard
Node
Node0-7.5B is a permissionless, multi-participant, model-parallel pretraining run occurring over the Internet. Anyone with a single 16GB+ GPU can join. Node0 allows participants to collaboratively train a model much larger than any individual could train alone.
Node0-7.5B Training Complete
We're pleased to announce that Node0-7.5B has successfully completed training after 36B tokens over 3 weeks. This run featured over 300 active participants, with 1,642 individual GPUs contributing from 198 cities worldwide.
Thank you to everyone who participated and contributed to this project!
Global Node Distribution
Switch between signal transmission and power consumption views
Signal Transmission
- GPU Power Capacity
- Continents: 6
- Countries: 44
- Cities: 198
Click points for detailed information
Computing Nodes
Top Countries
- USA: 1038 nodes (0% active)
- Canada: 148 nodes (0% active)
- Romania: 62 nodes (0% active)
Metrics
- Training Loss: 2.816
- Tokens per Second: 3891.2
Revenue Ownership
Top 35 users
Training Details
Pretraining
- Dataset: FineWeb-Edu
- Architecture: OLMo2-7.5B
- Dataset Tokens: 1.3T
- Batch Size: 4M tokens
Node Requirements
- GPU: ≥ 16GB
- RAM: ≥ 32GB
Training Config
Protocol Model
Node0-7.5B is the first pretraining run open to the public where anyone can join and contribute with a consumer-grade (16GB+) GPU and an Internet connection. This run is a proof of concept — because it is using completely novel networking and distributed training implementations, many things are being tested for the first time. The purpose of this run is to evaluate stability and convergence in an open multi-party setting. A detailed technical report will follow after the run concludes.
Model Parallelism
Model parallelism is the collective name for strategies that break up a large model and distribute it across multiple GPUs, enabling training at scales impossible on a single device. It is the standard approach used to train today's largest models. Traditionally, this requires extremely fast communication between devices, as activations and gradients must be transferred at every step — something only feasible in datacenter environments with high-speed connections (≥ 100 Gbps). Node0's unique contribution is that, for the first time, model-parallel training is being carried out over Internet connections rather than within a datacenter.
Compression
To split the model itself over participants, we make use of a novel compression algorithm that constrains the output projection weights of Transformer blocks to a shared, learned low-dimensional subspace. Leveraging these constrained weights alongside the recursive structure of Transformers, we achieve over 99% compression in both forward and backward passes, while preserving convergence. For further details, see our Protocol Models paper.
| Rank | User | Joined | Active Time | TFLOP/s | Score |
|---|---|---|---|---|---|
| 1 | S skre0 |
9/17/2025 | 585.55 h | 17.14 | 36136.70 |
| 2 | F fr00000 |
9/16/2025 | 617.93 h | 12.82 | 28522.97 |
| 3 | rpslzero |
9/17/2025 | 783.80 h | 7.78 | 21958.13 |
| 4 | omk-pluralis |
9/19/2025 | 542.58 h | 9.85 | 19238.83 |
| 5 | Herb0x |
9/19/2025 | 541.13 h | 8.80 | 17139.09 |
| 6 | hellopluralis |
9/17/2025 | 266.58 h | 16.74 | 16061.30 |
| 7 | solodolo8 |
9/17/2025 | 707.28 h | 5.89 | 15008.58 |
| 8 | X xfr00 |
9/17/2025 | 508.28 h | 6.33 | 11580.18 |
| 9 | skre-support-battalion |
9/19/2025 | 481.58 h | 6.23 | 10803.99 |
| 10 | S sampluralis |
9/16/2025 | 681.77 h | 4.39 | 10776.26 |
Frequently Asked Questions
- What is Node0?
- How is this possible?
- What models could be trained in this way?
- Isn't the bandwidth too small?
- What if nodes drop out mid run?
- What if many nodes join?
- **Does this method have similar scaling laws to centralized development?