https://github.com/databricks/megablocks

Last synced: 2 months ago
JSON representation

Host: GitHub
URL: https://github.com/databricks/megablocks
Owner: databricks
License: apache-2.0
Created: 2023-01-26T00:24:56.000Z (over 2 years ago)
Default Branch: main
Last Pushed: 2025-04-29T17:29:57.000Z (3 months ago)
Last Synced: 2025-05-06T16:07:46.607Z (2 months ago)
Language: Python
Size: 3.94 MB
Stars: 1,347
Watchers: 16
Forks: 193
Open Issues: 43
Metadata Files:
- Readme: README.md
- Contributing: CONTRIBUTING.md
- License: LICENSE

Awesome Lists containing this project

StarryDivineSky - databricks/megablocks - LM 集成，支持 MoE 的数据、专家和流水线并行训练。MegaBlocks 的 dMoE性能优于使用 Tutel 训练的 MoE，速度提升高达 40%。MegaBlocks dMoE 通过将 MoE 重构为块稀疏操作，避免了令牌丢弃，同时保持了硬件效率。与使用 Megatron-LM 训练的密集 Transformer 相比，MegaBlocks dMoE 可以将训练速度提高 2.4 倍。安装 MegaBlocks可以使用 `pip install megablocks` 命令，并使用提供的脚本进行 Transformer MoE 和 dMoE 语言模型的预训练。 (A01_文本生成_文本对话 / 大语言对话模型及数据)
awesome-adaptive-computation - pytorch code
awesome-open-source-lms - Megablocks (MoE Training)
jimsghstars - databricks/megablocks - (Python)

README

        # :robot: MegaBlocks

MegaBlocks is a light-weight library for mixture-of-experts (MoE) training. The core of the system is efficient "dropless-MoE" ([dMoE](megablocks/layers/dmoe.py), [paper](https://arxiv.org/abs/2211.15841)) and standard [MoE](megablocks/layers/moe.py) layers.

MegaBlocks is integrated with [Megatron-LM](https://github.com/NVIDIA/Megatron-LM), where we support data, expert and pipeline parallel training of MoEs. Stay tuned for tighter integration with Databricks libraries and tools!

# :rocket: Performance

![MegaBlocks Performance](media/dropping_end_to_end.png)

MegaBlocks dMoEs outperform MoEs trained with [Tutel](https://github.com/microsoft/tutel) by up to **40%** compared to Tutel's best performing `capacity_factor` configuration. MegaBlocks dMoEs use a reformulation of MoEs in terms of block-sparse operations, which allows us to avoid token dropping without sacrificing hardware efficiency. In addition to being faster, MegaBlocks simplifies MoE training by removing the `capacity_factor` hyperparameter altogether. Compared to dense Transformers trained with [Megatron-LM](https://github.com/NVIDIA/Megatron-LM), MegaBlocks dMoEs can accelerate training by as much as **2.4x**. Check out our [paper](https://arxiv.org/abs/2211.15841) for more details!

# :building_construction: Installation

NOTE: This assumes you have `numpy` and `torch` installed.

**Training models with Megatron-LM:** We recommend using NGC's [`nvcr.io/nvidia/pytorch:23.09-py3`](https://catalog.ngc.nvidia.com/orgs/nvidia/containers/pytorch/tags) PyTorch container. The [Dockerfile](Dockerfile) builds on this image with additional dependencies. To build the image, run `docker build . -t megablocks-dev` and then `bash docker.sh` to launch the container. Once inside the container, install MegaBlocks with `pip install .`. See [Usage](#steam_locomotive-usage) for instructions on training MoEs with MegaBlocks + Megatron-LM.

**Using MegaBlocks in other packages:** To install the MegaBlocks package for use in other frameworks, run `pip install megablocks`. For example, [Mixtral-8x7B](https://mistral.ai/news/mixtral-of-experts/) can be run with [vLLM](https://github.com/vllm-project/vllm) + MegaBlocks with this installation method.

**Extras:** MegaBlocks has optional dependencies that enable additional features.

Installing `megablocks[gg]` enables dMoE computation with grouped GEMM. This feature is enabled by setting the `mlp_impl` argument to `grouped`. This is currently our recommended path for Hopper-generation GPUs.

Installing `megablocks[dev]` allows you to contribute to MegaBlocks and test locally. Installing `megablocks[testing]` allows you to test via Github Actions. If you've installed megablocks[dev], you can run pre-commit install to configure the pre-commit hook to automatically format the code.

MegaBlocks can be installed with all dependencies (except for `testing`) via the `megablocks[all]` package.

# :steam_locomotive: Usage

We provide scripts for pre-training Transformer MoE and dMoE language models under the [top-level directory](megablocks/). The quickest way to get started is to use one of the [experiment launch scripts](exp/). These scripts require a dataset in Megatron-LM's format, which can be created by following their [instructions](https://github.com/NVIDIA/Megatron-LM#data-preprocessing).

# :writing_hand: Citation

```

@article{megablocks,

  title={{MegaBlocks: Efficient Sparse Training with Mixture-of-Experts}},

  author={Trevor Gale and Deepak Narayanan and Cliff Young and Matei Zaharia},

  journal={Proceedings of Machine Learning and Systems},

  volume={5},

  year={2023}

}

```

ecosyste.ms

Data

Tools

Indexes

Applications

Experiments

Awesome

https://github.com/databricks/megablocks

Awesome Lists containing this project

README