Skip to main content

Dolma

Dolma dataset​

Dolma is a large-scale, open English corpus containing trillions of tokens from a diverse mix of web content, academic publications, code, and more. Dolma was designed for large language model pretraining and is used as the training dataset for Olmo. Dolma 3, the latest version of Dolma, is openly available for download from the Hugging Face Hub under the ODC-BY license. All previous versions of Dolma can be downloaded from the original Dolma repository in the Hugging Face Hub.

Dolma toolkit​

The Dolma toolkit is an open-source, high-performance framework designed to enable scalable, efficient, and reproducible dataset curation for pretraining large AI models. It is built to handle billions of documents and hundreds of terabytes of raw text data, and supports both small-scale and production-grade pipelines. Below are the features of Dolma toolkit:

  • ⚡️ High performance. Optimized for speed and scalability, the Dolma toolkit can process massive datasets billions of documents in parallel.

  • 🧳 Portable. Runs seamlessly on a single machine, compute cluster, or in cloud environments.

  • 🏷 Built-in taggers. Includes prebuilt taggers for language detection, toxicity filtering, perplexity scoring, and more. Common filtering recipes such as those used in Gopher and C4 are included.

  • 🗑 Fast deduplication. Uses a Rust based Bloom filter for fast, memory-efficient duplicate detection across large corpora.

  • 🧩 Extensible. Fully modular and extensible, easily integrate your own taggers or filtering logic.

  • ☁️ Cloud support. Supports reading from and writing to local disk or S3-compatible cloud storage.

Dataset curation using the Dolma toolkit typically follows these four core steps:

  1. Tagging

    • Documents are tagged with metadata using built-in or custom taggers.
    • Tags can include attributes like language, toxicity level, text perplexity, regex patterns, and more.
  2. Deduplication (Optional)

    • Identical or near-duplicate documents are removed using a Bloom filter or other deduplication strategies.
    • Deduplication can be based on content or metadata.
  3. Mixing & Filtering

    • Documents are mixed and filtered based on tag values.
    • This includes up/down-sampling, test set decontamination, and complex filter configurations.
  4. Tokenization

    • Cleaned and filtered text is tokenized using any Hugging Face-compatible tokenizer to prepare it for model training.

The Dolma Toolkit is designed for reproducibility and extensibility—supporting practitioners in curating custom datasets or replicating Dolma’s construction for further research. The Dolma Toolkit can be used either as a Python library or a command-line tool. The CLI is accessible via the dolma command. To view all available commands and usage options, simply run:

pip install dolma
dolma --help

This will display a list of supported commands along with descriptions and usage flags.

usage: dolma [command] [options]

Command line interface for the DOLMa dataset processing toolkit

positional arguments:
{dedupe,mix,tag,list,stat,tokens}
dedupe Deduplicate documents or paragraphs using a bloom filter.
mix Mix documents from multiple streams.
tag Tag documents or spans of documents using one or more taggers. For a
list of available taggers, run `dolma list`.
list List available taggers.
stat Analyze the distribution of attributes values in a dataset.
tokens Tokenize documents using the provided tokenizer.

options:
-h, --help show this help message and exit
-c CONFIG, --config CONFIG
Path to configuration optional file

Dolma variants​

There are currently two main version of Dolma. Each version was designed for large language model pretraining on a set of Olmo models.

Dolma 3 Dolma 3 is the most recent open-source dataset in the Dolma family. The Olmo 3 model series is trained on Dolma 3.

Dolma Dolma is the original open-source dataset of the Dolma family and was used to train the Olmo and Olmo 2 model series.

Visit our GitHub repository to get started with the Dolma Toolkit.