Latest Release: Olmo3 (November 2025)
We introduce Olmo 3, a new family of 7B and 32B models with Instruct (7B) and Think (7B, 32B) variants available for use! These models are pre-trained on the Dolma 3 dataset and post-trained on the Dolci datasets. Long chain-of-thought thinking improves reasoning tasks like math and coding. We have released all code, checkpoints, and associated training details.
Installation
Create or activate a Python virtual environment with a Python version ≥ 3.10, then install PyTorch. We use Dolma Toolkit for tokenizing data, OLMo-core for training, and OLMES for evaluation. We recommend using a separate virtual environment for each stage.
Training
We recommend installing OLMo-core from source with the following commands.
git clone https://github.com/allenai/OLMo-core.git
cd OLMo-core
pip install -e .[all]
Alternatively, you can install OLMo-core from PyPI with pip install ai2-olmo-core.
We will run the official pretraining script found in src/examples/official/OLMo3/OLMo-3-1025-7B-pretrain-1.py. You will need to install flash-attn.
Defining the experiment config
The key components and hyperparameters are defined through the ExperimentConfig class in src/olmo_core/script_utils.py. To override any fields in the config at runtime, we can simply pass them in as command-line options. For instance, adding --train_module.rank_microbatch_size=8192 would update the rank_microbatch_size field within the train_module part of the config.
To validate that our overrides are applied correctly, we can print the config without actually launching training using the --dry-run flag. Note that the script also requires a --save-folder argument, which specifies where model checkpoints should be saved.
python src/scripts/official/OLMo3/OLMo-3-1025-7B-pretrain-1.py \
--save-folder=models/tutorial-run-01 \
--train_module.rank_microbatch_size=8192 \
--dry-run
The model architecture being trained in the script is olmo3_7B. We can replace this with a number of different preset model configurations defined by classmethods of TransformerConfig. To construct a new model config, we recommend creating a new classmethod under TransformerConfig. Keep in mind that as you change the model size and architecture you'll likely want to adjust hyperparameters and performance settings such as the learning rate and microbatch size.
Note that because the Olmo team decided to extend the number of train tokens partway through training, our official pretraining run is split into OLMo-3-1025-7B-pretrain-1.py (with a train duration of 5T tokens and a hard stop at step 597046) and OLMo-3-1025-7B-pretrain-2.py (which resumes training with a new train duration to 7T tokens). You should specifying a train duration that makes sense for your own use case using trainer_config.max_duration and trainer_config.hard_stop (e.g., ~150B tokens at 1 epoch for a compute-optimal run).
Launching the run
To launch our first run, we'll use overrides to disable the in-loop perplexity evaluator, in-loop downstream task evaluator, and terminate the training at step 100.
torchrun --nproc-per-node=gpu src/scripts/official/OLMo3/OLMo-3-1025-7B-pretrain-1.py \
--save-folder=models/tutorial-run-01 \
--train_module.rank_microbatch_size=8192 \
--trainer.callbacks.lm_evaluator.enabled=false \
--trainer.callbacks.downstream_evaluator.enabled=false \
--trainer.hard_stop='{value: 100, unit: steps}'
This should take about 2 hours on 8 NVIDIA H100s. You should see model checkpoints saved to models/tutorial-run-01/. Note that the script automatically attempts to resume from the latest checkpoint; you can modify this behavior if desired through trainer_config.load_strategy.
Converting models to HuggingFace format
To convert our trained model to HuggingFace format, we can use the following conversion script. Note that you will need to have transformers installed.
python src/examples/huggingface/convert_checkpoint_to_hf.py \
--checkpoint-input-path models/tutorial-run-01/step100 \
--huggingface-output-dir hf_models/tutorial-run-01/step100
Evaluation
We use the OLMES (Open Language Model Evaluation System) repository for evaluation.
git clone https://github.com/allenai/olmes.git
cd olmes
pip install -e . # for vLLM support, use pip install -e ".[gpu]"
The following command evaluates a model on a given set of tasks. The output predictions and performance will be stored in workspace.
olmes \
--model models/tutorial-run-01/step100 \
--task arc_challenge::olmes hellaswag::olmes \
--output-dir workspace
You can use the --inspect flag to see a sample prompt, or --dry-run to inspect the launch command.