Olmo
Olmo is a family of open-source language models that delivers powerful AI capabilities while maintaining complete transparency into its training data and methods. Unlike many "open" models that obscure their training data, Olmo offers full transparency with permissive licensing for unrestricted commercial use. The most recent release is Olmo 3, which includes 7B and 32B parameter models along with instruction-tuned variants optimized for chat interfaces, complex reasoning, and instruction following. With all training code, data, and evaluation methods publicly available, the Olmo family is fully reproducible, making it an ideal choice for developers seeking a self-hosted solution they can trust and verify.
Key features across Olmo variants
- Complete transparency: Full access to training data, code, and evaluation methods
- Permissive commercial licensing: Unrestricted commercial use without hidden restrictions
- Instruction-tuned variants: Chat-optimized models with enhanced instruction-following capabilities
- Reproducible training: All components needed to reproduce the models are publicly available
- Efficient deployment: Support for quantization and optimization techniques
Possible use cases
Olmo's capabilities make it suitable for a wide range of applications:
- Content Generation: Create high-quality text for blogs, articles, and marketing copy
- Question Answering: Build chatbots and virtual assistants that provide accurate, context-aware responses
- Text Analysis: Perform sentiment analysis, summarization, and content classification
- Research & Development: Use as a foundation for building specialized AI applications
- Education: Create interactive learning tools and educational content
- Enterprise Applications: Develop internal tools for document processing and knowledge management
Model variants
- Olmo 3 32B
- Olmo 3 7B
- Olmo 2 32B
- Olmo 2 13B
- Olmo 2 7B
Olmo 3 32B is the largest model in the most recent suite of open-source models in the Olmo family. The model is trained on up to 5.50 trillion tokens.
Olmo 3 7B is in the most recent suite of open-source models in the Olmo family. The model is trained on up to 5.93 trillion tokens.
Olmo 2 32B is the most capable and largest open-source model in the Olmo 2 family. The model is trained on up to 6 trillion tokens and further post-trained using Tülu 3.1.
Performance: Outperforms GPT-3.5 Turbo and GPT-4o Mini on average across standardized benchmarks
Olmo 2 13B is our mid-sized Olmo 2 model, offering an excellent balance of performance and efficiency. It demonstrates strong capabilities across a wide range of tasks while being more accessible than the 32B variant.
Performance: Strong mid-size model that rivals GPT-4o Mini on many reasoning tasks
Olmo 2 7B is our most efficient model, designed for applications where resource constraints are a primary consideration. Despite its smaller size, it maintains competitive performance across many benchmarks.
Performance: Competitive with GPT-3.5 Turbo and Gemini Flash on key benchmarks
Inference code example
You can use Olmo 3 with the Hugging Face transformers library:
from transformers import AutoModelForCausalLM, AutoTokenizer
olmo = AutoModelForCausalLM.from_pretrained("allenai/Olmo-3-1025-7B")
tokenizer = AutoTokenizer.from_pretrained("allenai/Olmo-3-1025-7B")
message = ["Language modeling is "]
inputs = tokenizer(message, return_tensors='pt', return_token_type_ids=False)
# optional verifying cuda
# inputs = {k: v.to('cuda') for k,v in inputs.items()}
# olmo = olmo.to('cuda')
response = olmo.generate(**inputs, max_new_tokens=100, do_sample=True, top_k=0, top_p=0.7, temperature=1)
print(tokenizer.batch_decode(response, skip_special_tokens=True)[0])
# >> 'Language modeling is a key component of any text-based application, but its effectiveness...'
For faster performance, you can quantize the model using the following method:
AutoModelForCausalLM.from_pretrained("allenai/Olmo-3-1025-7B",
torch_dtype=torch.float16,
load_in_8bit=True) # Requires bitsandbytes
Understanding the training pipeline and model types
We release two post-trained variants of Olmo 3 models: Olmo-3-Think and Olmo-3-Instruct. The post-training process for Olmo-3-Think and Olmo-3-Instruct involved three stages: Supervised Finetuning (SFT), Preference Finetuning with Direct Preference Optimization (DPO), and Reinforcement Learning with Verifiable Rewards (RLVR). The two models largely differ in the settings used for these training stages and the data they were trained on. We describe the details of all stages, for both models, in this section.
Olmo 3 Instruct Models
Olmo-3-Instruct models are general purpose chat models with a large variety of skills, such as instruction following or knowledge. They go through a multi-stage training pipeline, with each stage building on the previous one. For most software developers, you'll want the final Instruct models, but understanding this pipeline helps explain why other model checkpoints exist and when you might use them:
For most software developers, you'll want the final Instruct models, but understanding this pipeline helps explain why other model checkpoints exist and when you might use them.
- Instruct Model (Final)
- DPO Model (Stage 3)
- SFT Models (Stage 2)
- Base Models (Stage 1)
Use for: Most production applications
- Final stage of training, using Reinforcement Learning with Verifiable Rewards (RLVR)
- Optimized for helpfulness, following complex instructions, and function calling.
- Default choice for software developers
- Unless you have specific research needs, start here
Models: Olmo-3-7B-Instruct
Use for: Research and advanced training scenarios
- Third stage of training pipeline - preference-optimized but before final RLVR
- Useful for researchers studying Direct Preference Optimization techniques
- Starting point if you want to build your own RLVR training on top
- Replicating specific research results that used DPO checkpoints
- Most developers should skip this and use Instruct models
Models: Olmo-3-7B-DPO
Use for: Building your own preference optimization or studying instruction tuning
- Second stage of training pipeline - instruction-tuned but before preference optimization
- Good starting point if you want to train your own DPO or RLVR on top
- Useful for researchers studying supervised fine-tuning techniques
- Replicating research that used SFT checkpoints as baselines
- Applications should use Instruct models which perform better
Models: Olmo-3-7B-Instruct-SFT
Use for: Custom fine-tuning from scratch or research on pre-trained models
- First stage - raw pre-trained models without any instruction tuning
- Starting point when you want to do your own instruction tuning from scratch
- Domain-specific fine-tuning for specialized applications (legal, medical, etc.)
- Research on pre-training, studying model capabilities before instruction tuning
- Most developers should skip this unless doing custom training
Models: Olmo-3-1025-7B, Olmo-3-1125-32B
Function Calling with Instruct Models
Olmo-3-Instruct models are optimized for function calling. If you are using vLLM to serve these models, the following is how you can run inference with function calling. Make sure you have the right version of vLLM.
vllm>=0.11.1
You can serve the models with the following command
vllm serve allenai/Olmo-3-7B-Instruct --enable-auto-tool-choice --tool-call-parser olmo3
This will give you an OpenAI-compatible endpoint that you can use with OpenAI Chat Completions API to output function calls.
from openai import OpenAI
tools = [
{
"type": "function",
"function": {
"name": "get_current_weather",
"description": "Get the current weather in a given location",
"parameters": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "The city and state, e.g. San Francisco, CA",
},
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"]},
},
"required": ["location"],
},
},
}
]
client = OpenAI(
base_url="http://localhost:[port]/v1",
api_key="EMPTY",
)
messages = [
{"role": "system", "content": "You are a helpful assistant that can use tools."},
{"role": "user", "content": "What's the weather like in Seattle?"},
]
chat_completion = client.chat.completions.create(
model="allenai/Olmo-3-7B-Instruct",
messages=messages,
tools=tools,
)
Use with MCP servers
Olmo-3-Instruct models are also optimized for use with MCP servers. Once you serve the model with vLLM as described above, you can use an agents framework to run inference with MCP servers. Here is some example code with OpenAI Agents SDK using the Fetch MCP Server:
import asyncio
from agents import Agent, Runner, trace
from agents.mcp import MCPServer, MCPServerStdio
from agents.extensions.models.litellm_model import LitellmModel
async def run(mcp_server: MCPServer):
model = LitellmModel(
model=f"hosted_vllm/allenai/Olmo-3-7B-Instruct",
base_url="http://localhost:[port]/v1",
api_key="EMPTY",
)
agent = Agent(
name="Assistant",
instructions="You are a helpful assistant with access to a tool for fetching web pages.",
model=model,
mcp_servers=[mcp_server],
)
message = "Fetch the content of the page at https://allenai.org/ and summarize it."
result = await Runner.run(starting_agent=agent, input=message)
print(result.final_output)
async def main():
async with MCPServerStdio(
cache_tools_list=True,
params={
"command": "uvx",
"args": [
"mcp-server-fetch"
]
},
) as server:
with trace(workflow_name=f"Example workflow"):
await run(server)
if __name__ == "__main__":
asyncio.run(main())
Olmo 3 Think Models
Olmo-3-Think models are reasoning models that are good at targeted skills like math, code, and precise instruction following. They go through a multi-stage training pipeline, with each stage building on the previous one. For most software developers, you'll want the final Instruct models, but understanding this pipeline helps explain why other model checkpoints exist and when you might use them:
- Instruct Models (Final)
- DPO Models (Stage 3)
- SFT Models (Stage 2)
- Base Models (Stage 1)
Use for: Most production applications
- Final stage of training, using Reinforcement Learning with Verifiable Rewards (RLVR)
- Optimized for reasoning capabilities, math, code, and precise instruction following.
- Default choice for software developers
- Unless you have specific research needs, start here
Models: Olmo-3-7B-Think, Olmo-3-32B-Think
Use for: Research and advanced training scenarios
- Third stage of training pipeline - preference-optimized but before final RLVR
- Useful for researchers studying Direct Preference Optimization techniques
- Starting point if you want to build your own RLVR training on top
- Trained on preference pairs with reasoning chains.
- Most developers should skip this and use Instruct models
Models: Olmo-3-7B-Think, Olmo-3-32B-Think-DPO
Use for: Building your own preference optimization or studying instruction tuning
- Second stage of training pipeline - instruction-tuned but before preference optimization
- Good starting point if you want to train your own DPO or RLVR on top
- Useful for researchers studying supervised fine-tuning techniques
- Trained on Instruction - Completions, with reasoning chains
- Applications should use Instruct models which perform better
Models: Olmo-3-7B-Think-SFT, Olmo-3-32B-Think-SFT
Use for: Custom fine-tuning from scratch or research on pre-trained models
- First stage - raw pre-trained models without any instruction tuning
- Same base model as the one for the Instruct model.
Models: Olmo-3-1025-7B, Olmo-3-1125-32B
Visit the Olmo 3 artifacts page to find all model varients on Hugging Face.
Training Data
Olmo 3 is trained on Dolma 3, the most recent version of Dolma. Dolma is a large-scale, high-quality dataset for language model pre-training. Dolma is designed to be transparent and accessible to the development and research communities.