Skip to main content

Olmo

Olmo is a family of open-source language models that delivers powerful AI capabilities while maintaining complete transparency into its training data and methods. Unlike many "open" models that obscure their training data, Olmo offers full transparency with permissive licensing for unrestricted commercial use. The most recent release is Olmo 3, which includes 7B and 32B parameter models along with instruction-tuned variants optimized for chat interfaces, complex reasoning, and instruction following. With all training code, data, and evaluation methods publicly available, the Olmo family is fully reproducible, making it an ideal choice for developers seeking a self-hosted solution they can trust and verify.

Key features across Olmo variants​

  • Complete transparency: Full access to training data, code, and evaluation methods
  • Permissive commercial licensing: Unrestricted commercial use without hidden restrictions
  • Instruction-tuned variants: Chat-optimized models with enhanced instruction-following capabilities
  • Reproducible training: All components needed to reproduce the models are publicly available
  • Efficient deployment: Support for quantization and optimization techniques

Possible use cases​

Olmo's capabilities make it suitable for a wide range of applications:

  • Content Generation: Create high-quality text for blogs, articles, and marketing copy
  • Question Answering: Build chatbots and virtual assistants that provide accurate, context-aware responses
  • Text Analysis: Perform sentiment analysis, summarization, and content classification
  • Research & Development: Use as a foundation for building specialized AI applications
  • Education: Create interactive learning tools and educational content
  • Enterprise Applications: Develop internal tools for document processing and knowledge management

Model variants​

Olmo 3 32B is the largest model in the most recent suite of open-source models in the Olmo family. The model is trained on up to 5.50 trillion tokens.

Inference code example​

You can use Olmo 3 with the Hugging Face transformers library:

from transformers import AutoModelForCausalLM, AutoTokenizer

olmo = AutoModelForCausalLM.from_pretrained("allenai/Olmo-3-1025-7B")
tokenizer = AutoTokenizer.from_pretrained("allenai/Olmo-3-1025-7B")

message = ["Language modeling is "]

inputs = tokenizer(message, return_tensors='pt', return_token_type_ids=False)
# optional verifying cuda
# inputs = {k: v.to('cuda') for k,v in inputs.items()}
# olmo = olmo.to('cuda')
response = olmo.generate(**inputs, max_new_tokens=100, do_sample=True, top_k=0, top_p=0.7, temperature=1)
print(tokenizer.batch_decode(response, skip_special_tokens=True)[0])
# >> 'Language modeling is a key component of any text-based application, but its effectiveness...'

For faster performance, you can quantize the model using the following method:

AutoModelForCausalLM.from_pretrained("allenai/Olmo-3-1025-7B",
torch_dtype=torch.float16,
load_in_8bit=True) # Requires bitsandbytes

Understanding the training pipeline and model types​

We release two post-trained variants of Olmo 3 models: Olmo-3-Think and Olmo-3-Instruct. The post-training process for Olmo-3-Think and Olmo-3-Instruct involved three stages: Supervised Finetuning (SFT), Preference Finetuning with Direct Preference Optimization (DPO), and Reinforcement Learning with Verifiable Rewards (RLVR). The two models largely differ in the settings used for these training stages and the data they were trained on. We describe the details of all stages, for both models, in this section.

Olmo 3 Instruct Models​

Olmo-3-Instruct models are general purpose chat models with a large variety of skills, such as instruction following or knowledge. They go through a multi-stage training pipeline, with each stage building on the previous one. For most software developers, you'll want the final Instruct models, but understanding this pipeline helps explain why other model checkpoints exist and when you might use them:

For most software developers, you'll want the final Instruct models, but understanding this pipeline helps explain why other model checkpoints exist and when you might use them.

Use for: Most production applications

  • Final stage of training, using Reinforcement Learning with Verifiable Rewards (RLVR)
  • Optimized for helpfulness, following complex instructions, and function calling.
  • Default choice for software developers
  • Unless you have specific research needs, start here

Models: Olmo-3-7B-Instruct

Function Calling with Instruct Models​

Olmo-3-Instruct models are optimized for function calling. If you are using vLLM to serve these models, the following is how you can run inference with function calling. Make sure you have the right version of vLLM.

vllm>=0.11.1

You can serve the models with the following command

vllm serve allenai/Olmo-3-7B-Instruct --enable-auto-tool-choice --tool-call-parser olmo3

This will give you an OpenAI-compatible endpoint that you can use with OpenAI Chat Completions API to output function calls.

from openai import OpenAI

tools = [
{
"type": "function",
"function": {
"name": "get_current_weather",
"description": "Get the current weather in a given location",
"parameters": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "The city and state, e.g. San Francisco, CA",
},
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"]},
},
"required": ["location"],
},
},
}
]

client = OpenAI(
base_url="http://localhost:[port]/v1",
api_key="EMPTY",
)

messages = [
{"role": "system", "content": "You are a helpful assistant that can use tools."},
{"role": "user", "content": "What's the weather like in Seattle?"},
]

chat_completion = client.chat.completions.create(
model="allenai/Olmo-3-7B-Instruct",
messages=messages,
tools=tools,
)

Use with MCP servers​

Olmo-3-Instruct models are also optimized for use with MCP servers. Once you serve the model with vLLM as described above, you can use an agents framework to run inference with MCP servers. Here is some example code with OpenAI Agents SDK using the Fetch MCP Server:

import asyncio
from agents import Agent, Runner, trace
from agents.mcp import MCPServer, MCPServerStdio
from agents.extensions.models.litellm_model import LitellmModel

async def run(mcp_server: MCPServer):
model = LitellmModel(
model=f"hosted_vllm/allenai/Olmo-3-7B-Instruct",
base_url="http://localhost:[port]/v1",
api_key="EMPTY",
)
agent = Agent(
name="Assistant",
instructions="You are a helpful assistant with access to a tool for fetching web pages.",
model=model,
mcp_servers=[mcp_server],
)

message = "Fetch the content of the page at https://allenai.org/ and summarize it."
result = await Runner.run(starting_agent=agent, input=message)
print(result.final_output)


async def main():
async with MCPServerStdio(
cache_tools_list=True,
params={
"command": "uvx",
"args": [
"mcp-server-fetch"
]
},
) as server:
with trace(workflow_name=f"Example workflow"):
await run(server)


if __name__ == "__main__":
asyncio.run(main())

Olmo 3 Think Models​

Olmo-3-Think models are reasoning models that are good at targeted skills like math, code, and precise instruction following. They go through a multi-stage training pipeline, with each stage building on the previous one. For most software developers, you'll want the final Instruct models, but understanding this pipeline helps explain why other model checkpoints exist and when you might use them:

Use for: Most production applications

  • Final stage of training, using Reinforcement Learning with Verifiable Rewards (RLVR)
  • Optimized for reasoning capabilities, math, code, and precise instruction following.
  • Default choice for software developers
  • Unless you have specific research needs, start here

Models: Olmo-3-7B-Think, Olmo-3-32B-Think

Visit the Olmo 3 artifacts page to find all model varients on Hugging Face.

Training Data​

Olmo 3 is trained on Dolma 3, the most recent version of Dolma. Dolma is a large-scale, high-quality dataset for language model pre-training. Dolma is designed to be transparent and accessible to the development and research communities.