One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

AI GlossaryWords you meet while using AI

Dolma 3

An open training dataset of about 1.26 billion documents built by the Allen Institute for AI to train the Olmo 3 model

In plain words

Think of Dolma 3 as the entire textbook set the Allen Institute for AI used to teach its language model, Olmo 3. To use a cooking analogy: instead of just serving the finished dish (the model), they released the full shopping receipt showing exactly what ingredients went in and how much of each. Inside are over 1.2 billion documents, sorted into categories, ranging from literary works and everyday conversations to customer service text and technical documentation.

Most language models release only the finished product and keep this ingredient list hidden, making it essentially impossible to trace which texts produced which abilities. Dolma 3, by contrast, is open from start to finish, letting researchers remove specific documents and observe how the model's abilities change as a result. It was precisely because this data was fully public that a Georgia Tech research team was able to trace which texts gave rise to the model's ability to read human emotions and social context.

In the end, Dolma 3 is more than just training material — it also functions as an experimental tool for pinpointing where a language model's capabilities actually come from.

How it shows up in the news

The article explains that "Olmo 3, built by the Allen Institute for AI, has its training data (Dolma 3), intermediate checkpoints, and evaluation tools all fully released." It's important not to confuse Dolma 3 with the Olmo 3 model itself — Dolma 3 refers specifically to the document dataset used to train that model. The model and the material it learned from are two separate things.

See also

Stories using this term

Browse every entry