principles.fyi · the brain · concept

The Pile

A curated 800GB open dataset that mixes web text with books, code, papers, and more.

web + books + code + papers + ... -> one diverse 800GB mixture

The Pile (2020, by EleutherAI) is an influential open pretraining dataset built by deliberately combining 22 sources — not just web text but also books, GitHub code, arXiv papers, Wikipedia, PubMed, and so on. The idea is that a richer, more diverse mix teaches a broader model than web text alone. It's a landmark example of curating a corpus on purpose rather than just scraping, and it helped make open LLM research possible. Newer open datasets like Dolma follow the same diverse-mixture spirit at a larger scale.

Appears in

Nearby in the brain