principles.fyi · the brain · concept
The Pile
A curated 800GB open dataset that mixes web text with books, code, papers, and more.
web + books + code + papers + ... -> one diverse 800GB mixture
The Pile (2020, by EleutherAI) is an influential open pretraining dataset built by deliberately combining 22 sources — not just web text but also books, GitHub code, arXiv papers, Wikipedia, PubMed, and so on. The idea is that a richer, more diverse mix teaches a broader model than web text alone. It's a landmark example of curating a corpus on purpose rather than just scraping, and it helped make open LLM research possible. Newer open datasets like Dolma follow the same diverse-mixture spirit at a larger scale.
Appears in
- What they eat LLMs in the Wild · pt 3