Datasets
← Home
4 datasets found (tag “pretraining”)
Filter by tag: sft (17) · reasoning (12) · distilled (9) · agentic (8) · tool-calling (8) · code (5) · instruction-tuning (5) · web-text (4) · pretraining (4) · math (4) · filtered (3) · multi-turn (3) · multilingual (3) · open-harness (3) · deepseek-r1 (2) · reinforcement-learning (2) · dpo (2) · deepseek-v4 (2) · preference (2) · deepseek-v3.2 (2) · synthetic (2) · function-calling (2) · conversations (2) · human-written (2) · opencode (1) · openhands (1) · qwen3-coder (1) · real-user (1) · research (1) · rlhf (1)
Sort: popular · downloads · stars · newest
- RedPajama-Data-1T — Together AI's 1.2T-token open reproduction of the LLaMA pretraining mix. (0 downloads, 0 stars, apache-2.0, pretraining, web-text, open-reproduction)
- Dolma 3 Pool — AI2's 9T-token pretraining pool considered for training OLMo 3. (0 downloads, 0 stars, odc-by, pretraining, web-text, multi-source)
- Falcon RefinedWeb — ~600B-token filtered, deduplicated CommonCrawl extract used to pretrain Falcon. (0 downloads, 0 stars, apache-2.0, pretraining, web-text, filtered)
- FineWeb-Edu — Education-filtered subset of FineWeb web text for pretraining. (0 downloads, 0 stars, odc-by, pretraining, web-text, filtered)