Datasets
← Home
36 datasets found
Filter by tag: sft (17) · reasoning (12) · distilled (9) · agentic (8) · tool-calling (8) · code (5) · instruction-tuning (5) · web-text (4) · pretraining (4) · math (4) · filtered (3) · multi-turn (3) · multilingual (3) · open-harness (3) · deepseek-r1 (2) · reinforcement-learning (2) · dpo (2) · deepseek-v4 (2) · preference (2) · deepseek-v3.2 (2) · synthetic (2) · function-calling (2) · conversations (2) · human-written (2) · opencode (1) · openhands (1) · qwen3-coder (1) · real-user (1) · research (1) · rlhf (1)
Sort: popular · downloads · stars · newest
- Nemotron-SFT-OpenCode-v1 — Agentic instruction-tuning data for the OpenCode CLI framework. (1 downloads, 1 stars, cc-by-4.0 (per Nemotron post-training family; not independently re-verified on this specific card), tool-calling, agentic, skills)
- Nemotron-SWE-v1 — 59k agent trajectories targeting SWE-Bench-style repo navigation and issue-fixing. (0 downloads, 0 stars, cc-by-4.0 (a few subsets are bsd-3-clause — see HF dataset viewer), tool-calling, agentic, code)
- When2Call — Teaches when NOT to call a tool: 15k SFT + 9k DPO preference pairs. (0 downloads, 0 stars, cc-by-4.0, tool-calling, agentic, preference)
- Nemotron-SFT-Agentic-v2 — ~1.2M synthetic tool-use trajectories seeded from UltraTool, ToolEyes, and more. (0 downloads, 0 stars, cc-by-4.0, tool-calling, agentic, multi-turn)
- Nemotron-Agentic-v1 — 335k synthetic multi-turn trajectories for interactive tool use. (0 downloads, 0 stars, cc-by-4.0, tool-calling, agentic, multi-turn)
- Glaive Function Calling v2 — 113k function-calling examples for training tool-use models. (0 downloads, 0 stars, apache-2.0, tool-calling, function-calling, sft)
- xlam-function-calling-60k — 60k function-calling examples across 3,673 APIs via Salesforce's APIGen. (0 downloads, 0 stars, mixed signal (cited as cc-by-4.0, but source repo says research-only; gated), tool-calling, function-calling, agentic)
- Hermes 3 Dataset — NousResearch's official post-training dataset behind the Hermes model line. (0 downloads, 1 stars, apache-2.0, sft, instruction-tuning, reasoning)
- kimi-k3-coding-and-debugging-traces — Agentic coding/tool-use trajectories from Kimi K3, generated via moonshiner. (4 downloads, 3 stars, unspecified (source: Kimi K3, custom Kimi K3 License — permits derivatives, no distillation restriction), tool-calling, agentic, coding)
- KIMI-K2.5-550000x — 550,000 high-depth reasoning traces distilled from Kimi K2.5. (0 downloads, 0 stars, unspecified (source: Kimi K2.5, Modified MIT + no distillation restriction), reasoning, distilled, kimi-k2.5)
- qwen3-coder-480b-distill-mini — 9,543 cleaned code-reasoning samples distilled from Qwen3-Coder-480B-A35B-Instruct. (1 downloads, 3 stars, apache-2.0, code, reasoning, distilled)
- deepseek-v3.2-speciale-openr1-math-3k — Math reasoning traces distilled from DeepSeek-V3.2-Speciale. (0 downloads, 1 stars, unspecified (source: DeepSeek-V3.2-Speciale, MIT + explicit distillation permission), reasoning, distilled, math)
- deepseek-v3.2-speciale-1000x — High-reasoning-depth traces from DeepSeek-V3.2-Speciale for distillation. (0 downloads, 0 stars, unspecified (source: DeepSeek-V3.2-Speciale, MIT + explicit distillation permission), reasoning, distilled, deepseek-v3.2)
- Deepseek-v4-pro-max-distill-1500x — Coding and math reasoning traces distilled from DeepSeek V4 Pro Max. (0 downloads, 0 stars, unspecified (source: DeepSeek V4 Pro Max, MIT + explicit distillation permission), reasoning, distilled, code)
- DeepSeek-V4-Distill-8000x — Reasoning SFT distilled from DeepSeek-V4-Flash as teacher. (12 downloads, 5 stars, unspecified (source: DeepSeek-V4-Flash, MIT + explicit distillation permission), reasoning, distilled, deepseek-v4)
- OpenAssistant Conversations (oasst1) — Crowdsourced human-written multi-turn conversation trees from Open Assistant. (1 downloads, 0 stars, apache-2.0, sft, human-written, conversations)
- RedPajama-Data-1T — Together AI's 1.2T-token open reproduction of the LLaMA pretraining mix. (0 downloads, 0 stars, apache-2.0, pretraining, web-text, open-reproduction)
- Aya Dataset — Cohere For AI's human-curated multilingual instruction dataset across 65+ languages. (0 downloads, 0 stars, apache-2.0, sft, multilingual, human-curated)
- Dolma 3 Pool — AI2's 9T-token pretraining pool considered for training OLMo 3. (0 downloads, 0 stars, odc-by, pretraining, web-text, multi-source)
- Falcon RefinedWeb — ~600B-token filtered, deduplicated CommonCrawl extract used to pretrain Falcon. (0 downloads, 0 stars, apache-2.0, pretraining, web-text, filtered)
- Nemotron-Post-Training-Dataset-v2 — Multilingual Nemotron post-training data supporting Nemotron-Nano-9B-v2. (0 downloads, 0 stars, cc-by-4.0, sft, reinforcement-learning, multilingual)
- Llama-Nemotron-Post-Training-Dataset — NVIDIA's 30M+ example SFT+RL post-training dataset behind Llama-3-Nemotron. (0 downloads, 0 stars, cc-by-4.0, sft, reinforcement-learning, reasoning)
- AM-DeepSeek-R1-Distilled-1.4M — 1.4M verified general reasoning problems, ~0.9M distilled from DeepSeek-R1-671B. (0 downloads, 0 stars, cc-by-nc-4.0, reasoning, distilled, deepseek-r1)
- OpenR1-Math-220k — 220k verified math reasoning traces distilled from DeepSeek-R1 on NuminaMath 1.5. (0 downloads, 0 stars, apache-2.0, reasoning, math, distilled)
- FineWeb-Edu — Education-filtered subset of FineWeb web text for pretraining. (0 downloads, 0 stars, odc-by, pretraining, web-text, filtered)
- LMSYS-Chat-1M — 1M real user-model conversations from 25+ LLMs collected via Chatbot Arena. (0 downloads, 0 stars, custom (gated), conversations, real-user, research)
- Capybara — High-quality multi-turn synthetic conversations emphasizing reasoning and depth. (0 downloads, 0 stars, apache-2.0, sft, multi-turn, reasoning)
- Orca-AgentInstruct-1M-v1 — Microsoft's ~1M-example synthetic agentic-instruction dataset. (0 downloads, 0 stars, cdla-permissive-2.0, sft, agentic, synthetic)
- Magpie-Pro-300K (Filtered) — Filtered Magpie-Pro synthetic instructions self-synthesized from Llama-3-70B-Instruct. (0 downloads, 0 stars, llama3, sft, synthetic, filtered)
- Databricks Dolly 15k — 15k human-written instruction/response pairs across 8 task categories. (0 downloads, 0 stars, cc-by-sa-3.0, sft, human-written)
- SlimOrca — Deduplicated, cleaned subset of Orca with GPT-4-quality reasoning traces. (2 downloads, 2 stars, mit, sft, reasoning, deduplicated)
- OpenHermes-2.5 — ~1M GPT-4-generated instruction pairs deduplicated to ShareGPT/ChatML format. (0 downloads, 0 stars, unknown, sft, instruction-tuning, gpt4-distilled)
- UltraFeedback (binarized) — Binarized chosen/rejected preference pairs from UltraFeedback, ready for DPO. (0 downloads, 0 stars, mit, preference, dpo, rlhf)
- Tulu 3 SFT Mixture — AI2's ~939k-example multilingual SFT corpus with full provenance labels. (1 downloads, 0 stars, odc-by, sft, instruction-tuning, multilingual)
- SmolTalk — Curated Magpie-Ultra-based instruction mix used to train the SmolLM2 family. (0 downloads, 0 stars, mixed (apache-2.0 + source licenses), sft, instruction-tuning, magpie)
- SmolTalk2 — The 2025 SFT mixture behind SmolLM3: decontaminated multi-task instruction data spanning chat, reasoning, and tool use. (0 downloads, 0 stars, mixed (apache-2.0 + source licenses), sft, instruction-tuning, multi-task)