Datasets
← Home
17 datasets found (tag “sft”)
Filter by tag: sft (17) · reasoning (12) · distilled (9) · agentic (8) · tool-calling (8) · code (5) · instruction-tuning (5) · web-text (4) · pretraining (4) · math (4) · filtered (3) · multi-turn (3) · multilingual (3) · open-harness (3) · deepseek-r1 (2) · reinforcement-learning (2) · dpo (2) · deepseek-v4 (2) · preference (2) · deepseek-v3.2 (2) · synthetic (2) · function-calling (2) · conversations (2) · human-written (2) · opencode (1) · openhands (1) · qwen3-coder (1) · real-user (1) · research (1) · rlhf (1)
Sort: popular · downloads · stars · newest
- Nemotron-SFT-Agentic-v2 — ~1.2M synthetic tool-use trajectories seeded from UltraTool, ToolEyes, and more. (0 downloads, 0 stars, cc-by-4.0, tool-calling, agentic, multi-turn)
- Nemotron-Agentic-v1 — 335k synthetic multi-turn trajectories for interactive tool use. (0 downloads, 0 stars, cc-by-4.0, tool-calling, agentic, multi-turn)
- Glaive Function Calling v2 — 113k function-calling examples for training tool-use models. (0 downloads, 0 stars, apache-2.0, tool-calling, function-calling, sft)
- Hermes 3 Dataset — NousResearch's official post-training dataset behind the Hermes model line. (0 downloads, 1 stars, apache-2.0, sft, instruction-tuning, reasoning)
- OpenAssistant Conversations (oasst1) — Crowdsourced human-written multi-turn conversation trees from Open Assistant. (1 downloads, 0 stars, apache-2.0, sft, human-written, conversations)
- Aya Dataset — Cohere For AI's human-curated multilingual instruction dataset across 65+ languages. (0 downloads, 0 stars, apache-2.0, sft, multilingual, human-curated)
- Nemotron-Post-Training-Dataset-v2 — Multilingual Nemotron post-training data supporting Nemotron-Nano-9B-v2. (0 downloads, 0 stars, cc-by-4.0, sft, reinforcement-learning, multilingual)
- Llama-Nemotron-Post-Training-Dataset — NVIDIA's 30M+ example SFT+RL post-training dataset behind Llama-3-Nemotron. (0 downloads, 0 stars, cc-by-4.0, sft, reinforcement-learning, reasoning)
- Capybara — High-quality multi-turn synthetic conversations emphasizing reasoning and depth. (0 downloads, 0 stars, apache-2.0, sft, multi-turn, reasoning)
- Orca-AgentInstruct-1M-v1 — Microsoft's ~1M-example synthetic agentic-instruction dataset. (0 downloads, 0 stars, cdla-permissive-2.0, sft, agentic, synthetic)
- Magpie-Pro-300K (Filtered) — Filtered Magpie-Pro synthetic instructions self-synthesized from Llama-3-70B-Instruct. (0 downloads, 0 stars, llama3, sft, synthetic, filtered)
- Databricks Dolly 15k — 15k human-written instruction/response pairs across 8 task categories. (0 downloads, 0 stars, cc-by-sa-3.0, sft, human-written)
- SlimOrca — Deduplicated, cleaned subset of Orca with GPT-4-quality reasoning traces. (2 downloads, 2 stars, mit, sft, reasoning, deduplicated)
- OpenHermes-2.5 — ~1M GPT-4-generated instruction pairs deduplicated to ShareGPT/ChatML format. (0 downloads, 0 stars, unknown, sft, instruction-tuning, gpt4-distilled)
- Tulu 3 SFT Mixture — AI2's ~939k-example multilingual SFT corpus with full provenance labels. (1 downloads, 0 stars, odc-by, sft, instruction-tuning, multilingual)
- SmolTalk — Curated Magpie-Ultra-based instruction mix used to train the SmolLM2 family. (0 downloads, 0 stars, mixed (apache-2.0 + source licenses), sft, instruction-tuning, magpie)
- SmolTalk2 — The 2025 SFT mixture behind SmolLM3: decontaminated multi-task instruction data spanning chat, reasoning, and tool use. (0 downloads, 0 stars, mixed (apache-2.0 + source licenses), sft, instruction-tuning, multi-task)