Datasets
← Home
12 datasets found (tag “reasoning”)
Filter by tag: sft (17) · reasoning (12) · distilled (9) · agentic (8) · tool-calling (8) · code (5) · instruction-tuning (5) · web-text (4) · pretraining (4) · math (4) · filtered (3) · multi-turn (3) · multilingual (3) · open-harness (3) · deepseek-r1 (2) · reinforcement-learning (2) · dpo (2) · deepseek-v4 (2) · preference (2) · deepseek-v3.2 (2) · synthetic (2) · function-calling (2) · conversations (2) · human-written (2) · opencode (1) · openhands (1) · qwen3-coder (1) · real-user (1) · research (1) · rlhf (1)
Sort: popular · downloads · stars · newest
- Hermes 3 Dataset — NousResearch's official post-training dataset behind the Hermes model line. (0 downloads, 1 stars, apache-2.0, sft, instruction-tuning, reasoning)
- KIMI-K2.5-550000x — 550,000 high-depth reasoning traces distilled from Kimi K2.5. (0 downloads, 0 stars, unspecified (source: Kimi K2.5, Modified MIT + no distillation restriction), reasoning, distilled, kimi-k2.5)
- qwen3-coder-480b-distill-mini — 9,543 cleaned code-reasoning samples distilled from Qwen3-Coder-480B-A35B-Instruct. (1 downloads, 3 stars, apache-2.0, code, reasoning, distilled)
- deepseek-v3.2-speciale-openr1-math-3k — Math reasoning traces distilled from DeepSeek-V3.2-Speciale. (0 downloads, 1 stars, unspecified (source: DeepSeek-V3.2-Speciale, MIT + explicit distillation permission), reasoning, distilled, math)
- deepseek-v3.2-speciale-1000x — High-reasoning-depth traces from DeepSeek-V3.2-Speciale for distillation. (0 downloads, 0 stars, unspecified (source: DeepSeek-V3.2-Speciale, MIT + explicit distillation permission), reasoning, distilled, deepseek-v3.2)
- Deepseek-v4-pro-max-distill-1500x — Coding and math reasoning traces distilled from DeepSeek V4 Pro Max. (0 downloads, 0 stars, unspecified (source: DeepSeek V4 Pro Max, MIT + explicit distillation permission), reasoning, distilled, code)
- DeepSeek-V4-Distill-8000x — Reasoning SFT distilled from DeepSeek-V4-Flash as teacher. (12 downloads, 5 stars, unspecified (source: DeepSeek-V4-Flash, MIT + explicit distillation permission), reasoning, distilled, deepseek-v4)
- Llama-Nemotron-Post-Training-Dataset — NVIDIA's 30M+ example SFT+RL post-training dataset behind Llama-3-Nemotron. (0 downloads, 0 stars, cc-by-4.0, sft, reinforcement-learning, reasoning)
- AM-DeepSeek-R1-Distilled-1.4M — 1.4M verified general reasoning problems, ~0.9M distilled from DeepSeek-R1-671B. (0 downloads, 0 stars, cc-by-nc-4.0, reasoning, distilled, deepseek-r1)
- OpenR1-Math-220k — 220k verified math reasoning traces distilled from DeepSeek-R1 on NuminaMath 1.5. (0 downloads, 0 stars, apache-2.0, reasoning, math, distilled)
- Capybara — High-quality multi-turn synthetic conversations emphasizing reasoning and depth. (0 downloads, 0 stars, apache-2.0, sft, multi-turn, reasoning)
- SlimOrca — Deduplicated, cleaned subset of Orca with GPT-4-quality reasoning traces. (2 downloads, 2 stars, mit, sft, reasoning, deduplicated)