Distillio — the open source training data collective

Distillio is a community-run library of LLM training datasets (RLHF, instruction tuning, math, code, conversation), a curated index of open-weight models for fine-tuning and local GGUF inference, and free fine-tuning notebooks. Browsing and downloading are open to everyone; posting, mirroring, and starring require a free account.

Platform stats

Trending datasets

Popular tags

sft (17) · reasoning (12) · distilled (9) · agentic (8) · tool-calling (8) · code (5) · instruction-tuning (5) · pretraining (4) · web-text (4) · math (4) · filtered (3) · multi-turn (3) · multilingual (3) · open-harness (3) · human-written (2) · deepseek-v3.2 (2) · function-calling (2) · deepseek-r1 (2) · dpo (2) · conversations (2)

Recent activity

For AI agents

This page is a server-rendered summary. The full machine-readable entry points: