llms.txt Content
# YJTOON
> YJTOON is a free, open catalog of categorized reference datasets for AI agents and humans.
> 185 datasets across 27 categories, published as CC0 public domain data. Every dataset is
> available as JSON, YAML, and TOON (a compact encoding aimed at LLM prompts) from one URL,
> with no authentication, no API keys, and no signup. Rate limit: 120 requests per minute per IP.
Measured context (o200k_base tokenizer, all 185 datasets as served): TOON uses 166,870 tokens versus 246,704 for the pretty-printed JSON files (−32%), and 179,087 for minified JSON (−7%). 161 of 185 datasets are smaller in TOON than in minified JSON; table-shaped data gains most. In bytes: 1,199,518 B JSON vs 702,167 B TOON. Humans can read any dataset as a rendered HTML table at https://yjtoon.com/catalog/.
## Human-readable catalog (HTML)
- [Catalog index](https://yjtoon.com/catalog/): all 185 datasets, grouped by category
- [Per-category pages](https://yjtoon.com/catalog/{category}): e.g. https://yjtoon.com/catalog/programming-languages/
- [Per-dataset pages](https://yjtoon.com/catalog/{category}/{dataset}): the data rendered as tables, with exact sizes in each format, format links, suggested citation, and API example
- [Homepage](https://yjtoon.com/): browse by category, live three-format comparison, interactive voxel view of the catalog
- [TOON Studio](https://yjtoon.com/#studio): paste JSON and it is encoded in the browser by the same encoder the API uses; byte counts are exact, token counts are labelled estimates
## API endpoints (dynamic)
- [List categories](https://yjtoon.com/api/): all 27 categories with dataset counts
- [Category datasets](https://yjtoon.com/api/category/{slug}): datasets in a category (e.g. /api/category/programming-languages) plus a topic guide: overview, key ideas, pitfalls, agent tips and where to start
- [Single dataset](https://yjtoon.com/api/dataset/{slug}): full dataset content (e.g. /api/dataset/http-status-codes-complete) with a summary, use_w