Dataset Development Standard
Covers dataset provenance, licensing, validation, integrity, and release.
Applies to: Collected, transformed, generated, preserved or released datasets and corpora, including training and evaluation sets.
Explanation reviewed 2026-10-01. Catalog identity and source review have separate dates.
Read public MarkdownPurpose and applicability
A dataset can be readable while its origin, allowed use or split method is unknown. DDS preserves those facts and keeps release blockers visible.
Collected, transformed, generated, preserved or released datasets and corpora, including training and evaluation sets.
How it works
Classify the dataset and link its sources. Record collection and transformation, filtering, licensing/usage limits, schema, validation and splits. Keep evaluation sets out of training where applicable. Resolve release licensing uncertainty before publication.
Outputs are the dataset manifest, provenance and license records, validation and split notes, hashes and limitations. Raw, working, candidate, released and archived readiness describe different stages.
In practice
The lineage below is synthetic and omits the canonical examples’ internal-use dataset. It shows a transformation and source-group split without borrowing example counts or declaring redistribution rights.
Editorial synthetic lineage based on the split-record fields.
Source: DDS/examples/Example-Split-Record.mdSynthetic source groups → UTF-8 normalization → empty-record filter
→ source-group split
├─ validation: reviewed source groups
└─ test: held-out source groups
Split purpose: evaluation only; no training split.
Leakage control: related records stay in the same group.
Seed/counts: must be recorded for the actual run.
Validation: pending.Illustrative provenance/license boundary, preserving the requirement for an actual source review.
Source: DDS/examples/Example-License-Record.mdProvenance: synthetic teaching fixture; no collected records.
Transformation: normalization and empty-record filtering (illustrative).
Source terms: not established for a real dataset.
Redistribution: not approved by this example.
Release blocker: actual source/rights review missing.
Known limitation: no real validation, counts or integrity evidence.Adopt one part
Start with a bounded surface or record. Complete the relevant adopter checks before extending the claim.
- Record one source and transformation, including removed records and generated-data labels.
- Record source terms and usage limits; leave public release blocked while rights are unresolved.
- Define validation and split/leakage controls, then compute real release hashes and retain findings.
Sources and limits
Canonical license and split examples are illustrative and include publication restrictions. They are not licenses for this page or a dataset release. No real source data, personal record, example count or hash is published here.
These are reviewed public explanations, not the normative specifications. Suite references are relative to the canonical collection; site/ references identify committed website sources and webserver/ references identify serving configuration. Illustrative examples demonstrate record shape; they do not establish compliance. Suite checks and adopter validation are separate.
DDS/DDS.manifest.tomlDDS/Adoption-Guide.mdDDS/Validation-Checklist.mdDDS/Dataset Development Standard.mdDDS/examples/Example-Provenance-Record.mdDDS/examples/Example-License-Record.mdDDS/examples/Example-Split-Record.md