activev0.2.1Catalog reviewed 2026-09-23

Dataset Development Standard

Covers dataset provenance, licensing, validation, integrity, and release.

Applies to: Collected, transformed, generated, preserved or released datasets and corpora, including training and evaluation sets.

Source maturity: candidate

Explanation reviewed 2026-10-01. Catalog identity and source review have separate dates.

Read public Markdown

Purpose and applicability

A dataset can be readable while its origin, allowed use or split method is unknown. DDS preserves those facts and keeps release blockers visible.

Collected, transformed, generated, preserved or released datasets and corpora, including training and evaluation sets.

How it works

Classify the dataset and link its sources. Record collection and transformation, filtering, licensing/usage limits, schema, validation and splits. Keep evaluation sets out of training where applicable. Resolve release licensing uncertainty before publication.

Outputs are the dataset manifest, provenance and license records, validation and split notes, hashes and limitations. Raw, working, candidate, released and archived readiness describe different stages.

In practice

The lineage below is synthetic and omits the canonical examples’ internal-use dataset. It shows a transformation and source-group split without borrowing example counts or declaring redistribution rights.

Source-to-split lineageillustrative · teaching example, not a verification result

Editorial synthetic lineage based on the split-record fields.

Source: DDS/examples/Example-Split-Record.md
Synthetic source groups → UTF-8 normalization → empty-record filter
  → source-group split
     ├─ validation: reviewed source groups
     └─ test: held-out source groups

Split purpose: evaluation only; no training split.
Leakage control: related records stay in the same group.
Seed/counts: must be recorded for the actual run.
Validation: pending.
Inspect Source-to-split lineage
Provenance and license specimenillustrative · teaching example, not a verification result

Illustrative provenance/license boundary, preserving the requirement for an actual source review.

Source: DDS/examples/Example-License-Record.md
Provenance: synthetic teaching fixture; no collected records.
Transformation: normalization and empty-record filtering (illustrative).
Source terms: not established for a real dataset.
Redistribution: not approved by this example.
Release blocker: actual source/rights review missing.
Known limitation: no real validation, counts or integrity evidence.
Inspect Provenance and license specimen

Adopt one part

Start with a bounded surface or record. Complete the relevant adopter checks before extending the claim.

  1. Record one source and transformation, including removed records and generated-data labels.
  2. Record source terms and usage limits; leave public release blocked while rights are unresolved.
  3. Define validation and split/leakage controls, then compute real release hashes and retain findings.

Sources and limits

Canonical license and split examples are illustrative and include publication restrictions. They are not licenses for this page or a dataset release. No real source data, personal record, example count or hash is published here.

These are reviewed public explanations, not the normative specifications. Suite references are relative to the canonical collection; site/ references identify committed website sources and webserver/ references identify serving configuration. Illustrative examples demonstrate record shape; they do not establish compliance. Suite checks and adopter validation are separate.

  • DDS/DDS.manifest.toml
  • DDS/Adoption-Guide.md
  • DDS/Validation-Checklist.md
  • DDS/Dataset Development Standard.md
  • DDS/examples/Example-Provenance-Record.md
  • DDS/examples/Example-License-Record.md
  • DDS/examples/Example-Split-Record.md
Reviewed manifest facts