Quick Guide

Fundamentals

The data model: samples, labels, amendments and validations, the fingerprint that records each one, and the dataset version that makes a set of them ownable.

Four kinds of work go into the protocol, and each one is recorded the same way.

Sample — a raw observation a model will learn from: an image, an audio clip, a text span, a time-series window, a multi-sensor frame.

Label — a structured interpretation of a sample, or of another label: a class, a bounding box, a segmentation mask, a span, a rating, a relation, events over time.

Amendment — a correction. Nothing is ever edited in place, so a fix is a new piece of work pointing back at what it corrects.

Validation — a quality judgment on an earlier piece of work: a verdict, and a grade.

These four are the protocol's SAMPLE, LABEL, AMENDS and VALIDATION — the complete set, defined on-chain and not extensible off it.

Why quality matters to models

  • Signal-to-noise — mislabelled or low-information data reduces effective batch size and slows convergence.
  • Bias and leakage — inconsistent schema, shortcut features, or label leakage harm generalisation and fairness.
  • Heterogeneous tasks — multi-task and chain-of-thought models depend on clear, consistent instructions and traceable provenance to debug and improve.

Better samples, clearer labels and verifiable validations mean more useful gradient steps and fewer surprises in production. That is the argument for recording provenance at all, and it is why the unit of record is one piece of work rather than one file.

The data model

recorded as

task reference

chosen membership

claim allocation

owner anchors

Sample · Label
Amendment · Validation

Contribution fingerprint
one atomic record

Task → Campaign → Frontier
where the work belongs

Dataset version
a committed set of contributions

Ownership shares
ERC-1155

Access grant
per dataset version

One atomic contribution, one fingerprint. A Contribution Fingerprint (CF) is the on-chain record of a single piece of work: a hash of what was contributed, who contributed it, which task it was done for, and the validator's assessment of it. Fingerprints are what make contributions discoverable and auditable, and the fingerprint id is derived from those first three together — so resubmitting the same content as the same contributor under the same task is rejected as a duplicate.

A fingerprint is never revised. A correction or a review is a new fingerprint referencing the earlier one, so the history is append-only by construction.

A dataset version is a chosen set of fingerprints. Assembly anchors a commitment to that set, not the set itself: what goes on-chain is a Merkle root over the share allocation and a hash of a metadata document, and the ordered list of fingerprint ids stays off-chain behind a URI in that document. That version is the unit of ownership and the unit of access: ownership fractions are an ERC-1155 balance under one token id per version, and a grant of access names one version.

There is no layer between a fingerprint and a dataset version. The same fingerprint can appear in any number of versions, and a version can draw fingerprints from more than one frontier.

What the model lets you express

One sample, several interpretations

A sample is fingerprinted once by convention, not by protocol rule. A fingerprint id is derived from the content hash together with the contributor and the task, so the same bytes submitted by a second contributor, or by the same contributor under a second task, is a distinct fingerprint the contract accepts. Uniqueness is per content-contributor-task; one fingerprint per sample is a workflow the tooling keeps to.

Any number of label fingerprints can then be made against a sample, under different tasks and by different contributors.

Because a dataset version is just a selected set of fingerprints, two versions can share that sample and take different labels: one assembles the sample with the labels from one task, another with the labels from a second. Each version has its own share allocation, its own owners and its own access grants. The same raw sample can power different products without being contributed twice.

Labels on labels

A label fingerprint may reference another label rather than a sample. That is what lets you annotate interpretations — rubrics, explanations, confidence judgments, evaluator notes — with the chain of references preserved.

What the protocol does not do is route earnings along that chain. Royalties are not propagated from a child fingerprint up to its parents. A dataset version's revenue is divided by the share allocation fixed at assembly, and whether an upstream labeller appears in that allocation is a decision the assembler makes when building the list — not something the parent reference causes on its own.

Corrections without rewriting

An amendment is a fingerprint whose parent is the thing it corrects. The original stays exactly as it was, and both are visible.

Which kinds may reference which parents is not enforced on-chain — the contract accepts any combination. That policy lives with the frontier and is applied off-chain, which is worth knowing before assuming a parent link means what you expect it to.

Why put any of this on a chain

  • Provenance — a fingerprint fixes who did what, when, and to which payload. The content hash is what is attested; where the content is stored is a repairable hint.
  • Ownership — fractions are held against a dataset version, so a set of contributors and validators can share in what it earns rather than a single vendor owning the files.
  • Attribution that resolves — a fingerprint names its task, and the contract rejects a task that does not exist, so CF → Task → Campaign → Frontier resolves on-chain rather than by convention.
  • Append-only history — corrections and reviews add records instead of replacing them, so the record of a dataset cannot be quietly rewritten after the fact.

Payload confidentiality is an off-chain concern, and it is designed rather than running: one key per dataset version, wrapped per grant and held custodially, with no deployed service releasing it — see Access Control. Keeping data confidential while it is being computed on is a separate problem the protocol does not address at all — see Future Directions.

Where to go next

Last updated

On this page