Products

Data Lineage

The two derivation graphs the protocol records — between dataset versions and between contribution fingerprints — and what the Explorer shows.

Lineage in XnY Protocol is narrower than the word usually implies, and knowing its exact shape is the difference between a query that answers your question and one that returns an empty graph you misread.

There are two derivation graphs, they are separate, and both have exactly one kind of edge: derived_from.

Three questions, two lineage graphs

Between dataset versions. A version's metadata document can name the version it derives from. Walking those links upward gives the chain of versions a dataset came through; walking downward gives what was built on it.

Between contribution fingerprints. A CF can name a parent CF. A validation points at what it validates, so a CF's ancestry is also its review history. The registry accepts any (kind, parentCfId) pairing on chain and leaves per-kind parent policy to the frontier's own schema, so an AMENDS record can name what it amends the same way — whether it may do so on a given frontier is that frontier's rule, not the protocol's. The Explorer knows AMENDS as a kind but renders no view of an amendment's parent.

Nothing connects the two. A dataset version's lineage will not tell you which CFs went into it; that is membership, not derivation, and it is a different query.

The first two panels answer “where did this record come from?” Solid arrows point from a record to its parent, matching the API. The third answers “what is in this dataset?” Its dashed links illustrate membership from a separate lookup; they are not lineage edges.

03 · Which contributions are included? · separate lookup

includes

includes

Dataset B

Sample CF

Label CF

02 · Which contribution is this based on?

derived_from

derived_from

Validation CF

Label CF

Sample CF

01 · Which dataset did this derive from?

derived_from

derived_from

Dataset C

Dataset B

Dataset A

These are illustrative records. A dataset's ancestry does not imply membership of any CF shown in another panel, and a version number does not itself create a derivation edge.

Your questionRead
Which datasets are ancestors or descendants?Dataset lineage: GET /v1/datasets/{id}/lineage
What is this CF based on, or what builds on it?CF lineage: GET /v1/cfs/{cf_id}/lineage
Which CFs belong to this dataset version?Membership: GET /v1/datasets/{id}/cfs

See the dataset lineage, CF lineage, and membership references for parameters and pagination.

One edge type, not a typology

Both graphs return edges typed derived_from and nothing else. There is no selected_into, no validates, no supersedes. What kind of derivation an edge represents is read from the nodes — a CF's kind tells you whether it is an amendment or a validation — not from the edge.

An empty graph does not mean nothing is there

The dataset edge is not on chain. It is a field the publisher writes into the version's metadata document, which the indexer reads while fetching that document — so a version whose metadata has not been fetched yet returns a single-node graph even when an ancestor exists. An empty dataset lineage means nothing known yet, not nothing there; check the version's metadata status before concluding a dataset has no ancestry.

It is also optional in a second sense worth knowing: the field is not part of the published dataset-metadata schema, so a publisher who omits it produces a version with no recorded ancestry and no error. Absence of an edge is not evidence of absence of a parent.

The CF graph has neither caveat. A parent CF is a field on the record itself, written when the contribution is anchored.

Both graphs are bounded and say so. A traversal takes a depth, a direction (up, down or both) and a node limit, and a result that hit a limit is flagged as truncated rather than silently trimmed.

What the Explorer shows

The Explorer has a page per entity — frontier, campaign, task, CF, DID and dataset — and lineage appears on two of them.

A dataset version page carries seven sections: its details, the contribution fingerprints that are members of it, its contributors, a sample of example contributions, the access grants anchored against it, its recorded derivation relations, and its royalty activity. Those relations render as a flat list of child–parent rows rather than a drawn graph. The examples are a bounded sample of approved members, which is a display filter and not the membership rule — membership comes from the version's frozen snapshot and does not re-check on-chain verdicts.

A CF page shows its immediate relations on a single rail: its author, the label it certifies if it is a validation, the validation that certifies it if one exists, and each dataset version it is a member of. Note what that means in practice — a label's own parent sample is not on the rail, and the rail mixes membership in with derivation. The full parent walk is available from the API, not from the page.

What lineage does not track

A lineage graph is often expected to run all the way to commercial use — published here, adopted by that model, paid out to whoever holds the rights now. This one stops earlier:

Sometimes expectedReality
Publication tracked as a lineage stageA version's metadata document carries a publisher and a published-at timestamp. Nothing tracks a push to an external host, and no page shows one
Adoption by a model, as "proof of utility"No adoption record exists anywhere — not in the contracts, not in the indexer's schema, not in the Explorer
An edge from a payout back to the contributions that earned itLineage does not model payouts at all. Royalty goes to whoever holds the ERC-1155 shares, which contributors claim rather than buy — see Royalty Engine

For what a contribution fingerprint is and how it is anchored, see Contribution Fingerprint. For how versions are assembled from members, see Data Assembly.

References

Last updated

On this page