Data Lineage
The two derivation graphs the protocol records — between dataset versions and between contribution fingerprints — and what the Explorer shows.
Lineage in XnY Protocol is narrower than the word usually implies, and knowing its exact shape is the difference between a query that answers your question and one that returns an empty graph you misread.
There are two derivation graphs, they are separate, and both have exactly one kind of edge: derived_from.
Three questions, two lineage graphs
Between dataset versions. A version's metadata document can name the version it derives from. Walking those links upward gives the chain of versions a dataset came through; walking downward gives what was built on it.
Between contribution fingerprints. A CF can name a parent CF. A validation points at what it validates, so a CF's ancestry is also its review history. The registry accepts any (kind, parentCfId) pairing on chain and leaves per-kind parent policy to the frontier's own schema, so an AMENDS record can name what it amends the same way — whether it may do so on a given frontier is that frontier's rule, not the protocol's. The Explorer knows AMENDS as a kind but renders no view of an amendment's parent.
Nothing connects the two. A dataset version's lineage will not tell you which CFs went into it; that is membership, not derivation, and it is a different query.
The first two panels answer “where did this record come from?” Solid arrows point from a record to its parent, matching the API. The third answers “what is in this dataset?” Its dashed links illustrate membership from a separate lookup; they are not lineage edges.
These are illustrative records. A dataset's ancestry does not imply membership of any CF shown in another panel, and a version number does not itself create a derivation edge.
| Your question | Read |
|---|---|
| Which datasets are ancestors or descendants? | Dataset lineage: GET /v1/datasets/{id}/lineage |
| What is this CF based on, or what builds on it? | CF lineage: GET /v1/cfs/{cf_id}/lineage |
| Which CFs belong to this dataset version? | Membership: GET /v1/datasets/{id}/cfs |
See the dataset lineage, CF lineage, and membership references for parameters and pagination.
One edge type, not a typology
Both graphs return edges typed derived_from and nothing else. There is no selected_into, no validates, no supersedes. What kind of derivation an edge represents is read from the nodes — a CF's kind tells you whether it is an amendment or a validation — not from the edge.
An empty graph does not mean nothing is there
The dataset edge is not on chain. It is a field the publisher writes into the version's metadata document, which the indexer reads while fetching that document — so a version whose metadata has not been fetched yet returns a single-node graph even when an ancestor exists. An empty dataset lineage means nothing known yet, not nothing there; check the version's metadata status before concluding a dataset has no ancestry.
It is also optional in a second sense worth knowing: the field is not part of the published dataset-metadata schema, so a publisher who omits it produces a version with no recorded ancestry and no error. Absence of an edge is not evidence of absence of a parent.
The CF graph has neither caveat. A parent CF is a field on the record itself, written when the contribution is anchored.
Both graphs are bounded and say so. A traversal takes a depth, a direction (up, down or both) and a node limit, and a result that hit a limit is flagged as truncated rather than silently trimmed.
What the Explorer shows
The Explorer has a page per entity — frontier, campaign, task, CF, DID and dataset — and lineage appears on two of them.
A dataset version page carries seven sections: its details, the contribution fingerprints that are members of it, its contributors, a sample of example contributions, the access grants anchored against it, its recorded derivation relations, and its royalty activity. Those relations render as a flat list of child–parent rows rather than a drawn graph. The examples are a bounded sample of approved members, which is a display filter and not the membership rule — membership comes from the version's frozen snapshot and does not re-check on-chain verdicts.
A CF page shows its immediate relations on a single rail: its author, the label it certifies if it is a validation, the validation that certifies it if one exists, and each dataset version it is a member of. Note what that means in practice — a label's own parent sample is not on the rail, and the rail mixes membership in with derivation. The full parent walk is available from the API, not from the page.
What lineage does not track
A lineage graph is often expected to run all the way to commercial use — published here, adopted by that model, paid out to whoever holds the rights now. This one stops earlier:
| Sometimes expected | Reality |
|---|---|
| Publication tracked as a lineage stage | A version's metadata document carries a publisher and a published-at timestamp. Nothing tracks a push to an external host, and no page shows one |
| Adoption by a model, as "proof of utility" | No adoption record exists anywhere — not in the contracts, not in the indexer's schema, not in the Explorer |
| An edge from a payout back to the contributions that earned it | Lineage does not model payouts at all. Royalty goes to whoever holds the ERC-1155 shares, which contributors claim rather than buy — see Royalty Engine |
For what a contribution fingerprint is and how it is anchored, see Contribution Fingerprint. For how versions are assembled from members, see Data Assembly.
References
Last updated