Data Assembly
How a set of contribution fingerprints becomes a versioned, ownable dataset, and why every assembly produces a new identifier rather than mutating an old one.
Assembly is the step that turns a chosen set of contribution fingerprints into a dataset a consumer can be granted access to and a contributor can own a fraction of. It is a single on-chain call, and no frontier owner has to approve it — there is no manifest-ownership table to consult and no coupling to the frontier registry. What the call does check is that the caller controls the DID it assembles under, or is an operator on an owner-managed allowlist that may assemble on any DID's behalf.
What the call anchors is a commitment, not the data. The contributions, the list naming them, and the payload itself all stay off-chain; the chain holds hashes and a Merkle root over them.
What a version records
A dataset version record holds seven fields and no more: the DID that assembled it, the manifest id it was assembled under, its version number, the contributors Merkle root, the total shares it may ever issue, a metadata hash, and the URI where that metadata can be read.
Everything a reader usually wants to know about the dataset — how many fingerprints are in it, what encryption the payload uses, which frontier it claims to belong to — lives in that metadata document rather than in the version record. Which fingerprints are in it is one hop further out again: the document carries a URI and a hash for a separate CF list, not the list itself.
A version is an identity, not a revision
The identifier is derived, not assigned: datasetId = keccak256(assemblerDidId, manifestId, versionNumber).
So every version is its own datasetId. There is no separate dataset object that versions hang off, and nothing on-chain marks one version as superseding another. A series is just the pair (assemblerDidId, manifestId), and the contract keeps a counter per pair — the next assemble under the same pair gets the next version number, and therefore a different identifier.
manifestId is a bytes32 the caller chooses. Nothing derives it from the contributions, the payload or the metadata, and nothing reserves it: the same manifestId picked by two different assemblers produces two independent series whose identifiers can never collide, because the assembling DID is hashed in. That is why there is no manifest-ownership table to consult, and why another DID picking your manifestId gets a series of its own rather than a version inside yours.
That guarantee is against other DIDs, not against the platform. The caller check passes for an allowlisted operator under any DID, so an operator can assemble under a DID it does not control — and because the version number is the series counter incremented, that assemble extends the existing series rather than starting a new one. Whoever the owner puts on that allowlist can append to any series on the chain.
Predicting the version you are about to get
The version number is assigned by the contract, which is awkward for a publisher: the identifier is an input to work that has to happen before the transaction. Preparing the payload, uploading it, building the contributors Merkle root — all of it is keyed to the datasetId, and the Merkle leaves are keyed to it too, so they cannot be built after the fact.
The publish flow therefore reads the series counter, predicts the version number, derives the datasetId from the prediction, and does its work against that. It then passes the prediction back into the assemble call as expectedVersion. If a concurrent assemble for the same series advanced the counter in between, the assigned version no longer matches and the call reverts — which rolls the counter increment back, so the race strands no half-published version and anchors no metadata bound to an identifier that turned out to be someone else's.
Passing zero opts out of the check. Version numbers are 1-based, so zero is never a valid prediction.
The contributors Merkle root
The root is the whole ownership mechanism, and it has two layers.
Layer 1 is per contributor: a tree over that contributor's own fingerprint ids, sorted, producing one root per contributor. Layer 2 is the tree the chain stores: one leaf per contributor, committing to the DID, the share amount it is owed, that contributor's layer-1 root, the datasetId, the chain id, and the address of the ownership contract.
The example below has two contributors. Each has a separate CF tree; their allocations then become leaves in the shared ownership tree. The share amount is part of the assembler's allocation, not a count inferred from the number of CFs.
Scope binds each allocation to the datasetId, chain ID and ownership contract address. A share claim proves one allocation against the layer-2 root; it does not submit every CF individually.
Claiming shares means proving membership in layer 2. The layer-1 root travelling inside the leaf is what binds those shares to a specific set of fingerprints rather than to a bare number, so an allocation can be audited against the contributions it was paid for.
The chain id and contract address in the leaf are there so a root cannot be replayed: the same CF list committed on another chain, or against another deployment, yields different leaves and every proof against it fails.
The claimed frontier is verified, not trusted
Datasets are frontier-agnostic on-chain — the assemble event deliberately does not carry a frontierId, because a dataset may draw contributions from more than one frontier.
The metadata document does name a primary frontier, and it is recorded as a claim. An off-chain enricher walks every fingerprint in the CF list, resolves each one's task and follows task → campaign → frontier on-chain, and checks the claimed frontier against the set it derives. The result is tracked with a count of how many fingerprints have been checked so far, and a dataset is only marked verified when that count reaches the whole list. A claim the derived set does not contain is marked failed.
Consumers should read the derivation result, not the claim in the document.
What this page does not describe
Earlier versions of this page described an intermediate entity between fingerprints and datasets — a Data Asset, "the minimal commercial unit", where ownership and licensing were said to be enforced. It also described lineage as a typed graph with derived-from, selected-into, validates and supersedes edges carrying the manifest rule that caused each inclusion; queries over it for why a fingerprint was excluded and what a proposed amendment would affect; diffs between two versions; re-assembly triggered when a fingerprint is amended or deprecated; and payout reserves held back during disputes.
None of it exists. There is no asset entity in the contracts, and none in the indexer's schema either — thirty-one migrations, no such table. Ownership is an ERC-1155 balance with one token id per dataset version, and a grant names a datasetId; both operate on the version, and there is nothing in between for them to operate on instead. Membership is two flat tables, fingerprint-to-dataset and contributor-to-dataset, with no edge types and no graph. Nothing computes a diff between versions, nothing reacts to an amended fingerprint by re-assembling, no version can be marked deprecated, and the royalty engine has no reserve mechanism to hold anything back with.
The determinism claim was wrong in a more specific way, and worth naming separately: identifiers are keyed on the assembler, the manifest id and a counter — never on content. Re-running the same selection does not reproduce a version. It produces the next one.
Interfaces
- In: a CF list, a share allocation committed as the contributors Merkle root, and a metadata document — see Contribution Fingerprint for what a fingerprint is and Storage, Compute and Serving for where the payload and the document are stored.
- Out: a
datasetIdand version number, consumed by Tokenized Ownership Proofs for share claims and by Access Control for grants. - Identity: the assembling DID is authorised through the DID registry, or the caller is an allowlisted platform operator — see Identity.
Invariants
- The identifier is derived from its inputs.
keccak256(assemblerDidId, manifestId, versionNumber)— reachable by anyone holding the three, and not assignable. - Version numbers are contract-assigned and monotonic per series. A caller cannot choose one, and cannot skip one.
- A version never changes. The record is written once; new information is a new version, which is a new identifier.
- Total shares is a ceiling fixed at assembly. Claims are refused past it, and the royalty engine divides distributions by it — see Royalty Engine.
- A root is bound to one chain and one contract. Replaying it elsewhere invalidates every proof.
Last updated