Contribution Fingerprint (CF)
The on-chain record binding one atomic contribution to its contributor, its validator and a commitment to its content — and what it does not carry.
A Contribution Fingerprint is one on-chain record per atomic contribution. It is the protocol's smallest unit of authorship: everything downstream — dataset membership, ownership shares, royalty — resolves back to a CF.
It is a deliberately thin record. Nine fields:
| Field | What it holds |
|---|---|
cfId | Derived, not chosen — see below |
parentCfId | The CF this one descends from, or zero |
contributorDidId | Whose contribution it is — reassignable |
verifiedByDidId | The validator behind the current verdict |
metadataHash | A commitment to the metadata document |
metadataUri | Where that document can be fetched |
kind | SAMPLE, LABEL, AMENDS or VALIDATION |
verdict | PENDING, APPROVED or REJECTED |
qualityGrade | NONE, D, C, B, A or S |
The task a CF was submitted against is not one of the nine. It is held in a separate public mapping from CF to task, and the frontier above it is reached off-chain by following that task to its campaign. The task matters twice below — it is an input to the id, and it is where the acceptance rule a task declares lives, which is not the same as the rule a verdict was reached by — so it is worth knowing it is adjacent to the record rather than inside it.
The id is derived
cfId is not assigned. It is keccak256 over the metadata hash, the contributor's DID and the task — plus the parent CF when there is one. So the same contribution, submitted the same way, always lands on the same id.
Note which fields are not in that formula, because it decides what counts as a duplicate. kind is absent: the same document, contributor, task and parent submitted once as a SAMPLE and once as a LABEL derives one id, and the second submission is rejected as already existing rather than anchored beside the first. The id is the tuple, and the tuple does not include what you called it.
metadataUri is deliberately absent from that formula, and the reason is the interesting part: a CF is identified by its content commitment, not by where the content happens to live. That is what makes the location repairable. Where no commitment was made — metadataHash zero, which the registry allows — the id identifies the contributor, task and parent and nothing about the content at all. updateCFMetadataUri lets the CF's own contributor or an enabled validator point the record at a corrected address without changing the record's identity.
The same asymmetry runs through the rest of the record. Where metadataHash is non-zero it is a commitment — the document either hashes to it or it does not, and anyone can check. It may also be zero, which the registry accepts and which means "no metadata committed"; readers are expected to fail closed on that, since a body fetched without a commitment is a body nothing vouches for. metadataUri is attested by nothing either; on the batch submission path anyone relaying a signed set may substitute it, which is why that path caps its length rather than trusting it.
The line the protocol draws is between resolution and integrity: a wrong metadataUri can stop you finding the document, and can never make a wrong document pass. It tells you where to look without vouching for the address you are looking at.
State the resulting guarantee carefully, because it is weaker than "proof of authorship" in three separate ways. Content integrity is conditional — it holds where metadataHash is non-zero and there is nothing to check where it is zero. Attribution is attested, not proved: a validator signs that this DID contributed, and contributorDidId is a field a later transfer can re-point. And anchoring is open to relayers — the contract checks the validator's signature, never who submitted the transaction. What a CF gives you is a validator's attestation, tied to a content commitment when one was made.
The record has no time field
There are two timestamps available and they are not equivalent. The block that anchored a CF gives a time nobody chose, and it is always there. The metadata document carries a published_at its schema requires, which the publisher wrote — an assertion by the submitter rather than an observation by the chain. Whether it is checkable depends on the commitment: with a non-zero metadataHash it is fixed and anyone can verify it, and with a zero one there is nothing to verify it against.
Use the block when you need a time you can trust, and published_at when you need the time the submitter meant.
Why a CF is credible
Three of the nine fields carry that, and they arrive together from a validator rather than from the contributor:
verdict— whether the contribution was admitted, rejected, or is still pending.qualityGrade— five grades,DthroughS, plusNONEfor ungraded.verifiedByDidId— which validator stood behind those two, as a DID, so the assessment is attributable.
Every anchor requires a validator's EIP-712 signature. There is no unvalidated submission path: the bare submitCF that existed in v1 was removed, and the two entrypoints that remain — one per CF, one batched — both take a signature. That signature binds almost the whole record: the validator's DID, the contributor's, the task, the kind, the verdict, the grade, a deadline, and a hash of the content identity — metadataHash together with parentCfId. The one thing it does not bind is metadataUri, which is why a validator can sign before the metadata is even published and the location can be repaired afterwards. A verdict attests to what was contributed, never to where it is stored.
Read the guarantee narrowly, because two things are not in it. The validator set is an allowlist the contract owner curates, and nothing requires a validator's DID to differ from the contributor's — so a contributor who is also an enabled validator can admit their own work. And nothing constrains how a grade is chosen: the protocol defines the range and records what the signer picked, while the policy behind the choice is the validator's, not the chain's. Attributable is the property on offer here; independent is not.
The signature also covers only the moment of anchoring. validateCF can re-set all three of these fields later, and takes no signature at all — see below.
What a task declares about deciding a verdict lives in the task, not in the CF. A task's metadata carries an acceptance rule, and there are four shapes: consensus with a minimum agreement count, gold_standard with a match threshold, auto_accept, and agent_qa_with_spot — agent review with a spot-check percentage and a sidecar task for the checks.
Read that as a stated intention, not as a rule the protocol applies. Nothing in the protocol evaluates it: a validator supplies the verdict and grade directly at anchoring, validateCF can re-set them afterwards without reference to any task, and the indexer checks only that an acceptance rule is present — no service reads a min_agreement or a match threshold and counts against it. So the rule tells you what a task said it would do, and the CF tells you what a validator did.
Corrections add records; some fields also change in place
The history is append-only: an AMENDS or VALIDATION record is a new CF naming the earlier one as its parent, and nothing removes or supersedes what came before.
The parent link is a claim, not a checked reference. The registry has a CFNotFound and a TaskNotFound, and no parent equivalent — a non-zero parentCfId is fed into the id derivation without ever being looked up, so a CF naming a parent that does not exist anchors successfully. Read an ancestry chain as something to resolve rather than something the chain guaranteed.
The record itself is not immutable, and it is worth knowing exactly which parts move. Three functions write into an existing CF:
validateCFre-setsverdict,qualityGradeandverifiedByDidId. This is deliberate — the registry states that re-scoring before the anchor is unsupported and that a verdict change afterwards goes through this function. It takes no signature. The gate is that the caller owns an enabled validator's DID, so any enabled validator can overwrite what a previous one attested, and the record keeps no trace of which signature authorised the values it now holds.transferCFOwnershipre-setscontributorDidId. ThecfIddoes not change with it, because the id was derived from the original contributor.updateCFMetadataUrire-setsmetadataUri, as above.
So five of the nine fields can change after anchoring. What cannot change is cfId, parentCfId, metadataHash and kind — the identity and the content commitment. Read the record accordingly: whatever content it committed to is fixed, the assessment of that content is the current value of a mutable field.
The registry accepts any (kind, parentCfId) pairing on chain and leaves per-kind parent policy to the frontier's own schema, so which kinds may name a parent, and must, is a frontier's rule rather than the protocol's.
What a CF does not carry
| Sometimes assumed | Reality |
|---|---|
| Structured evidence artifacts — transaction ids, document hashes, signed credentials | The schema declares no evidence field and nothing reads one. It declares the submission and task ids, the contributor, an uploader, the payload location and size, and either inline data or an external {uri, sha256, provider}, alongside the usual document header. It is an open schema, so a producer can attach an extra property and have it hashed and anchored — but as an unverified annotation nothing validates or consumes |
| A source class and a validation class ("heuristic", "peer-reviewed", "staked verification", "adjudicated") | No such classification exists. verdict and qualityGrade are what the protocol records |
| Reviewer reputation scores feeding the assessment | A campaign or task may state a minimum contributor reputation score in its metadata, but no contract and no protocol service computes or enforces it — the indexer stores that whole qualification block as opaque JSON. The schema names a second reader outside this repository, so whether a platform enforces it is not answerable from here |
| A challenge window, or a dispute that deprecates a CF | No challenge or dispute mechanism exists in the contracts, the indexer or the executor. A correction is an AMENDS record, which adds rather than retracts |
| Metered usage flowing back to the CFs that earned it | No consumption is metered and no usage event feeds royalty. Royalty is distributed per dataset version against the share total it was assembled with — see Royalty Engine. The storage gateway does enforce per-caller upload quotas, but those are operational limits on writes, not billing, and they touch no CF — see Storage & Serving |
| An asset layer between a CF and a dataset | There is none. CFs are members of a dataset version directly — see Data Assembly |
Be careful with the boundary here, and in both directions. More of the review pipeline is recorded on chain than it looks: a task declares whether it expects human or agent execution, and that is an on-chain enum emitted with the task, as is whether it is a submission task or a validation task. But recorded is not enforced. No contract checks that an AGENT task's work came from an agent, and a task's duplicate_permission — the maximum submissions one contributor may make — is a metadata field that nothing counts against; the registry has no per-contributor submission counter to enforce it with.
The registry does refuse a second submission of the same derived id, and that is worth keeping separate: it is uniqueness of the tuple, not a submission quota. Two genuinely different contributions under the same task both anchor, whatever duplicate_permission says.
What is not the protocol's is the pipeline around a submission: how a format check is implemented, what a deduplication heuristic looks at, how an agent is prompted, how a human reviewer is queued. What the protocol fixes is narrower than a task's metadata makes it sound — which validator may authorise an anchor, that identical submissions collapse to one id, and that every anchor carries a validator's DID. Not who may anchor: both submission entrypoints are open to any relayer holding a valid signature. Who executed, how many duplicates are tolerated and how a verdict was reached are declared by a task and left to whoever runs it.
What CFs do enable
- Attributable authorship. One CF per atomic contribution, naming a DID — the DID holding it now, which a transfer can change.
- A derivation history, to the extent its links resolve. A record may name a parent, and nothing verified that the parent exists. Walking up is reading a field; walking down means finding the records that name this one, which the indexer does with a join — the chain itself only ever points backwards. See Data Lineage.
- Ownership. A dataset version commits to what each contributor is owed; contributors claim ERC-1155 shares against that commitment.
Note what is not on that list: confidentiality is not a CF-level property. The metadata document describes the payload as raw — the inline form is base64, which is an encoding, not encryption.
Confidentiality is designed to enter a layer up, at the dataset version — one key per version, wrapped per grant — but that model is not implemented, and no production service performs it today. Treat it as target architecture; the storage design is where its status is tracked. What does exist on chain now is the grant, and note the layer it names: a dataset version and a subject DID, never a CF.
So a CF is a commitment and a pointer. Whatever protects the payload behind it is the store it sits in, not the record.
Example: from contribution to royalties
Dr. Lin contributes a label for a medical image provided by the patient from a recent pathology diagnosis. The image creates one CF for the sample; Dr. Lin's label creates a second CF naming the first as its parent. Each was admitted by a validator whose DID is on the record, and each is anchored in a block. Independent reviewers add their own CFs to confirm or amend the label — a correction is a new CF linked to the original, never an edit.
These CFs become members of a versioned dataset, which fixes how many ownership shares it may issue and commits to what each contributor is owed. The patient and Dr. Lin each claim their shares against that commitment, and the shares mint to their DID-bound accounts. When an AI builder pays for the dataset, the Royalty Engine divides that payment across the shares — and because it divides by the total the version was assembled with, a contributor who claims months later still collects their part of everything paid in since.
Invariants
- Append-only history. One CF per atomic contribution, and later work is a new CF naming the earlier one. Nothing is removed or superseded — though five fields on an existing record can be re-set, as above.
- Fixed identity.
cfId,parentCfId,metadataHashandkindnever change. The id iskeccak256of the submission's inputs, so it is reproducible from those inputs by either path — but not from the record's current fields, since a transfer changescontributorDidIdand leaves the id derived from the original. - Content-committed where a commitment was made, location-hinted always. A non-zero
metadataHashis verifiable; a zero one is the registry's accepted way of saying nothing was committed, and then there is no content to verify and no content in the id either — two different uncommitted payloads under the same contributor, task and parent derive the samecfId, so the second is refused as a duplicate of a submission it does not match.metadataUriis attested by nothing and may be repaired. - Validator-gated. No CF is anchored without an enabled validator's signature, and the verdict fields are only ever written by an enabled validator — at anchor under signature, afterwards without one.
Reproducible identity stops at the CF, and it is worth knowing where. A dataset version's id is derived too, but from the assembler, the manifest id and a version counter — never from the members. Assembling the same CFs twice therefore yields two different versions. Data Assembly has that mechanism and what follows from it.
Last updated