Using the API
Base URL, identifiers, enum ordinals, cursors, consistency, proof verification and error codes — the conventions an OpenAPI document cannot state.
Read access to XnY Protocol lineage: datasets, contribution fragments (CFs), DIDs, frontiers, campaigns, tasks, and the Merkle proofs that back share claims.
Everything here is a projection of public Base chain state. There is no private data behind this API and no write path — you cannot create, modify, or claim anything through it.
The API Reference has one page per endpoint, generated from the service's own OpenAPI document: paths, parameters, response shapes. This page is the part a schema cannot express — how an identifier is written, what an integer enum means, how a cursor behaves, and how long after a write a read reflects it. Read it before your first call.
Audience: developers reading lineage from outside XnY.
Access
Base URL
The public API is served on the Explorer's host under the /api/v1 prefix, by a
service that exists only for external callers. Requests to xnyscan.com are
redirected here; use the address above directly.
Authentication: none. Requests are anonymous. No API keys are issued and none are accepted.
Methods: GET and OPTIONS. Any other verb is a 405 with
METHOD_NOT_ALLOWED and an Allow header — the surface is read-only by
construction, so a write never reaches the data layer at all.
CORS: Access-Control-Allow-Origin: *, methods GET, OPTIONS, headers
Accept, preflight cached for 24 hours. Browser calls from any origin work. The
wildcard is correct rather than lax: the surface is public chain state and no
cookie or credential is ever in play.
Rate limit: 120 requests per minute, per client, fixed window. Exceeding it
returns 429 with a Retry-After header.
The client is identified by the address the load balancer observed, taken from
X-Forwarded-For counting from the right — not by the first entry in that
header. Entries you send arrive to the left of it and do not change which bucket
you are counted against, so there is no way to widen your own budget by setting
the header, and no reason to try.
One caveat can make your effective budget differ from 120: the counter lives in the serving process's memory, so it is per pod rather than per cluster. One replica is deployed today, but that is not a guarantee.
Treat 120/min as a fair-use ceiling to stay under, not as a quota you have been granted.
Query parameters are snake_case, and unknown ones are ignored
Response fields are camelCase throughout. Query parameters are not: the
filters are campaign_id, frontier_id, owner_did, task_type, status,
token, from, to. The exceptions are the graph reads, where depth and
direction are single words and nodeLimit is camelCase.
This matters more than an inconsistency usually would, because an unrecognised
query parameter is silently ignored rather than rejected. Asking for
?campaignId=… on a list endpoint does not 400 — it returns the unfiltered
page, which looks like a filter that matched everything. Check that a filtered
response is actually narrower than the unfiltered one when you first wire a
filter up.
Identifiers
bytes32 ids
datasetId, cfId, frontierId, campaignId, taskId, all Merkle roots and
proof elements: 0x-prefixed with exactly 64 hex digits. Responses use lowercase; requests
also accept uppercase hexadecimal digits and an 0X prefix. A missing 0x prefix is a 400, not a guess:
DID ids: one value, two renderings
An XnY DID id is an on-chain uint128. The canonical DID string is
did:xny:<uuid>, where the UUID is that 128-bit number rendered as a UUID. Both
forms are the same value and conversion is lossless.
In requests, any of these four forms is accepted wherever {did} appears:
In responses, didId / ownerDidId / contributorDidId /
verifiedByDidId render as UUID — with one deliberate exception.
The proof endpoint's decimal contributorDidId
GET /v1/datasets/{id}/cf/{cfId}/proof returns contributorDidId as a
decimal string, not a UUID.
This is not an inconsistency to work around. The Merkle leaf is built over the
numeric uint128, and this endpoint's entire purpose is to hand you the exact
values that hash to the anchored root. Rendering it as a UUID here would mean
every caller converting it back before hashing, with a silent wrong-proof
failure for anyone who forgot. The decimal is the hashing input, so the endpoint
returns the hashing input.
Converting between the two is a base change over the same 16 bytes:
Large integers are strings
Share counts and token amounts are decimal JSON strings, as are
chainId and the numeric DID in a proof response:
These values can exceed IEEE-754's exact integer range, so a JSON parser that maps numbers
to doubles corrupts them silently — no exception, just a wrong value in the low
digits. This matters most for contributorDidId in decimal form, which is a
128-bit value that cannot safely be represented as a JavaScript Number.
In JavaScript use BigInt(value). In Python int(value) is fine. Do not let a
generic parser coerce these fields.
Block numbers, versionNumber, cfCount, limit, logIndex and leaf indices
are returned as JSON numbers. That encoding alone does not guarantee JavaScript
integer precision: use a lossless JSON parser when a numeric field can exceed
Number.MAX_SAFE_INTEGER; converting an already-rounded number to BigInt
cannot recover the original value.
Quickstart
Five requests, no credentials, real ids from the live host. They walk one lineage upward — a contribution to the frontier it belongs to — because that is the direction that always resolves: every id you need is a field on the record in front of you.
Downward is the direction that needs care, and step 5 shows why.
1. Read a contribution fingerprint
A CF is one atomic piece of contributed work, and the bottom of the hierarchy.
It carries its whole ancestry denormalised: taskId, campaignId and
frontierId. The next three requests each use one of them. kind: 0 is
SAMPLE, verdict: 1 is APPROVED and qualityGrade: 2 is C — on-chain
ordinals, not labels, tabulated under Enum values.
2. Read its task
Take taskId from the response above.
taskType: 0 is SUBMISSION and execution: 0 is HUMAN. Its campaignId
matches the one the CF carried, which is the check worth making at each hop.
3. Read its campaign
Take campaignId from the task.
4. Read its frontier
Take frontierId from the campaign. A frontier is a data domain — the top of
the hierarchy, and there is one on this host today.
The lineage closes: the frontierId here is the one the CF named in step 1,
four records down.
5. Now go the other way, and mind the filter
Downward, you have to ask for children and they may not exist yet. Pass the parent's id as a filter — without one this returns every campaign on the host, which is the trap the note above describes: a response that looks filtered because everything in it happened to match.
No response is reproduced here, and that is deliberate: a list is ordered
newest-first, so its first page changes whenever anyone creates a campaign —
one appeared between drafting this page and publishing it. What is stable is
the shape. Every item carries the frontierId you passed, which is how you
confirm the filter took effect rather than trusting that it did, and
pagination.hasMore tells you whether there is more behind it.
Two things to expect when you walk down from here:
- Newest first, so a small
limitis not a shortcut to a known record. The campaign steps 1 to 4 walked through is old enough that it is not on the first page of this listing. Reaching a specific record downward means paging to it. - Recent records are often empty. A newly created campaign usually has a task or two and no contributions against them yet, so a walk down to a CF has to keep going until it finds one. A walk up cannot dead-end that way, which is why the steps above are ordered as they are.
Enum values
Integer enums come straight from the protocol's Solidity types. They are ordinals, not labels.
| Field | 0 | 1 | 2 | 3 | 4 | 5 |
|---|---|---|---|---|---|---|
kind (AtomicKind) | SAMPLE | LABEL | AMENDS | VALIDATION | ||
verdict (Verdict) | PENDING | APPROVED | REJECTED | |||
qualityGrade (QualityGrade) | NONE | D | C | B | A | S |
taskType (TaskType) | SUBMISSION | VALIDATION | ||||
execution (Execution) | HUMAN | AGENT |
Note kind in particular: 0 is SAMPLE and 3 is VALIDATION. A CF with
kind: 0, verdict: 1 is an approved sample, not an approved label.
Some fields arrive pre-labelled as strings instead, and they come from different vocabularies. Do not map any of these through the table above:
| Field | Values | Origin |
|---|---|---|
status on a dataset | ACTIVE, REVOKED, ARCHIVED | The off-chain metadata document, validated on ingest against the published dataset-metadata schema. Passed through verbatim |
status on a grant | ACTIVE, REVOKED, UNKNOWN | On-chain grant status ordinal, labelled server-side. UNKNOWN means an ordinal this build does not recognise |
metadataStatus | PENDING, IN_PROGRESS, DONE, FAILED | Enrichment job state — see Eventual consistency |
proofStatus on claimability | pending, ready — lowercase, unlike every other status field | Whether the proof columns are populated |
A dataset's status is therefore a claim its publisher made off-chain, while a
grant's status is on-chain fact. They are not interchangeable despite sharing
two labels. The document schemas are under Metadata documents.
Pagination
Most list endpoints return an items array beside a pagination object:
Four endpoints do not use that exact envelope, so do not assume items is
always present:
| Endpoint | Shape |
|---|---|
/v1/datasets/{id}/holders | items and pagination, plus top-level datasetId and totalSupply |
/v1/datasets/{id}/royalty-flow | Two arrays — distributions and claims — plus pagination |
/v1/datasets/{id}/versions | versions array, plus assemblerDidId and manifestId. Not paginated |
/v1/datasets/{id}/frontier | Verification result. Not a list |
The cursor rules below apply wherever pagination appears.
limit— default 50, maximum 200. A larger value is silently clamped rather than rejected, so check thelimitechoed inpaginationif it matters. A non-integer or a value below 1 is a400. Both bounds are server-configured, so read them from the response rather than hardcoding them.nextCursor— present only whenhasMoreis true. Pass it back verbatim as?cursor=....- The cursor is opaque. It happens to be base64url JSON today; that is an implementation detail, not a contract. Do not parse it, construct it, or persist it as a bookmark across deploys.
GET /v1/datasets/{id}/lineage and GET /v1/cfs/{cf_id}/lineage are graph
reads rather than pages, and take different parameters:
depth— default 2, maximum 5.direction—up,down, orboth. Empty is treated asboth; any other value is a400.nodeLimit— default 50, maximum 100.
Unlike limit, these are rejected with a 400 when over the cap rather
than clamped. The response carries a truncated flag when the graph was cut off
by nodeLimit.
Eventual consistency
Two independent lags apply. Neither is an error state, and both are visible in the response.
1. The indexer's cursor, not chain head. Responses reflect blocks the
indexer has processed. A transaction confirmed seconds ago may not be visible
yet. Every entity carries the blockNumber it was observed at, which is what
you should use for ordering and for judging freshness.
2. Off-chain metadata enrichment lags further. Names and descriptions live
off-chain and are fetched by a separate job after the on-chain record is
indexed. metadataStatus tells you where that stands:
| Value | Meaning |
|---|---|
PENDING | Not fetched yet — including the case where no fetch job row exists |
IN_PROGRESS | Fetch running |
DONE | Fetched and stored |
FAILED | Fetch failed; will not resolve without intervention |
The consequence that catches people: "name": null means "not known yet", not
"has no name". Read it together with metadataStatus. Rendering "Untitled"
for a PENDING record is wrong; it will have a name shortly.
metadataFetchedAt is the timestamp of the successful fetch, when there was one.
The same rule reaches the graph reads: previousDatasetId is enrichment-derived,
so a dataset whose metadataStatus is not DONE can return a single-node
lineage graph even though an ancestor exists on chain.
Finding a starting point
There is no dataset-list endpoint. To reach a dataset without already having its id, walk down:
or, if you have a DID, /v1/dids/{did}/holdings gives dataset ids directly.
Verifying a proof
The claim proof is two layers. Understanding both is the difference between "the API says 500 shares" and "500 shares is provably what the chain anchored".
Layer 1 — the CF is in the contributor's CF set
One tree per contributor per dataset version. Leaves are raw cfId values,
used directly — a cfId is already a keccak256 hash and is not re-hashed.
Leaves are sorted by cfId byte-ascending. The root is cfMerkleRoot.
Verify cfProof for cfId at cfLeafIndex against cfMerkleRoot.
Layer 2 — the contributor is in the dataset's contributor set
One leaf per contributor. The root, contributorsMerkleRoot, is anchored
on-chain in the dataset version record.
abi.encode left-pads each field into a 32-byte slot — 192 bytes total. The
hash is applied twice. Leaves are sorted by contributorDidId ascending.
Verify claimProof for that leaf at contributorLeafIndex against
contributorsMerkleRoot.
The last two fields are what bind a proof to one deployment, so take both from the response rather than hardcoding either. The endpoint returns them for exactly this reason:
Internal nodes, both layers
Sorted-pair hashing, matching OpenZeppelin's MerkleProof:
Comparison is as 256-bit big-endian unsigned integers. On an odd layer, the last node is duplicated before pairing.
Using the SDK instead
The repository's consumer SDK includes a local Layer-2 verifier. With Python 3.12 or later, install it from a checkout of the protocol repository (run these commands at the repository root):
This example uses only the proof helper; the SDK's legacy PRE access-resolution examples do not describe the current access workflow. No private key, RPC connection or write SDK is needed to verify this response.
Save a successful JSON body from
GET /datasets/{id}/cf/{cfId}/proof
as proof.json, then run the following with .venv/bin/python.
contributor_did_id and share_amount must be ints, not strings:
Trusting the result
The service rebuilds both trees from stored rows on every proof request and
compares the result against the anchored roots. A mismatch is reported as a
500 with INTERNAL_ERROR rather than a proof you might act on. So a 200
from this endpoint means the served proof reproduces the on-chain root — but
verify it yourself anyway if you are about to submit a claim; that is the point
of the proof.
The normative specification for both trees — leaf formats, ordering, the
assembly flow, and the verification steps a claimant should run — is
Ownership Merkle tree specification in the protocol repository. This page
covers what a caller of the proof endpoint needs; that document covers the
construction. One thing it is worth reading there before writing a checker:
shareAmount is cfCount * 100 only when the source carries no weight. A
publication snapshot may supply a contributor-level amount that is used
verbatim, so do not assert the derived formula unconditionally.
Errors
Non-2xx responses carry a code you can match on. Every layer emits a single shape:
Read body.error?.code. There is no flat shape to fall back to — describing one
would teach a parser branch for a body that is never produced.
Codes
From the public API service:
| Code | Status | Meaning |
|---|---|---|
NOT_FOUND | 404 | Path is not part of this API |
INVALID_PATH | 400 | Percent-encoded path separator (%2F) in the path |
METHOD_NOT_ALLOWED | 405 | Any verb other than GET or OPTIONS. Carries an Allow header |
RATE_LIMITED | 429 | Over 120/min. Honour Retry-After |
UPSTREAM_UNAVAILABLE | 502 | Lineage service unreachable, or slower than the upstream budget |
From the lineage service behind it:
| Code | Status | Meaning |
|---|---|---|
BAD_REQUEST | 400 | Malformed id, limit, cursor, or filter |
NOT_FOUND | 404 | Entity does not exist at the indexer's cursor |
PROOF_NOT_READY | 409 | Proof columns not yet populated. Retry later |
UNAVAILABLE | 503 | Endpoint not configured |
INTERNAL_ERROR | 500 | Server fault, including a failed proof integrity check |
NOT_FOUND is deliberately the same code for "this route does not exist here"
and "this record does not exist": from outside, both are 404s, and a second code
would only describe the service's internals.
404 deserves care given eventual consistency: for an entity created very
recently it may mean "not indexed yet" rather than "does not exist".
PROOF_NOT_READY (409) is the explicit "exists but not ready" signal, and it
applies only to the proof endpoint.
One code is declared in the service's error package but never emitted by any
handler: INVALID_PARAMETER (422). Do not write branches for it.
What is deliberately not here
Five paths exist upstream and are withheld from this API:
/v1/datasets/{id}/overview, /v1/cfs/{cf_id}/overview,
/v1/dids/{did}/owned, /v1/frontiers/{id}/activity, and /v1/search.
All are shaped for a consuming page — one-call rollups over granular reads, a
merged registry list, a merged event timeline, and an id lookup for a search box
— so they follow a UI rather than the protocol, and building against them would
mean building against a moving target. Everything a rollup carries is composable
from this surface: the dataset overview is /datasets/{id} plus /contributors
plus /holders, and the CF overview is /cfs/{cf_id} plus a per-dataset
/datasets/{id}/cf/{cfId}/proof call.
One more once existed and has been removed rather than withheld:
/v1/datasets/{id}/activity, a merged share-movement and royalty timeline no
client ever adopted. Nothing was lost with it: it was /circulation plus
/royalty-flow.
There is also no rate-limit header, no cache header, no conditional-request support, no webhook, and no write path.
Planned changes
The endpoint shapes track on-chain primitives, and the v2 storage layout is frozen — so they are the stable part of this document. Two things are not:
- Auth. Anonymous today. API keys are under discussion; if introduced, anonymous access to this surface may be reduced.
- Rate limits. The key is settled — the address the load balancer observed — but the counter is still per pod. A shared-store limiter would change the effective numbers.
Coverage will also grow rather than change shape: the dataset, ownership and royalty endpoints documented above begin returning data once those registries are deployed to Base mainnet, without their responses changing.
Last updated