ECAI Index NFTs: Give a Knowledge Release a Verifiable Identity
What ECAI index artifacts and indexing-job NFTs represent, how IPFS and lineage fit together, and what remains before an automated paid indexing market.
Make the release itself something people can reference
Preparing a useful knowledge collection takes work: choosing sources, normalising records, building an index and checking whether it answers the questions users actually ask. If another team wants that collection, a URL labelled "latest index" provides little assurance about what it will receive.
ECAI's artifact path gives the release a more precise description. A manifest records source identity, pipeline information, counts and output-file digests. Optional IPFS publication supplies content-addressed references. NFT metadata can then describe that release with identifiers that other applications can retain and check.
The potential is a reusable publication unit for knowledge infrastructure. A curator could distribute an evaluated reference corpus, an application could pin the version it uses, and a later release could identify its predecessor. The current implementation supplies parts of that workflow. Automatic minting, verification of useful work and payment must be assessed separately.
Separate the finished artifact from the job
The source uses NFT-related concepts in two different places.
| Concept | What it describes | Current implementation boundary |
|---|---|---|
| Index artifact and NFT metadata | A completed index release and its published manifest | Finalisation, hashes, optional IPFS publication and metadata generation |
| Indexing-job NFT | A dataset chunk, line range and job identity | Metadata adapter and owner-controlled contract entry point; the reviewed mint helper dry-runs its call |
A token describing work to perform is not evidence that the work was successfully completed. Conversely, a useful index can be built and used locally without minting a token.
What the artifact commits to
The artifact finaliser constructs a canonical manifest from source and pipeline
descriptions, target settings, record counts and file roles, lengths and
SHA-256 digests. Local paths and build timestamps are excluded from its
identity material. The resulting index_root is a digest of that canonical
identity; it is distinct from the per-term posting roots used for search
membership proofs.
The contents depend on the indexing mode. Wikimedia visibility jobs include a full search snapshot and their corpus-selection material. Disk artifacts include posting segments and the document store. Legacy live JSONL and Yelp artifacts contain shared-context term headers, so they should not be sold as self-contained copies of a single job's dataset.
A previous_manifest_cid can connect a release to an earlier one. The source
frontier digest incorporates that predecessor reference and source identity.
This records lineage; it does not prove that every upstream change was
processed or that the build reused previous work incrementally.
Publication adds an address, not permanent availability
With IPFS publication enabled, the finaliser publishes files and then the manifest. It reads the manifest back and compares the returned bytes before reporting it ready for NFT metadata. A local-only artifact can complete without entering that ready-to-mint condition.
IPFS content addressing identifies content independently of a particular hosting location. A CID should not be confused with the manifest's ordinary file SHA-256 digest: encoding and import choices participate in IPFS addressing. Keeping the material retrievable still requires a persistence and pinning policy. A token reference does not supply that storage on its own.
The generated NFT metadata carries the manifest reference and digest, index root, namespace, pipeline, counts and predecessor information. The schema version varies with the artifact path. A consumer should validate the stated schema and required files instead of assuming every index NFT has the same payload.
Where the on-chain workflow currently stops
The knowledge NFT contract includes ownership, transfer and approval behaviour,
plus an owner-controlled mint_index_job entry point. Its job key supports
idempotent lookup, and the contract checks token identifiers and stores
dataset, chunk and kind indexes with the payload.
The reviewed Erlang mint_index_job helper constructs the chunk metadata and calls
damage_ae:contract_call_dry. That path simulates a contract call; it does
not submit the mint transaction. The durable index queue exposes ready
metadata, but no automatic artifact-to-mint submission flow was established
by this review.
There is also a separate chunk-job manager with claim, submit and pay operations.
It keeps jobs in memory. Its pay operation marks the job paid locally;
the chain creation, claim, submission and payment calls are comments or
integration placeholders. That status is not evidence that money moved.
These distinctions matter for adoption. The reviewed paths support describing artifacts and representing job identities, but do not establish a complete paid indexing market. A production workflow would need transaction submission and confirmation, work acceptance, durable payment state and recovery from partial failure.
The 7 October integration review found no change to these artifact, mint-helper or chunk-job paths. The mint submission and settlement gaps remain open.
What this could enable
Versioned reference collections are the nearest opportunity. A team could publish a tested corpus and let several assistants use the same identifiable release. A research group could keep an evaluation collection fixed while comparing retrieval methods. A specialist curator could publish its selection policy and evidence alongside an index for others to inspect.
A paid indexing service is a further possibility, provided the buyer can verify the delivered material and the payment workflow is completed. The valuable work would be collection quality, coverage, maintenance and reliable delivery. Token ownership alone does not make the source text exclusive or replace its existing attribution and reuse conditions.
Freshness also needs a publication policy. New content creates a new release; an older token continues to identify its older material. A service needs an explicit way to discover and approve successors instead of treating every historical identifier as "latest".
Prove delivery before building a market around it
A useful pilot begins with a public fixture collection. Publish its artifact, retrieve only the published files on a clean consumer, check their hashes, reload the index and run a fixed set of queries. Where version 2 proofs apply, verify them against the intended release and reject altered records.
Only then add a test-chain mint with a confirmed transaction and independently checked ownership events. Payment needs its own end-to-end test; neither a dry run nor a local paid flag is sufficient. The artifact tests provide source-level examples of identity and readiness checks, not a claim that this complete pilot has already passed.
See Wikipedia corpus releases, indexing workflows and search evidence and its verification limits for the components that make such a release useful.