Blockchain data indexing
Ask a blockchain node for every transfer a wallet ever made and it cannot tell you, at least not in a way anyone can wait for. Blockchains are built for verification and immutability, not for search, so every onchain product a bank or a fund touches sits on an index built from the raw chain.
What follows is why a block is the wrong shape for a query, what an indexer does with events and traces, how a chain reorganization is handled, what contract upgrades do to decoding, and the gap between an indexed figure and an audited one.
Why a block is the wrong shape for a question
A chain stores blocks in sequence, each holding transactions and the logs they emitted. Nothing in that structure answers a question like "all transfers involving this address" or "the balance of this token across these hundred accounts", because the data is organized by time and not by subject. Answering such a question from the chain alone means scanning millions of blocks in order, which is why standard RPC methods do not offer it.
An indexer solves that by inverting the organization. It reads blocks once, extracts the events and the execution traces, decodes them into records with named fields, and writes them into a database with indexes on the fields people query, which is an extract, transform and load pipeline of the kind a bank already runs for other data. The query then takes milliseconds because the work was done in advance.
Events against traces, and why both are needed
An event log is what a contract deliberately emitted, which makes it cheap to read and dependent on the contract author having emitted it. Most token transfers appear as events, which is why a transfer history is the easiest thing to index.
Execution traces are the record of what actually happened inside a transaction, including internal calls between contracts that emitted nothing. They are needed exactly where events are absent or misleading: a native currency transfer inside a contract call, a failed sub-call in a transaction that otherwise succeeded, or a contract that simply does not log an action. The practical consequence for a firm reading a provider's index is to ask which of the two it is built from, because an index built on events alone has blind spots that only show up when a balance fails to reconcile.
Subgraphs and the indexing services behind them
The Graph is the most widely used decentralized indexing protocol, and the unit of work in it is a subgraph: a specification naming which contracts to watch, which events to listen for and how to map that event data into a schema that can be queried over GraphQL. Independent operators called indexers run the processing and serve the queries.
Around that sit two other shapes. A managed indexing provider runs the pipeline and hands over an API, which removes the operations and adds a supplier. A push or streaming service inverts the direction, delivering filtered chain data into the firm's own destination as it happens, which suits a firm that wants the data in its own warehouse next to everything else it reports from. Each of the three is a different dependency, and the DORA question on the RPC providers page applies to all of them.
Reorgs: when the chain changes its mind
A chain reorganization happens when the canonical chain shifts because competing blocks were produced at the same height and the network settles on a different fork from the one the indexer already processed. An indexer that only processes blocks past the finality threshold never has to deal with this, and one that follows the chain head always does. Everything the indexer wrote from the orphaned blocks is now wrong: transactions that it recorded as having happened did not happen on the chain everyone else now follows.
A correct indexer therefore detects the reorg, rolls back the affected records and reprocesses from the fork point, and an indexer that does not do this reports phantom transactions that never settle. For a financial firm this is not an engineering footnote. A figure read before finality can be revised, so a reporting or a valuation process has to state which confirmation depth it treats as final and has to be able to restate when a reorg crosses that line. Any index whose provider cannot describe its reorg handling should not be the source for a reported number.
Decoding a contract: ABIs, proxies and upgrades
Chain data arrives as hexadecimal. Turning it into a named field requires the contract's ABI, the interface description that says which event has which parameters in which order, and without the right ABI an event is an undecodable blob. That makes an index dependent on a piece of information that lives outside the chain.
Proxy contracts make this harder in a way that catches reporting processes out. A proxy holds the storage and forwards calls to an implementation contract that can be replaced, so the address stays the same while the logic and the event signatures change under it. An index that decoded against the old implementation keeps producing plausible records from the wrong schema after an upgrade, and nothing in the data announces the change. The question to ask a provider is how it detects an implementation change, and the question for a firm's own pipeline is the same one.
An indexed figure is not an audited one
This is the point where the subject stops being technical. A number from an index is the output of a pipeline with choices in it: which chains were included, which confirmation depth counted as final, which addresses were attributed to which entity, which ABI version was used to decode, and what happened to the records from a reorg. Two providers can produce two different totals for the same portfolio without either being wrong about the chain.
For a fund's net asset value or a bank's reported holding, the figure needs a documented derivation and a reconciliation against a second source, which for custodied assets is the custodian's own records. The useful discipline is the one already used for market data: name the source, name the time, name the method, and keep the inputs so the number can be reproduced later. The onchain analytics in finance page covers the attribution layer on top, where the error is larger.
Should a firm build or buy an index?
Buy for breadth, build for the figures the firm is accountable for. A custom indexer gives full control over decoding, reorg handling and schema, and it costs a team to build and keep running: block ingestion, decoding, schema design, reorg recovery, monitoring and the backfill whenever the schema changes. A provider removes all of that and makes the firm's reported numbers depend on a method it does not control.
The split most firms arrive at is a provider for discovery and broad coverage, plus a narrow own pipeline for the handful of contracts and addresses that feed a report. The narrow pipeline is small enough to validate and audit, which is the property that matters.
How long does it take to index a chain from scratch?
Days to weeks for a busy chain, and the figure is dominated by the backfill and not by keeping up. Reading every historical block, decoding it and writing it to a database is bounded by the archive node's read throughput and the pipeline's write throughput, and a chain with years of history and high throughput means terabytes to process. The operational trap is that any schema change or ABI correction requires a backfill of the affected range again, so a firm plans for reindexing as a normal event and not an incident.
Onchain data and Finance Loop
Finance Loop brings the data engineers who build these pipelines together with the fund accountants and reporting teams who have to defend the numbers that come out of them, in its Investment & Digital Assets and Digital Infrastructure & Sovereignty tracks. Finance Loop runs these sessions in Frankfurt, next to the funds and banks that report onchain holdings.
Finance Loop is a professional network and has the goal of driving the adoption of emerging technologies in finance, such as AI, tokenization, stablecoins, and DeFi. Finance Loop helps its members build skills and personal networks in these fields: Investment & Digital Assets, Payments & Digital Money, Digital Infrastructure & Sovereignty, and Risk & Compliance.