Thegraph

Thegraph subgraphs map blockchain records into entities for GraphQL queries

Thegraph subgraphs turn selected blockchain data into structured entities for GraphQL queries. A manifest selects indexing inputs, mapping code creates or updates records, and a schema defines their fields. The API exposes that indexed data, so its answers depend on the mapping's meaning and the blocks that indexing has processed.

Key takeaway: A valid GraphQL response is useful only when its deployment, entity meanings, and indexed block match the application's data needs.

Manifest, mappings, and schema connect the API to its data

The manifest connects source data and handlers to the schema that mappings populate. In subgraph.yaml, a data source identifies the network and the triggers that matter. The contract ABI, or application binary interface, describes the contract's events and functions. AssemblyScript mappings transform those inputs into records that match schema.graphql. An activity history needs persistent event details; a changing state view needs records that mappings update. Naming a field in the schema defines its type, but the mapping still has to assign its value.


Subgraph coverage and entity meanings determine whether reuse fits the application

Reuse fits when an existing subgraph covers the required contracts, history, entity meanings, and query fields. A matching protocol name alone doesn't establish that coverage. Source addresses and block ranges describe which activity contributes to the dataset. Field definitions describe what the API exposes, while mapping logic determines what those fields mean. Graph Explorer helps locate published subgraphs, but each deployment has its own data model. Query changes can retrieve existing records differently; they can't manufacture history that the mapping never indexed.


Entity design separates changing state from immutable records

Entity types should reflect the questions that the application needs to answer. Their IDs distinguish records, scalar types preserve values, and references connect related objects without requiring every query to reconstruct those relationships.

Mutable state and immutable records

A status that changes across blocks belongs in a mutable entity. A record that won't change can use @entity(immutable: true), reducing the database work associated with versioned updates. Graph Node still allows changes within the block in which the mapping creates an immutable entity. Every entity needs an ID unique within its type. Bytes suits binary identifiers; string identifiers suit readable text. Required scalar fields carry !, and saving an entity without their values causes an error.

References and derived collections

An entity reference stores the related record's ID. The @derivedFrom directive provides a reverse lookup using a reference on the other entity. This design avoids repeatedly loading and saving a growing array on the parent record. Queries can traverse the relationship and select fields from related entities. The relationship still depends on mappings storing the correct references, and large nested collections still need bounded queries.


Handler types determine which changes enter the store

For Ethereum Virtual Machine (EVM) contract sources, handlers can respond to emitted events, matching function calls, or block triggers. The data source kind determines the available triggers, and each handler type has different indexing requirements.

Event inputs and contract reads

An event handler receives typed event data and can create, load, update, or remove entities. Mappings use AssemblyScript, which compiles to WebAssembly. They can also read contract state through eth_call. Those reads can add requests to the blockchain data provider and can slow indexing. If event parameters already contain the required value, using them avoids that extra read. Event signatures and the ABI must match the activity that the handler expects.

Call tracing and block frequency

Call handlers and call-filtered block handlers require compatible tracing support. Without it, a subgraph configured with these handlers won't start syncing. General support for a network doesn't establish support for those triggers. An unfiltered block handler runs for every block. Polling filters change that frequency where the manifest version and data source kind support them. NEAR data sources use receipt and block handlers, so handler settings must match the selected source kind.


Templates and composition extend the indexing inputs

When contracts emerge during indexing, a supported data source template can provide their shared indexing definition. A mapping supplies the discovered address to create a dynamic contract source. This fits factory or registry patterns without listing every future contract address in advance. Multiple ordinary contract sources can contribute to one subgraph on the same chain.

Visual outline: Templates and composition extend the indexing inputs (Thegraph subgraphs)

View full-size image

Subgraph composition uses entities from source subgraphs as triggers for a dependent subgraph. Composition requires manifest specVersion 1.3.0 or later, immutable source entities, and sources on the same chain. This mode doesn't support nested composition or mixing direct contract handlers with subgraph-source triggers. Source deployment IDs identify the inputs, so a changed source deployment requires an updated reference to use its changes. Composition therefore reuses indexed records while adding a separate mapping and schema.

Client-side composition combines queries from separate subgraphs in an application; it doesn't create the same entity-triggered indexing dependency.

GraphQL selections, filters, and pagination shape the response

A query can retrieve only fields that the deployed schema exposes. Graph Node generates fields for individual entities and entity collections, with arguments for filtering and ordering. Selection sets specify the returned fields, including fields on related entities. A subgraph's generated query API exposes read operations without GraphQL mutations. Static query documents with typed variables separate the query structure from changing input values. This also lets tooling validate the operation against its schema.

Collection arguments include where, orderBy, orderDirection, and first. A collection response is a page, so a successful request doesn't establish that it contains every matching entity. Large skip offsets can perform poorly. Attribute-based pagination can continue from the last returned value when the filter and ordering agree.

Diagram: Thegraph subgraphs: GraphQL selections, filters, and pagination shape the response

View full-size image

An ID-based order doesn't establish chronological order. A history view needs a suitable stored ordering field.


How do paginated queries stay on the same indexed snapshot?

Paginated queries need the same available indexed block, deployment, filters, and ordering to traverse a stable dataset. Otherwise, writes between requests can change the records that later pages encounter.

The _meta field exposes indexing information, including the block, deployment identifier, and indexing-error flag. An HTTP response alone doesn't show that the requested activity has reached the entity store.

Block-specific query arguments support a block number or hash. Historical reads require the deployment to retain the corresponding entity state. Pinning a block therefore depends on both indexing progress and retention; a requested snapshot isn't available merely because its block exists on the chain.

A block hash identifies a particular block, while a block number identifies a height. Stable historical reads need a final block on the canonical chain. A height alone doesn't distinguish competing blocks during a reorganization.

Pagination also needs a continuation rule consistent with the sort order, including how it handles equal ordering values.


Build checks and mapping tests cover different defects

A successful build establishes that the subgraph's components compile; it doesn't establish that their records have the intended meaning. Graph CLI uses graph codegen to generate types from the schema and contract ABIs, then graph build to compile mappings and prepare deployment artifacts. The manifest specification version controls supported manifest features, while the mapping API version controls available host behavior.

Matchstick tests mapping logic with mocked events, contract calls, and assertions about stored entities. Such tests can cover missing initial records or repeated updates. Their fixtures still determine what they exercise. Indexing in a staging or local Graph Node environment exposes different issues, including data-provider compatibility and failures that occur while processing actual chain history.


Query scope and request volume affect resource use

A query is usable when it returns the required records at the intended indexed snapshot without relevant errors. Request volume and query structure affect different parts of the resource budget.

For an illustrative comparison, keep the deployment, available block, and filters unchanged. A bounded request selects the fields and collection size that an application needs. Expanding nested collections can increase returned records and database work even when the application sends a single request. A timeout leaves that larger response unavailable. Splitting the retrieval into pages changes request volume, which can affect metered usage.

  • Use fields and relationships that the deployment's schema exposes.
  • Require an indexed, retained block when the task needs a fixed snapshot.
  • Bound nested collections and paginate when the task needs all matching records.
  • Handle relevant query errors before using the response as complete data.
  • Match the projected request volume to the endpoint's access limits and billing model.

A gateway's consumer pricing can differ from its underlying costs of serving queries. A larger response alone doesn't prove a higher monetary charge. The endpoint's metering rules determine how changed usage affects payment.

Query failures and indexing failures need different repairs

The failing layer determines the remedy. A request can fail before query execution, fail schema validation, or reach an entity store that indexing errors have left incomplete.

Transport, access, and schema errors

Network timeouts and rejected credentials concern request delivery or access. Unknown fields and incompatible variable types concern the deployed schema. Repeating a malformed query preserves its defect. A transient transport failure can justify a bounded retry, while an access restriction requires the appropriate credential or endpoint policy. GraphQL responses can contain errors alongside data, so clients must interpret the response body before accepting the required fields.

Visual outline: Thegraph subgraphs: Transport, access, and schema errors

View full-size image

Failed handlers and skipped writes

A fatal mapping error can stop indexing at the affected block. Where the deployment supports and enables non-fatal deterministic errors, Graph Node can discard the failing handler's changes and continue. Queries must opt in through subgraphError: allow to read potentially inconsistent data. The Graph Network doesn't support non-fatal errors, and some errors remain fatal. An advancing indexed block can therefore coexist with skipped writes. Repeated query attempts won't reconstruct those writes. Recovering the records requires corrected mapping logic and reprocessing of the affected inputs.

Thegraph subgraphs: frequently asked questions

Why can events with the same transaction hash need different entity IDs?

A transaction can emit multiple events, so its hash alone may collide when each event creates a record of the same entity type. Combining the transaction hash with the log index distinguishes those events. If one handler creates several entities of that type from a single event, their IDs need an additional distinguishing component.

When does a dynamic EVM contract data source begin collecting events?

A newly instantiated EVM contract data source processes its creation block and subsequent blocks. It doesn't replay earlier blocks for that contract. The creation block here means the block in which the mapping instantiates the data source, which may differ from the contract's deployment block. Reading contract state at that point can initialize state, but it doesn't recreate earlier event history.

Does @derivedFrom calculate totals across related records?

The @derivedFrom directive resolves reverse relationships; it doesn't calculate sums or averages. A total needs separate calculation logic or a supported aggregation definition. Built-in aggregation entities operate over defined timeseries data and configured intervals.

Is BigDecimal suitable for every fixed-point amount in a mapping?

BigDecimal has a finite precision of 34 significant decimal digits, so it can't preserve every wider fixed-point value exactly. The graph-ts implementation uses decimal128 representation. Storing the original integer quantity as BigInt and keeping its scale separate can preserve exact input data, while decimal conversion or arithmetic may introduce rounding.

Can ordinary JavaScript packages run inside a subgraph mapping?

Ordinary JavaScript packages aren't automatically compatible with AssemblyScript mappings. Their code and dependencies must fit the mapping's compiled WebAssembly environment and supported APIs. A library that assumes a JavaScript runtime isn't a drop-in dependency. Logic that requires such a library can remain in the application that consumes the indexed records.

What changes when a mapping update keeps the same schema?

Changing a mapping can change the meaning of stored fields even when their names and types stay the same. A changed compiled mapping produces a different deployment identity. An application therefore needs the intended deployment or published version, and that deployment needs the relevant indexed data. Schema compatibility alone doesn't establish equivalent results.

Will pruning discard the latest values of mutable entities?

Pruning removes historical versions of mutable entities, rather than their latest stored state. Its retention setting determines which earlier snapshots remain available. Removing an old version doesn't erase a separate immutable record merely because that record originated in an old block. Historical queries, grafting, and rewinding require the entity state at their target block to remain available.

How can a browser application keep a privileged query credential private?

A browser can't keep a credential confidential when public frontend code or browser requests expose it. A server-side proxy can hold privileged credentials and enforce application access rules. Public query keys can have domain, subgraph, or usage restrictions where the gateway supports them. Those restrictions limit permitted use; they don't turn an exposed key into a secret.
Updated on