Skip to main content

EVM Data

Bitquery provides blockchain data dumps for EVM-base chains like Ethereum, BSC, Base, Polygon/Matic, Optimism, Robinhood, etc. in parquet format that you can host directly in your own cloud (for example AWS S3) and plug into your analytics stack or data lake.

Available Topics

For EVM chains we currently provide the following topics:

Sample Ethereum Cloud Dataset

To explore the schema and test your tooling, use our public sample EVM datasets on GitHub:

The GitHub repository includes one sample file. The complete list of Parquet files is stored in our public S3 bucket and can be accessed directly. For example:
https://bitquery-blockchain-dataset.s3.us-east-1.amazonaws.com/ethereum/balance_updates/24053500_24053549.parquet


bitquery-blockchain-dataset/
└── ethereum/
├── balances/
│ ├── 2025-01-01.parquet
│ ├── 2025-01-02.parquet
│ └── ...
├── balance_updates/
│ ├── 24053500_24053549.parquet
│ ├── 24053550_24053599.parquet
│ ├── 24053600_24053649.parquet
│ ├── 24053650_24053699.parquet
│ ├── 24053700_24053749.parquet
│ ├── 24053750_24053799.parquet
│ ├── 24053800_24053849.parquet
│ ├── 24053850_24053899.parquet
│ ├── 24053900_24053949.parquet
│ └── 24053950_24053999.parquet
├── blocks/
│ ├── 24053500_24053549.parquet
│ ├── 24053550_24053599.parquet
│ ├── 24053600_24053649.parquet
│ ├── 24053650_24053699.parquet
│ ├── 24053700_24053749.parquet
│ ├── 24053750_24053799.parquet
│ ├── 24053800_24053849.parquet
│ ├── 24053850_24053899.parquet
│ ├── 24053900_24053949.parquet
│ └── 24053950_24053999.parquet
├── calls/
│ ├── 24053500_24053549.parquet
│ ├── 24053550_24053599.parquet
│ ├── 24053600_24053649.parquet
│ ├── 24053650_24053699.parquet
│ ├── 24053700_24053749.parquet
│ ├── 24053750_24053799.parquet
│ ├── 24053800_24053849.parquet
│ ├── 24053850_24053899.parquet
│ ├── 24053900_24053949.parquet
│ └── 24053950_24053999.parquet
├── dex_trades/
│ ├── 24053500_24053549.parquet
│ ├── 24053550_24053599.parquet
│ ├── 24053600_24053649.parquet
│ ├── 24053650_24053699.parquet
│ ├── 24053700_24053749.parquet
│ ├── 24053750_24053799.parquet
│ ├── 24053800_24053849.parquet
│ ├── 24053850_24053899.parquet
│ ├── 24053900_24053949.parquet
│ └── 24053950_24053999.parquet
├── events/
│ ├── 24053500_24053549.parquet
│ ├── 24053550_24053599.parquet
│ ├── 24053600_24053649.parquet
│ ├── 24053650_24053699.parquet
│ ├── 24053700_24053749.parquet
│ ├── 24053750_24053799.parquet
│ ├── 24053800_24053849.parquet
│ ├── 24053850_24053899.parquet
│ ├── 24053900_24053949.parquet
│ └── 24053950_24053999.parquet
├── miner_rewards/
│ ├── 24053500_24053549.parquet
│ ├── 24053550_24053599.parquet
│ ├── 24053600_24053649.parquet
│ ├── 24053650_24053699.parquet
│ ├── 24053700_24053749.parquet
│ ├── 24053750_24053799.parquet
│ ├── 24053800_24053849.parquet
│ ├── 24053850_24053899.parquet
│ ├── 24053900_24053949.parquet
│ └── 24053950_24053999.parquet
├── transactions/
│ ├── 24053500_24053549.parquet
│ ├── 24053550_24053599.parquet
│ ├── 24053600_24053649.parquet
│ ├── 24053650_24053699.parquet
│ ├── 24053700_24053749.parquet
│ ├── 24053750_24053799.parquet
│ ├── 24053800_24053849.parquet
│ ├── 24053850_24053899.parquet
│ ├── 24053900_24053949.parquet
│ └── 24053950_24053999.parquet
├── transfers/
│ ├── 24053500_24053549.parquet
│ ├── 24053550_24053599.parquet
│ ├── 24053600_24053649.parquet
│ ├── 24053650_24053699.parquet
│ ├── 24053700_24053749.parquet
│ ├── 24053750_24053799.parquet
│ ├── 24053800_24053849.parquet
│ ├── 24053850_24053899.parquet
│ ├── 24053900_24053949.parquet
│ └── 24053950_24053999.parquet
└── uncle_blocks/
├── 15535500_15535549.parquet
├── 15535550_15535599.parquet
├── 15535600_15535649.parquet
├── 15535650_15535699.parquet
├── 15535700_15535749.parquet
├── 15535750_15535799.parquet
├── 15535800_15535849.parquet
└── 15535850_15535899.parquet

Use these samples to:

  • Validate your ETL / analytics pipeline against realistic EVM data.
  • Inspect column names and types before connecting to full buckets.
  • Benchmark query performance on your preferred engines and hardware.

Balances (Daily Snapshots)

The Balances topic is a daily snapshot of account balances, both native (ETH) and token. It differs from Balance Updates in two ways:

  • Balance Updates are deltas — one row per balance change, which you must sum to reconstruct a balance.
  • Balances are levels — one row per account/token with the balance as of the end of that day, already aggregated.

Files are partitioned by date, not block range:

https://bitquery-blockchain-dataset.s3.us-east-1.amazonaws.com/ethereum/balances/<YYYY-MM-DD>.parquet

Sample Parquet download (public S3)

Each daily file covers only the accounts whose balance changed that day, not the entire chain state. The 2025-01-01 sample holds about 954,000 rows across roughly 514,000 addresses and 13,000 currencies. To reconstruct full chain state at a date, carry forward the last known balance per account from earlier files.

Balances Schema

ColumnTypeDescription
Balance_AddressstringAccount holding the balance
Block_DatedateSnapshot date, matches the file name
Currency_SmartContractstringToken contract address; 0x for native ETH
Currency_SymbolstringToken symbol as declared by the contract
Currency_NamestringToken name as declared by the contract
Currency_ProtocolNamestringToken standard, e.g. erc20, erc721, erc1155, erc404_v1
Balance_AmountstringBalance at end of day, decimal string — see the precision note below
Balance_FirstChangeTimedatetimeFirst balance change on this date (UTC)
Balance_LastChangeTimedatetimeLast balance change on this date (UTC)
Balance_UpdateCountuint64Number of balance changes on this date
Balance_RowCountuint64Number of underlying aggregate rows merged into this row

Filter on the Contract Address, Never the Symbol

Token symbols are set by the contract, so anyone can deploy a token claiming any symbol. In the 2025-01-01 sample, 25 different contracts report the symbol ETH and 34 report USDT.

Filtering Currency_Symbol = 'ETH' picks up impostor ERC-20s and inflates the native ETH total by several orders of magnitude. Native ETH is identified by the contract address:

-- correct: native ETH
WHERE Currency_SmartContract = '0x'

-- correct: real USDT
WHERE Currency_SmartContract = '0xdac17f958d2ee523a2206206994597c13d831ec7'

-- wrong: matches impostor tokens too
WHERE Currency_Symbol = 'USDT'

Filtered correctly, the largest native ETH holders in the sample are the Beacon Deposit Contract (0x00000000219ab540356cbb839cbe05303d7705fa) and the WETH contract (0xc02aaa39b223fe8d0a0e5c4f27ead9083c756cc2), which you can cross-check on any block explorer.

Balance_Amount Is a String

Balances are stored as decimal strings, not floats, so that 18-decimal values survive intact. Casting to a 64-bit float silently loses precision on large balances. Parse to a decimal type instead:

from decimal import Decimal
df["amount"] = df["Balance_Amount"].map(Decimal)

In SQL, cast to a wide decimal — for example CAST(Balance_Amount AS DECIMAL(38, 18)) — rather than DOUBLE. Note that a few scam tokens carry balances near 2^256, which overflow even a DECIMAL(38, 18); filter those out or cast to DECIMAL(76, 18) if your engine supports it.

Row Grain and Deduplication

The grain is (Balance_Address, Currency_SmartContract, Currency_ProtocolName). Hybrid tokens such as ERC-404 emit under more than one standard, so the same address and contract can appear on multiple rows — about 4,600 such pairs in the sample.

These rows often repeat the same balance under a different Currency_ProtocolName, so summing across them double-counts. Pick one Currency_ProtocolName, or deduplicate on the address and contract before aggregating.

Reading a File in Python

import pandas as pd
from decimal import Decimal

url = "https://bitquery-blockchain-dataset.s3.us-east-1.amazonaws.com/ethereum/balances/2025-01-01.parquet"
df = pd.read_parquet(url)

# native ETH only, identified by contract address rather than symbol
native = df[df.Currency_SmartContract == "0x"].copy()
native["amount"] = native.Balance_Amount.map(Decimal)

print(len(native), "accounts changed ETH balance on this date")

# sort_values, not nlargest: pandas cannot rank an object/Decimal column
top = native.sort_values("amount", ascending=False).head(10)
print(top[["Balance_Address", "amount"]])

Other Ways to Access EVM Data

Cloud data dumps are ideal for batch analytics and historical workloads.
If you need low-latency real-time data, you can also consume Bitquery streams via Kafka and GraphQL subscriptions.