Base Archive Bootstrap
Index Base history that peers no longer serve, straight from the official snapshot archives.
Why
Many Base peers prune old receipts, so a deep backfill over P2P alone often has no one to ask. Base publishes V2 history snapshots split into aligned archive groups. Sieve can read those directly: it downloads one group at a time, indexes only what your config asks for, deletes the group's transaction and receipt files, and moves on.
No history node, no execution client, no second database. The normal sieve command does the import, and with handoff = true the same process continues over P2P into live follow.
archive groups (download, verify, index, clean up)
|
v
peer history (from archive end + 1)
|
v
live follow1. Plan the import
Download the snapshot's small manifest.json and get its SHA-256 from a source you trust. Then run the offline planner on your exact range:
sieve archive-plan --manifest base-manifest.json \ --manifest-sha256 <64-hex-digit-digest> \ --start-block 500123 --end-block 1000010
It reads only the local manifest. No config, database, peers, or network. The JSON plan on stdout lists the selected archive groups and their files, which blocks are decoded only as dependencies, total download and extracted bytes, the peak staging space needed, and the trust inputs still required. Blocks before your start inside the first group are decoded to establish transaction positions but are never indexed.
Size your budgets from the plan. Early groups are small. Recent groups are large: the group covering blocks 51.5M to 52M is about 25 GB compressed. Staging is temporary and released per group: once a group is indexed, its transaction and receipt files are deleted and only its headers are kept. archive-plan reports the exact download and staging needs for your range. See CLI for every flag.
2. Configure
Add an [archive] section to the root config, next to chain = "base" and your contracts. Paths are relative to the root config file. Like the other globals, it cannot go in a fragment.
chain = "base" [archive] manifest = "base-manifest.json" manifest_sha256 = "<64-hex-digit-digest>" end_block = 500999 checkpoint_hash = "0x<trusted-end-block-hash>" staging_dir = "archive-stage" max_staging_bytes = 8000000000 max_download_bytes = 300000000 handoff = false
The budget values above fit only a couple of small early groups. Take yours from archive-plan. Sieve refuses to start if the planned staging requirement (peak files plus working space and reserve) does not fit max_staging_bytes or the selected archives do not fit max_download_bytes.
Required keys
| Key | Meaning |
|---|---|
| manifest | Local copy of the snapshot's manifest.json |
| manifest_sha256 | SHA-256 of the exact manifest bytes, 64 hex digits, no 0x |
| end_block | Last block imported from the archive (inclusive) |
| checkpoint_hash | Trusted block hash at end_block |
| staging_dir | Dedicated directory for downloads and retained headers |
| max_staging_bytes | Hard cap on staging disk use |
| max_download_bytes | Cap on total bytes downloaded, retries included |
Optional keys
| Key | Meaning |
|---|---|
| handoff | Continue through peer history into live follow |
| handoff_retry_secs | Wait between handoff attempts |
| min_free_bytes | Free-space reserve (5 GiB) |
| working_space_bytes | Working allowance (1 GiB) |
| retries | Download retries per archive |
| timeout_secs | Per-request timeout |
| max_transactions_per_block | Sanity limit while decoding |
3. Run
Finite import
sieve --config sieve.toml --start-block 499000 --end-block 500999
The start comes from your contract, factory, and transfer start blocks, or --start-block. With handoff = false, --end-block is optional and, if given, must equal archive.end_block. The run exits once that block is committed and cleanup is done. It never connects to peers or starts the API. Add --verbose for per-component and per-batch progress.
Archive to live follow
Set handoff = true and pick a recent snapshot endpoint that ordinary Base peers can bridge.
# archive, then peer history, then live follow sieve --config sieve.toml # archive, then peer history, stop at this height sieve --config sieve.toml --end-block <HEIGHT_AFTER_ARCHIVE_END>
An end above archive.end_block bounds the peer tail. An end equal to it does only the import. An end below it is rejected.
The archive is downloaded and indexed without any peers. After that, Sieve asks peers only for the first unindexed block, end_block + 1. Its header needs the usual three-peer quorum, its parent must equal your trusted archive hash, and its body and receipts must pass the normal payload checks. From there it is ordinary P2P catch-up and then follow mode, all in the same process with the same checkpoint.
If no peer can serve that next block yet, Sieve keeps the finished import and retries every handoff_retry_secs. It never skips missing blocks to reach a newer height. Ctrl+C stops the wait, and the next run picks up from the saved progress. Once the import finishes, an enabled API serves the imported data while the handoff waits.
The archive and peer catch-up phases count as backfill for streams with backfill = false. Follow mode uses the live policy.
What is verified
The manifest must match the SHA-256 you supply. Every extracted file must match the size and BLAKE3 hash the manifest declares.
The header chain is authenticated link by link, from the first selected group's boundary up to your trusted end hash, before a single payload is indexed.
Each reconstructed block is checked against its authenticated header: transaction and receipt alignment, transaction types, cumulative gas, transaction and receipt roots, ommers, and Base withdrawal rules.
Archive blocks go through the same ordered factory discovery, filters, decoders, and transactional writer as P2P blocks. Indexed rows, block hashes, checkpoint, and archive progress commit together.
The manifest digest and the end-block hash are your trust inputs. Get both from sources you trust, such as a block explorer and a node you run. Checksums and header links prove the files match that anchor. They do not prove OP derivation or L1 finality. A database whose frontier came from an archive records that provenance separately from peer-quorum verification.
Resume and cleanup
Rerun the same command after an interruption. PostgreSQL is the source of truth: archive progress commits in the same transaction as indexed rows, block hashes, factory coverage, and the checkpoint. An interrupted group is rebuilt from its boundary and indexing resumes at the next uncommitted block. Completed groups only finish any pending cleanup.
Only one payload group is staged at a time. Its transaction and receipt files are deleted after its last requested block commits. Header files are kept, so retained headers grow with the selected range.
The retained header files, manifest, and job identity authenticate every restart, including after handoff. Keep them, and keep the [archive] section in the config. Removing it is not a supported restart path.
The job binds the manifest, end hash, requested range, indexing config, and ABI contents. Changing any of them refuses resume. Transfer limits, retries, timeout, and worker count can change between runs.
The staging directory must be dedicated: Sieve refuses to adopt an unrelated non-empty directory and holds an exclusive lock on it. The database is also locked to a single Sieve writer.
Interrupted transfers are not resumed mid-file. Every attempt is charged in full against max_download_bytes, across retries and restarts. Raise it if retries use up the budget.
Measured on Base
One run on Base mainnet (chain 8453), indexing the WETH/AAVE Uniswap v3 pool from the archive through peers into live follow.
| Archive import | 7 days, 302,400 blocks |
| Peer history | 31,336 blocks, then live follow |
| Chain | 333,736 consecutive blocks, every parent hash linked, including the seam from archive block 51,968,527 to peer block 51,968,528 |
| Pool events | 9,588 (Swap 7,401, Burn 838, Collect 832, Mint 481, CollectProtocol 36) |
The archive anchor, the first peer block, and the decoded pool events around the seam were checked against an independent RPC. That RPC was used only to audit the result, not to collect it.
Supported sources
Base mainnet only. Sieve reads Base V2 aligned static-file snapshots from the producers 2.5.2-dev (76a8261) and 2.5.2-dev (5877708). Any other producer, layout, or chain is rejected before anything is downloaded. Unsafe archive paths, links, size or hash mismatches, and truncated streams stop the import.
The end-block hash is always explicit. Sieve does not look up checkpoints on its own and never swaps the snapshot under a running job.