Base Archive Bootstrap

Index Base history that peers no longer serve, straight from the official snapshot archives.

Why

Many Base peers prune old receipts, so a deep backfill over P2P alone often has no one to ask. Base publishes V2 history snapshots split into aligned archive groups. Sieve can read those directly: it downloads one group at a time, indexes only what your config asks for, deletes the group's transaction and receipt files, and moves on.

No history node, no execution client, no second database. The normal sieve command does the import, and with handoff = true the same process continues over P2P into live follow.

flow
archive groups (download, verify, index, clean up)
       |
       v
peer history (from archive end + 1)
       |
       v
live follow

1. Plan the import

Download the snapshot's small manifest.json and get its SHA-256 from a source you trust. Then run the offline planner on your exact range:

terminal
sieve archive-plan --manifest base-manifest.json \
  --manifest-sha256 <64-hex-digit-digest> \
  --start-block 500123 --end-block 1000010

It reads only the local manifest. No config, database, peers, or network. The JSON plan on stdout lists the selected archive groups and their files, which blocks are decoded only as dependencies, total download and extracted bytes, the peak staging space needed, and the trust inputs still required. Blocks before your start inside the first group are decoded to establish transaction positions but are never indexed.

Size your budgets from the plan. Early groups are small. Recent groups are large: the group covering blocks 51.5M to 52M is about 25 GB compressed. Staging is temporary and released per group: once a group is indexed, its transaction and receipt files are deleted and only its headers are kept. archive-plan reports the exact download and staging needs for your range. See CLI for every flag.

2. Configure

Add an [archive] section to the root config, next to chain = "base" and your contracts. Paths are relative to the root config file. Like the other globals, it cannot go in a fragment.

sieve.toml
chain = "base"

[archive]
manifest = "base-manifest.json"
manifest_sha256 = "<64-hex-digit-digest>"
end_block = 500999
checkpoint_hash = "0x<trusted-end-block-hash>"
staging_dir = "archive-stage"
max_staging_bytes = 8000000000
max_download_bytes = 300000000
handoff = false

The budget values above fit only a couple of small early groups. Take yours from archive-plan. Sieve refuses to start if the planned staging requirement (peak files plus working space and reserve) does not fit max_staging_bytes or the selected archives do not fit max_download_bytes.

Required keys

KeyMeaning
manifestLocal copy of the snapshot's manifest.json
manifest_sha256SHA-256 of the exact manifest bytes, 64 hex digits, no 0x
end_blockLast block imported from the archive (inclusive)
checkpoint_hashTrusted block hash at end_block
staging_dirDedicated directory for downloads and retained headers
max_staging_bytesHard cap on staging disk use
max_download_bytesCap on total bytes downloaded, retries included

Optional keys

KeyMeaning
handoffContinue through peer history into live follow
handoff_retry_secsWait between handoff attempts
min_free_bytesFree-space reserve (5 GiB)
working_space_bytesWorking allowance (1 GiB)
retriesDownload retries per archive
timeout_secsPer-request timeout
max_transactions_per_blockSanity limit while decoding

3. Run

Finite import

terminal
sieve --config sieve.toml --start-block 499000 --end-block 500999

The start comes from your contract, factory, and transfer start blocks, or --start-block. With handoff = false, --end-block is optional and, if given, must equal archive.end_block. The run exits once that block is committed and cleanup is done. It never connects to peers or starts the API. Add --verbose for per-component and per-batch progress.

Archive to live follow

Set handoff = true and pick a recent snapshot endpoint that ordinary Base peers can bridge.

terminal
# archive, then peer history, then live follow
sieve --config sieve.toml

# archive, then peer history, stop at this height
sieve --config sieve.toml --end-block <HEIGHT_AFTER_ARCHIVE_END>

An end above archive.end_block bounds the peer tail. An end equal to it does only the import. An end below it is rejected.

The archive is downloaded and indexed without any peers. After that, Sieve asks peers only for the first unindexed block, end_block + 1. Its header needs the usual three-peer quorum, its parent must equal your trusted archive hash, and its body and receipts must pass the normal payload checks. From there it is ordinary P2P catch-up and then follow mode, all in the same process with the same checkpoint.

If no peer can serve that next block yet, Sieve keeps the finished import and retries every handoff_retry_secs. It never skips missing blocks to reach a newer height. Ctrl+C stops the wait, and the next run picks up from the saved progress. Once the import finishes, an enabled API serves the imported data while the handoff waits.

The archive and peer catch-up phases count as backfill for streams with backfill = false. Follow mode uses the live policy.

What is verified

▸ Pinned manifest

The manifest must match the SHA-256 you supply. Every extracted file must match the size and BLAKE3 hash the manifest declares.

▸ Headers first

The header chain is authenticated link by link, from the first selected group's boundary up to your trusted end hash, before a single payload is indexed.

▸ Payload checks

Each reconstructed block is checked against its authenticated header: transaction and receipt alignment, transaction types, cumulative gas, transaction and receipt roots, ommers, and Base withdrawal rules.

▸ Same pipeline

Archive blocks go through the same ordered factory discovery, filters, decoders, and transactional writer as P2P blocks. Indexed rows, block hashes, checkpoint, and archive progress commit together.

The manifest digest and the end-block hash are your trust inputs. Get both from sources you trust, such as a block explorer and a node you run. Checksums and header links prove the files match that anchor. They do not prove OP derivation or L1 finality. A database whose frontier came from an archive records that provenance separately from peer-quorum verification.

Resume and cleanup

Rerun the same command after an interruption. PostgreSQL is the source of truth: archive progress commits in the same transaction as indexed rows, block hashes, factory coverage, and the checkpoint. An interrupted group is rebuilt from its boundary and indexing resumes at the next uncommitted block. Completed groups only finish any pending cleanup.

▸ Rolling disk use

Only one payload group is staged at a time. Its transaction and receipt files are deleted after its last requested block commits. Header files are kept, so retained headers grow with the selected range.

▸ Keep the staging directory

The retained header files, manifest, and job identity authenticate every restart, including after handoff. Keep them, and keep the [archive] section in the config. Removing it is not a supported restart path.

▸ The job is pinned

The job binds the manifest, end hash, requested range, indexing config, and ABI contents. Changing any of them refuses resume. Transfer limits, retries, timeout, and worker count can change between runs.

▸ One writer

The staging directory must be dedicated: Sieve refuses to adopt an unrelated non-empty directory and holds an exclusive lock on it. The database is also locked to a single Sieve writer.

▸ Downloads restart from zero

Interrupted transfers are not resumed mid-file. Every attempt is charged in full against max_download_bytes, across retries and restarts. Raise it if retries use up the budget.

Measured on Base

One run on Base mainnet (chain 8453), indexing the WETH/AAVE Uniswap v3 pool from the archive through peers into live follow.

Archive import7 days, 302,400 blocks
Peer history31,336 blocks, then live follow
Chain333,736 consecutive blocks, every parent hash linked, including the seam from archive block 51,968,527 to peer block 51,968,528
Pool events9,588 (Swap 7,401, Burn 838, Collect 832, Mint 481, CollectProtocol 36)

The archive anchor, the first peer block, and the decoded pool events around the seam were checked against an independent RPC. That RPC was used only to audit the result, not to collect it.

Supported sources

Base mainnet only. Sieve reads Base V2 aligned static-file snapshots from the producers 2.5.2-dev (76a8261) and 2.5.2-dev (5877708). Any other producer, layout, or chain is rejected before anything is downloaded. Unsafe archive paths, links, size or hash mismatches, and truncated streams stop the import.

The end-block hash is always explicit. Sieve does not look up checkpoints on its own and never swaps the snapshot under a running job.

Next steps

▸ Chains Base peers, trusted peers, and discovery
▸ Sync & Follow what happens after the archive