Should Your Raw Instrument Zone Become Iceberg?
- Shameem Abdul Salam
- Aws , Dev ops
- September 24, 2026
Table of Contents
The science team asked a reasonable question: “Can we query two years of run metadata without a week of Athena scans?” The raw prefix held forty million objects, LIST calls were timing out, and every dashboard paged through keys like it was 2013.
S3 Tables—managed Apache Iceberg—look like the answer, and often are. But “make raw a table” is usually the wrong first move. This post is a placement checklist: what becomes a table, what stays an object, and how to migrate without breaking the audit story. Assumes the ingestion patterns post.
Objects and tables solve different problems
| Property | Raw objects (raw/) | Iceberg tables (S3 Tables) |
|---|---|---|
| Immutability of first capture | Yes—required for original + attributable | No—tables are mutable by design (snapshots) |
| Row-level query/filter | Painful (partition scans) | Native (partition pruning, stats) |
| Schema evolution | Manual reprocessing | Built-in |
| Time travel / history | Via object versions | Via snapshots—cheap and queryable |
| Managed compaction | n/a | Handled by S3 Tables |
Rule of thumb: originals stay objects; anything you derive, join, or filter repeatedly becomes a table.
The pattern that works: raw objects + Iceberg index
Keep raw/{instrument}/{run}/... untouched and versioned. Then maintain tables that point into it:
- Run index table—one row per run: ids, timestamps, file counts, bytes, marker status. Replaces
LIST-based dashboards entirely. - Manifest/QC table—per-file checksums, validation outcomes, quarantine reasons. Your audit queries become SQL.
- Parsed metrics tables—whatever consumers recompute today from raw parses; computed once, versioned by snapshot.
The raw zone remains the system of record; tables are the queryable projection. When a table is wrong, you rebuild it from raw—enduring by construction.
Why S3 Tables specifically
- Managed compaction and snapshot maintenance—one less Spark cluster you own.
- Native access from Athena, Glue, EMR, and SageMaker—no custom catalog glue.
- Replication across regions/accounts and Intelligent-Tiering for table storage (both added at re:Invent 2025)—matters for the endurance story.
- Storage Lens visibility into table growth—hybrid systems grow quietly.
Migration checklist
- Start with the run index table only; one writer, rebuilt from manifests. Prove query latency before committing to more.
- Partition by
instrument_id+ date, not byrun_id—run-id partitioning explodes small-file counts. - Define snapshot expiry and retention deliberately; snapshots are your time travel, expiry is your cost valve.
- Table writes go through the same CI/CD and review as everything else—derived does not mean unregulated.
- Document the rebuild procedure: table from raw, end to end, timed. A table you can’t rebuild is a liability.
- Keep an eye on per-request costs—row-level access changes the request profile vs plain S3.
Failure modes
| Symptom | Likely cause |
|---|---|
| Queries slow again after months | Small-file accumulation in tables—check compaction |
| Table and raw disagree | Backfill job skipped the validation gate |
| “Original” questioned in audit | Someone started writing analysis results back into raw/ |
If you only do one thing: build a run index table over your raw zone this week—it is the highest-leverage table, it replaces your worst dashboards, and it teaches the team Iceberg before anything critical depends on it.
Next: CI/CD under GxP change control.
Previous: Instrument ingestion patterns.