Should Your Raw Instrument Zone Become Iceberg?

Table of Contents

The science team asked a reasonable question: “Can we query two years of run metadata without a week of Athena scans?” The raw prefix held forty million objects, LIST calls were timing out, and every dashboard paged through keys like it was 2013.

S3 Tables—managed Apache Iceberg—look like the answer, and often are. But “make raw a table” is usually the wrong first move. This post is a placement checklist: what becomes a table, what stays an object, and how to migrate without breaking the audit story. Assumes the ingestion patterns post.

Objects and tables solve different problems

PropertyRaw objects (raw/)Iceberg tables (S3 Tables)
Immutability of first captureYes—required for original + attributableNo—tables are mutable by design (snapshots)
Row-level query/filterPainful (partition scans)Native (partition pruning, stats)
Schema evolutionManual reprocessingBuilt-in
Time travel / historyVia object versionsVia snapshots—cheap and queryable
Managed compactionn/aHandled by S3 Tables

Rule of thumb: originals stay objects; anything you derive, join, or filter repeatedly becomes a table.

The pattern that works: raw objects + Iceberg index

Keep raw/{instrument}/{run}/... untouched and versioned. Then maintain tables that point into it:

  1. Run index table—one row per run: ids, timestamps, file counts, bytes, marker status. Replaces LIST-based dashboards entirely.
  2. Manifest/QC table—per-file checksums, validation outcomes, quarantine reasons. Your audit queries become SQL.
  3. Parsed metrics tables—whatever consumers recompute today from raw parses; computed once, versioned by snapshot.

The raw zone remains the system of record; tables are the queryable projection. When a table is wrong, you rebuild it from raw—enduring by construction.

Why S3 Tables specifically

  • Managed compaction and snapshot maintenance—one less Spark cluster you own.
  • Native access from Athena, Glue, EMR, and SageMaker—no custom catalog glue.
  • Replication across regions/accounts and Intelligent-Tiering for table storage (both added at re:Invent 2025)—matters for the endurance story.
  • Storage Lens visibility into table growth—hybrid systems grow quietly.

Migration checklist

  • Start with the run index table only; one writer, rebuilt from manifests. Prove query latency before committing to more.
  • Partition by instrument_id + date, not by run_id—run-id partitioning explodes small-file counts.
  • Define snapshot expiry and retention deliberately; snapshots are your time travel, expiry is your cost valve.
  • Table writes go through the same CI/CD and review as everything else—derived does not mean unregulated.
  • Document the rebuild procedure: table from raw, end to end, timed. A table you can’t rebuild is a liability.
  • Keep an eye on per-request costs—row-level access changes the request profile vs plain S3.

Failure modes

SymptomLikely cause
Queries slow again after monthsSmall-file accumulation in tables—check compaction
Table and raw disagreeBackfill job skipped the validation gate
“Original” questioned in auditSomeone started writing analysis results back into raw/

If you only do one thing: build a run index table over your raw zone this week—it is the highest-leverage table, it replaces your worst dashboards, and it teaches the team Iceberg before anything critical depends on it.

Next: CI/CD under GxP change control.

Previous: Instrument ingestion patterns.

Share :

Related Posts

Observing and Debugging Hybrid Data Transfers

Slack said “files aren’t in S3.” The agent host said “fine.” CloudWatch had three log groups and no shared id. Twenty minutes later someone found the run under a different instrument prefix—and a completion marker that never wrote.

Read More

The ALCOA+ Field Guide for Platform Engineers

The auditor asked a simple question: “How do you know this file was not modified after upload?” The team had versioning off, a shared transfer credential, and a fourteen-second silence that felt much longer.

Read More

Hybrid Lab-to-Cloud: Map the Data Path Before Building It

The transfer agent was “done.” S3 had objects. Downstream Batch jobs still saw empty prefixes for hours. Nobody could say whether the instrument was late, the poller was stuck, or a completion marker never arrived.

Read More