Hybrid Lab-to-Cloud: a Go-Live and Recovery Checklist

Table of Contents

The pilot worked for one instrument on a quiet Tuesday. Cutover week added three instruments, a VPN blip, and a KMS key rotation—and the only recovery plan was “call the person who built the agent.”

This is the closing post of the hybrid lab-to-cloud operations series. It assumes you have a mapped path, reliable S3 landing, IAM/KMS boundaries, idempotent retries, and observability. The goal is a go-live and recovery checklist you can hand to a teammate before cutover.

Readiness review (gate to production)

Do not schedule cutover until these are true.

Contracts

  • Source done definition and cloud completion marker documented and tested.
  • Object key scheme and run_id rules agreed with consumers.
  • SLO for lag (lab done → marker) written and accepted by lab + platform.

Security and compliance

  • Transfer, validator, and consumer identities are separate; no shared long-lived admin keys in the path.
  • SSE-KMS (if required) proven with prod-like roles in a pre-prod bucket.
  • Data classification, retention, and residency constraints recorded.
  • Secrets rotation owner and procedure named.

Reliability

  • State machine persisted; retry budget and quarantine path exercised.
  • Agent kill/restart and dual-marker tests passed.
  • Incomplete multipart lifecycle rule enabled.

Operability

  • Dashboards and alarms for stale pending, quarantine spike, and marker lag.
  • Runbook linked from the alert (not only from the wiki homepage).
  • On-call knows lab contact path for after-hours instrument issues.

Capacity and rehearsal

  • Load test at peak file count and peak bytes, not average day.
  • Rehearsal includes at least one injected failure (IAM deny, truncated file, agent restart).
  • Consumer path (Batch or other) tested on validated markers only.
  • Time-to-detect and time-to-mitigate measured during rehearsal; update SLO/runbook if fantasy numbers were used.

Cutover checklist

Day of go-live:

  1. Freeze non-essential changes to IAM, key policy, and agent config.
  2. Confirm monitoring green on pre-prod synthetic or canary instrument.
  3. Enable production instrument(s) one at a time when possible.
  4. Watch lag, quarantine, and marker metrics for the first full production window.
  5. Keep lab retention intact until the first N production runs validate end-to-end.
  6. Record actual cutover time and versions (agent, Terraform, job definitions).

Rollback

Define rollback before you need it:

TriggerAction
Sustained SLO breachPause new instruments; keep agent draining safe work
Systemic AccessDenied / KMSRevert last IAM/key change via IaC; do not “fix” with *
Data corruption patternStop consumers; quarantine; preserve raw and lab copies
Agent bugPin previous agent version; disable auto-update

Rollback is not always “turn off S3.” Often it is stop consumers and stop deleting lab copies while you repair the hop that broke.

Recovery and replay

  • Document how to replay a run_id (re-upload vs re-validate vs re-run consumer only).
  • Preserve evidence: logs, reason.json, object versions if enabled.
  • RPO/RTO for the landing zone stated in writing (even if imperfect).
  • Backup/lifecycle on control-plane state (agent DB, manifests) owned like production data.

Ownership (RACI lite)

AreaResponsibleConsulted
Instrument / source shareLab opsPlatform
Transfer agentPlatformLab ops
Bucket, KMS, IAMPlatform / securityCompliance
Consumer pipelinesData / science platformPlatform
Alert responseOn-call rotationLab contact

If two rows share the same unnamed human, you do not have a platform—you have a hero.

Recovery exercises

Run on a calendar, not only after incidents:

  • Quarterly: restore or rebuild agent state and prove catch-up.
  • Quarterly: revoke and re-issue transfer credentials in non-prod, then prod change window.
  • After major Terraform or key changes: repeat Put/Get/Decrypt proof.
  • Tabletop: VPN down for two hours—what pages, what pauses, what communicates to lab.

Series map

PostFocus
Map the data pathDiscovery, ownership, failure domains
Move data to S3 reliablyUpload integrity and markers
IAM and KMS boundaryIdentity and encryption
Idempotent pollersRetries and quarantine
Observe and debugCorrelation and first 15 minutes
This postGo-live and recovery

Related platform-notes: Terraform layout, Logs Insights, AWS Batch checklist.

Phase 2 begins here: the story continues with running the platform after go-live—starting with the ALCOA+ field guide for platform engineers.


If you only do one thing: refuse cutover until a named rollback, a replay procedure for one run_id, and a stale-pending alarm exist in writing—pilot success on a quiet day is not a go-live plan.

Share :

Related Posts

AWS Batch on EC2: a Deployment Checklist

This post is for teams running compute-heavy workloads on AWS Batch with EC2—genomics pipelines, media processing, ETL, simulations—not for “hello world” Fargate tutorials. It assumes you already have a container image, a VPC, and a reason to use Batch instead of ECS or Lambda.

Read More

GitHub Actions to ECR to ECS: a Deployment Checklist

This post is for teams already running workloads on Amazon ECS who want a reliable GitHub Actions pipeline without reinventing the wheel each release. It is not a greenfield Kubernetes guide, and it assumes you have a working cluster, service, and task definition—not a blank AWS account.

Read More

Terraform Module Layout I Use for Multi-Environment AWS

This post is for teams already running Terraform in AWS who have felt the pain of three “similar but not quite the same” environment folders. It is not a Terraform 101 tutorial, and it assumes you know what a module, variable, and remote state backend are.

Read More