Hybrid Lab-to-Cloud: a Go-Live and Recovery Checklist
- Shameem Abdul Salam
- Dev ops , Aws
- August 27, 2026
Table of Contents
The pilot worked for one instrument on a quiet Tuesday. Cutover week added three instruments, a VPN blip, and a KMS key rotation—and the only recovery plan was “call the person who built the agent.”
This is the closing post of the hybrid lab-to-cloud operations series. It assumes you have a mapped path, reliable S3 landing, IAM/KMS boundaries, idempotent retries, and observability. The goal is a go-live and recovery checklist you can hand to a teammate before cutover.
Readiness review (gate to production)
Do not schedule cutover until these are true.
Contracts
- Source done definition and cloud completion marker documented and tested.
- Object key scheme and
run_idrules agreed with consumers. - SLO for lag (lab done → marker) written and accepted by lab + platform.
Security and compliance
- Transfer, validator, and consumer identities are separate; no shared long-lived admin keys in the path.
- SSE-KMS (if required) proven with prod-like roles in a pre-prod bucket.
- Data classification, retention, and residency constraints recorded.
- Secrets rotation owner and procedure named.
Reliability
- State machine persisted; retry budget and quarantine path exercised.
- Agent kill/restart and dual-marker tests passed.
- Incomplete multipart lifecycle rule enabled.
Operability
- Dashboards and alarms for stale pending, quarantine spike, and marker lag.
- Runbook linked from the alert (not only from the wiki homepage).
- On-call knows lab contact path for after-hours instrument issues.
Capacity and rehearsal
- Load test at peak file count and peak bytes, not average day.
- Rehearsal includes at least one injected failure (IAM deny, truncated file, agent restart).
- Consumer path (Batch or other) tested on validated markers only.
- Time-to-detect and time-to-mitigate measured during rehearsal; update SLO/runbook if fantasy numbers were used.
Cutover checklist
Day of go-live:
- Freeze non-essential changes to IAM, key policy, and agent config.
- Confirm monitoring green on pre-prod synthetic or canary instrument.
- Enable production instrument(s) one at a time when possible.
- Watch lag, quarantine, and marker metrics for the first full production window.
- Keep lab retention intact until the first N production runs validate end-to-end.
- Record actual cutover time and versions (agent, Terraform, job definitions).
Rollback
Define rollback before you need it:
| Trigger | Action |
|---|---|
| Sustained SLO breach | Pause new instruments; keep agent draining safe work |
Systemic AccessDenied / KMS | Revert last IAM/key change via IaC; do not “fix” with * |
| Data corruption pattern | Stop consumers; quarantine; preserve raw and lab copies |
| Agent bug | Pin previous agent version; disable auto-update |
Rollback is not always “turn off S3.” Often it is stop consumers and stop deleting lab copies while you repair the hop that broke.
Recovery and replay
- Document how to replay a
run_id(re-upload vs re-validate vs re-run consumer only). - Preserve evidence: logs,
reason.json, object versions if enabled. - RPO/RTO for the landing zone stated in writing (even if imperfect).
- Backup/lifecycle on control-plane state (agent DB, manifests) owned like production data.
Ownership (RACI lite)
| Area | Responsible | Consulted |
|---|---|---|
| Instrument / source share | Lab ops | Platform |
| Transfer agent | Platform | Lab ops |
| Bucket, KMS, IAM | Platform / security | Compliance |
| Consumer pipelines | Data / science platform | Platform |
| Alert response | On-call rotation | Lab contact |
If two rows share the same unnamed human, you do not have a platform—you have a hero.
Recovery exercises
Run on a calendar, not only after incidents:
- Quarterly: restore or rebuild agent state and prove catch-up.
- Quarterly: revoke and re-issue transfer credentials in non-prod, then prod change window.
- After major Terraform or key changes: repeat Put/Get/Decrypt proof.
- Tabletop: VPN down for two hours—what pages, what pauses, what communicates to lab.
Series map
| Post | Focus |
|---|---|
| Map the data path | Discovery, ownership, failure domains |
| Move data to S3 reliably | Upload integrity and markers |
| IAM and KMS boundary | Identity and encryption |
| Idempotent pollers | Retries and quarantine |
| Observe and debug | Correlation and first 15 minutes |
| This post | Go-live and recovery |
Related platform-notes: Terraform layout, Logs Insights, AWS Batch checklist.
Phase 2 begins here: the story continues with running the platform after go-live—starting with the ALCOA+ field guide for platform engineers.
If you only do one thing: refuse cutover until a named rollback, a replay procedure for one run_id, and a stale-pending alarm exist in writing—pilot success on a quiet day is not a go-live plan.