Observability as Audit Evidence
Mid-investigation, someone asked for the transfer logs from eleven weeks ago—the ones that would show exactly what happened to run R. The retention policy was thirty days. The logs that could have answered the question had been rotated into nothing, and the finding closed as “probable, not demonstrable.”
Read MoreValidated Delivery: CI/CD Under GxP Change Control
The validation auditor asked for the change record behind a production change made two weeks earlier. The honest answer was a Slack thread, a terraform apply from someone’s laptop, and a Jira ticket with no attachments. None of that is a change record.
Read MoreShould Your Raw Instrument Zone Become Iceberg?
The science team asked a reasonable question: “Can we query two years of run metadata without a week of Athena scans?” The raw prefix held forty million objects, LIST calls were timing out, and every dashboard paged through keys like it was 2013.
Read MoreInstrument Ingestion Patterns: Poller, DataSync, or S3 Events
Three instruments, three capabilities. One writes finished runs to a Windows share. One can call the S3 API directly. One ships files over SFTP to wherever you point it. The team forced all three through one custom poller—and spent a year paying for it in special cases.
Read MoreThe ALCOA+ Field Guide for Platform Engineers
The auditor asked a simple question: “How do you know this file was not modified after upload?” The team had versioning off, a shared transfer credential, and a fourteen-second silence that felt much longer.
Read MoreAmazon Linux 2 Is EOL: a Migration Checklist for Batch, ECS, and Lambda
The security scan flagged eleven CVEs on the ECS instances. There was nothing to patch with—Amazon Linux 2 went end-of-life on June 30, 2026, and the fixes simply stopped coming. The AMI had been “temporary” for two years.
Read MoreHybrid Lab-to-Cloud: a Go-Live and Recovery Checklist
The pilot worked for one instrument on a quiet Tuesday. Cutover week added three instruments, a VPN blip, and a KMS key rotation—and the only recovery plan was “call the person who built the agent.”
Read MoreObserving and Debugging Hybrid Data Transfers
Slack said “files aren’t in S3.” The agent host said “fine.” CloudWatch had three log groups and no shared id. Twenty minutes later someone found the run under a different instrument prefix—and a completion marker that never wrote.
Read MoreIdempotent Pollers, Retries, and Quarantine Paths
The agent crashed after uploading 40 of 42 files. On restart it uploaded everything again, overwrote nothing useful, emitted two completion markers, and the pipeline ran twice—once on a partial set that somehow got marked ready during the race.
Read MoreIAM and KMS Across the Hybrid Lab–Cloud Boundary
Transfer logs said AccessDenied. The bucket policy “allowed S3.” Someone added s3:* on the role and it still failed—because the objects used SSE-KMS and the key policy never trusted the transfer identity.
Read MoreMoving Instrument Data to S3 Reliably
The object existed in S3. Size looked right. Downstream parsing failed because the transfer cut off mid-write and a later retry never ran—or ran into a different key. Metrics said “uploaded.” Science said “garbage.”
Read MoreHybrid Lab-to-Cloud: Map the Data Path Before Building It
The transfer agent was “done.” S3 had objects. Downstream Batch jobs still saw empty prefixes for hours. Nobody could say whether the instrument was late, the poller was stuck, or a completion marker never arrived.
Read MoreFrom Varnish and Drupal Tuning to Platform Engineering
In 2013 I wrote about Varnish in front of Drupal and where Drupal sites actually lose performance. The stack was Apache, PHP, MySQL, SSH, and a lot of hand-tuning. Today I work on AWS Batch, multi-region Terraform, and hybrid lab-to-cloud integrations in life sciences. The logos changed; the job did not.
Read More