About
I make cloud infrastructure reproducible, well-monitored, and dependable in production.
I'm a Senior DevOps and Site Reliability Engineer with over five years across cloud infrastructure, CI/CD, Kubernetes, and observability. I started in SRE — administering Splunk Cloud, running enterprise data migrations, and handling encryption, certificate lifecycles, and on-call support.
From there I moved into DevOps, where I provision infrastructure with Terraform, automate delivery through Jenkins and GitHub Actions, and run workloads on Kubernetes with Helm — so releases stay predictable and incidents stay short.
What I bring
Standardized AWS and Azure provisioning with reusable Terraform modules and remote state — cutting provisioning time ~35% and manual effort ~40%.
Jenkins CI/CD automating build, test, Docker, and Kubernetes deploys — reducing deployment time ~30% through pipeline parallelization.
Proactive monitoring and structured RCA across Splunk, SignalFx, and CloudWatch — improving MTTR ~35% and incident detection ~30%.
Experience
Cloud infrastructure and reliability work across banking, AML, and observability platforms.
Banking & AML platform
Client project · AML Partners (DevOps)
Splunk Cloud · government cloud
Client project · Splunk Observability (Wingman / on-call)
Selected automation
A few of the cloud automation tools I've built end to end across AWS and Azure.
AWS Lambda functions on EventBridge cron schedules that start and stop EC2 instances by tag (Type / Env) — powering off non-production environments outside working hours to cut compute spend, with least-privilege IAM.
A Lambda that finds the newest AMI, updates the EC2 launch template to it, promotes it to the default version, and prunes old versions — keeping Auto Scaling Group refreshes always on the latest image.
Lambda functions that create daily AMIs of EC2 instances tagged for backup, applying a retention policy and a delete-on date, then deregister expired AMIs and delete their snapshots. Reworked with boto3 paginators to handle accounts with 1,000+ snapshots.
A Lambda that copies AMIs to a second AWS region for disaster recovery, tagging each copy with an expiry date — paired with a cleanup function that deregisters expired copies and their snapshots on schedule.
A secure pipeline moving SQL backups and client files from AWS S3 to Azure Blob Storage with AzCopy — GPG/PGP encryption in transit and SAS-token, IP-allowlisted access control.
A scheduled PowerShell job that scans nightly batch logs for exception patterns and sends structured HTML alerts through AWS SES on failure — replacing manual log checks with proactive notifications via Windows Task Scheduler.
An automated ingestion pipeline on AWS Transfer Family: partners push GPG-encrypted files over SFTP into an encrypted S3 bucket, where a Lambda decrypts and routes them downstream — encrypted in transit and at rest, with IP-allowlisted access.
Toolbox