Senior DevOps / SRE

Dhruvil Raithatha

Five years building and operating cloud infrastructure on AWS and Azure for banking, AML, and observability platforms.

LinkedIn
5+ yrs
In production
99.9%
EKS availability
150+
Incidents resolved
~40%
Faster deploys

About

I make cloud infrastructure reproducible, well-monitored, and dependable in production.

I'm a Senior DevOps and Site Reliability Engineer with over five years across cloud infrastructure, CI/CD, Kubernetes, and observability. I started in SRE — administering Splunk Cloud, running enterprise data migrations, and handling encryption, certificate lifecycles, and on-call support.

From there I moved into DevOps, where I provision infrastructure with Terraform, automate delivery through Jenkins and GitHub Actions, and run workloads on Kubernetes with Helm — so releases stay predictable and incidents stay short.

What I bring

Highlights

Multi-cloud delivery

Standardized AWS and Azure provisioning with reusable Terraform modules and remote state — cutting provisioning time ~35% and manual effort ~40%.

Faster pipelines

Jenkins CI/CD automating build, test, Docker, and Kubernetes deploys — reducing deployment time ~30% through pipeline parallelization.

Reliability at scale

Proactive monitoring and structured RCA across Splunk, SignalFx, and CloudWatch — improving MTTR ~35% and incident detection ~30%.

Experience

Where I've worked

Cloud infrastructure and reliability work across banking, AML, and observability platforms.

Senior DevOps EngineerJan 2025 — Present

Banking & AML platform

  • Own full IaC buildout across AWS and Azure using Terraform for automated, repeatable provisioning.
  • Design and maintain CI/CD pipelines, accelerating delivery cycles by ~30% and cutting deployment time by ~40%.
  • Manage containerized workloads on EKS, sustaining 99.9% availability through proactive monitoring and capacity optimization.
  • Resolved 150+ production incidents across multi-client environments, reducing mean downtime by ~25% via structured root cause analysis.
  • Architected Splunk HEC log-ingestion workflows and proposed a Claude API chatbot stack (React / Node.js / PostgreSQL) for onboarding automation.

Client project · AML Partners (DevOps)

  • Streamlined infrastructure automation and patch compliance across Linux distributions using Jenkins & Terraform.
  • Eliminated ~85% of manual tasks while maintaining 100% adherence to security SLAs.
Senior Site Reliability EngineerNov 2020 — Dec 2024

Splunk Cloud · government cloud

  • Administered Splunk Cloud on government cloud platforms — complex data migrations, platform upgrades, and 24/7 on-call duties.
  • Automated infrastructure and CI/CD using Terraform, Ansible, Puppet, Jenkins, and GitLab, boosting delivery speed by ~50%.
  • Orchestrated EKS and Docker workloads, cutting deployment runtime and cloud infrastructure spend by ~40%.
  • Implemented automated DB replication, backups, and DR recovery plans, decreasing downtime by ~15%.
  • Led post-incident reviews with Splunk leadership and key stakeholders, reducing critical issue resolution time by ~15%.

Client project · Splunk Observability (Wingman / on-call)

  • Reduced cloud costs by ~30% through resource optimization.
  • Elevated uptime by ~20% and cut operational manual effort by ~50% using automated workflow tooling.

Selected automation

Things I've built

A few of the cloud automation tools I've built end to end across AWS and Azure.

Scheduled EC2 start / stop

AWS Lambda functions on EventBridge cron schedules that start and stop EC2 instances by tag (Type / Env) — powering off non-production environments outside working hours to cut compute spend, with least-privilege IAM.

LambdaEventBridgePythonboto3IAM

Launch template auto-update

A Lambda that finds the newest AMI, updates the EC2 launch template to it, promotes it to the default version, and prunes old versions — keeping Auto Scaling Group refreshes always on the latest image.

LambdaEC2AMIAuto Scaling

Tag-driven AMI backups & cleanup

Lambda functions that create daily AMIs of EC2 instances tagged for backup, applying a retention policy and a delete-on date, then deregister expired AMIs and delete their snapshots. Reworked with boto3 paginators to handle accounts with 1,000+ snapshots.

LambdaEC2AMIboto3CloudWatch Events

Cross-region AMI replication

A Lambda that copies AMIs to a second AWS region for disaster recovery, tagging each copy with an expiry date — paired with a cleanup function that deregisters expired copies and their snapshots on schedule.

LambdaAMIMulti-regionDR

Cross-cloud data migration

A secure pipeline moving SQL backups and client files from AWS S3 to Azure Blob Storage with AzCopy — GPG/PGP encryption in transit and SAS-token, IP-allowlisted access control.

AWS S3Azure BlobAzCopyGPG / PGP

Log monitoring & email alerts

A scheduled PowerShell job that scans nightly batch logs for exception patterns and sends structured HTML alerts through AWS SES on failure — replacing manual log checks with proactive notifications via Windows Task Scheduler.

PowerShellAWS SESTask SchedulerIAM

Secure SFTP file ingestion

An automated ingestion pipeline on AWS Transfer Family: partners push GPG-encrypted files over SFTP into an encrypted S3 bucket, where a Lambda decrypts and routes them downstream — encrypted in transit and at rest, with IP-allowlisted access.

AWS Transfer FamilySFTPLambdaS3GPG

Toolbox

Skills

CloudAWSAzureGCP
IaCTerraformAnsibleRemote stateEnv management
ContainersDockerKubernetes (EKS)HelmArgoCD
CI / CDJenkinsGitHub ActionsAzure DevOpsRelease automation
ObservabilitySplunkELK StackSignalFxCloudWatchPrometheusGrafana
ScriptingPythonBash
Config & VCSPuppetGitGitHubGitLab
SecuritySecrets ManagerKMSSSL/TLSAccess control
CollaborationJiraConfluence