Open to New Roles

Senior Cloud
Data Platform
Engineer

Python · SQL · PySpark · Apache Spark · AWS Glue · Airflow · Databricks · Snowflake · Redshift · AWS Bedrock

6+ years engineering and operating production cloud data platforms across financial services and healthcare, including governed data services that support analytics, machine learning, and generative AI. I specialize in secure ETL/ELT pipelines, AI/ML-ready datasets, modern data platforms, data quality, and reliable cloud infrastructure using Python, SQL, PySpark, AWS Glue, Airflow, Databricks, Snowflake, Redshift, AWS, and Azure.

Location
San Francisco Bay Area
Availability
Immediate · Open to Relocation
Industries
Banking · Healthcare · Tech
Experience
6+ Years
99.9%
Production Data Platform Uptime
Production ingestion services and ML inference APIs sustained through resilient EKS deployment patterns
45→15min
MTTR Reduced by 67%
Full-stack observability with Prometheus, Grafana, ELK & CloudWatch — alerts before users notice
~20%
Cloud Cost Reduction
Snowflake + Redshift right-sizing and auto-suspend strategies with zero SLA impact
85→98%
Deployment Reliability
Standardized container builds, Helm releases, and GitOps delivery across data and ML workloads
About Me

Enterprise data needs a
reliable platform.

I build secure, governed cloud data platforms that power analytics, machine learning, generative AI, and regulatory reporting.

At Fifth Third Bank, I engineer regulated ingestion pipelines, curated analytics layers, and governed feature and document datasets for AWS Bedrock-backed risk and compliance workflows. At CVS Health, I developed Azure Data Factory and Databricks pipelines that delivered validated healthcare claims data to analytics and ML consumers.

My strength is at the intersection of data engineering, AI-ready data delivery, and platform reliability: I design pipelines, data models, validation controls, governed datasets, cloud infrastructure, CI/CD, observability, and security as one integrated production system.

Current Role
Cloud Data Platform Engineer · Fifth Third Bank
Cloud Platforms
AWS · Azure
Core Stack
Python · SQL · PySpark · Airflow · Glue · Databricks · Snowflake · Redshift · Bedrock
Compliance Exp.
PCI-DSS · HIPAA · SOC2 — production, not lab
Certifications
4 certifications — AWS, Azure, HashiCorp Terraform, Databricks
Education
M.S. Computer Science · University of Central Missouri
Location
San Francisco Bay Area · Open to relocation · Remote, hybrid, or onsite
Contact
bsandya1926@gmail.com · +1 913-335-0346
Core Specializations

Six domains I deliver end-to-end.

Core data and platform engineering capabilities applied across regulated financial services and healthcare environments.

01
⚙️
Cloud Data Engineering

Production ETL and ELT pipelines using Python, SQL, PySpark, AWS Glue, Apache Airflow, Azure Data Factory, and Apache Spark for analytics and reporting.

PythonSQLPySparkAWS GlueAirflow
02
🤖
Modern Data Platforms

Cloud data warehouses and lakehouse platforms across Snowflake, Redshift, Databricks, and Azure Data Factory, with performance tuning and cost optimization.

SnowflakeRedshiftDatabricksAzure Data Factory
03
🏗️
Data Quality & Reliability

Embedded validation, reconciliation, exception handling, retries, alerts, dependency controls, and production monitoring into data workflows to protect reporting SLAs.

ValidationReconciliationObservabilitySLA Controls
04
🔄
Cloud Infrastructure & Automation

Infrastructure automation and deployment reliability using Terraform, Kubernetes, Docker, Helm, GitHub Actions, Jenkins, Azure DevOps, and ArgoCD.

TerraformKubernetesDockerHelmArgoCD
05
🔒
Data Security & Governance

Security controls for regulated data using IAM, RBAC, KMS, Key Vault, Secrets Manager, private networking, encryption, and audit logging aligned with PCI-DSS, HIPAA, and SOC2.

PCI-DSSHIPAASOC2IAM/RBACKMS
06
📡
AI & Generative AI Data Platforms

Governed feature and document datasets for AWS Bedrock, Databricks ML feature engineering pipelines, SageMaker inference support, model-serving data workflows, and AI/ML workload observability.

AWS BedrockDatabricks MLSageMakerGoverned Datasets
Career History

Where I've delivered.

Four roles across 6+ years, spanning cloud data engineering, analytics platforms, infrastructure automation, and production support.

Fifth Third Bank
CLOUD DATA PLATFORM ENGINEER
Banking · PCI-DSS · SOC2 Feb 2025 – Present 📍 Kentwood, MI
99.9%
EKS Uptime
–67%
MTTR
~20%
Cost Saved
85→98%
Deploy Reliability

I engineer and operate the bank's Enterprise Cloud Data Platform Modernization, building regulated ingestion pipelines, curated analytics layers, reliable orchestration, secure cloud infrastructure, and production observability for treasury, compliance, reporting, and AI-enabled workflows.

AWS Bedrock for LLM risk & compliance workflows — configured IAM-scoped model access, VPC endpoints, and CloudTrail audit logging to keep all inference traffic within private network boundaries per PCI-DSS, enabling risk summarization and compliance document review tooling for analytics teams.
Production EKS platform (3 node groups) — multi-AZ deployment, autoscaling, namespace-level workload isolation for financial ingestion services and containerized ML inference. Sustained 99.9% uptime through peak transaction windows.
Reusable Terraform and Kubernetes controls — reviewed modules, manifests, IAM changes, and platform architecture updates to improve scalability, resilience, security, and governance.
AWS Glue + Apache Airflow ETL pipelines — financial data from PostgreSQL/MySQL into Snowflake and Redshift for treasury and compliance reporting, with the same pipelines supplying curated feature datasets to Bedrock-backed AI workflows.
GitOps CI/CD (GitHub Actions + Jenkins + ArgoCD) — standardized container builds, Helm releases, and controlled deployments across data and ML workloads, improving deployment reliability from 85% to 98%.
Full observability stack (Prometheus + Grafana + ELK + CloudWatch) — covering Airflow DAG failures, pod health, memory pressure, API latency, and inference response times. MTTR improved from 45 minutes to under 15 minutes.
Snowflake + Redshift performance tuning — warehouse auto-suspend and right-sizing strategies reduced compute costs by ~20% without impacting reporting SLAs.
Hybrid cloud collaboration — coordinated with Azure platform teams on integration between AWS-hosted data services and Azure-based analytics and identity management systems, aligning governance and access controls across environments.
Environment
EKSBedrockGlueRedshiftLambdaECSKMSSecrets ManagerTerraformHelmArgoCDGitHub ActionsJenkinsSnowflakeAirflowPostgreSQLPrometheusGrafanaELKCloudWatch
CVS Health
DATA ENGINEER — HEALTHCARE ANALYTICS
Healthcare · HIPAA · SOC2 Sep 2023 – Jan 2025 · 1 yr 5 mo 📍 New York City, NY

I engineered Azure Data Factory and Databricks pipelines for healthcare claims ingestion, actuarial reporting, fraud detection, and utilization prediction, combining PySpark transformations, SQL validation, Snowflake administration, Terraform automation, and HIPAA-aligned security controls.

Secure PHI data processing — protected healthcare datasets with Azure Blob Storage, Key Vault, RBAC, namespace isolation, and network policies aligned with HIPAA and SOC2 controls.
Databricks ML feature pipelines — PySpark transformations supplying structured claims data to fraud detection and utilization prediction models. Tuned Spark cluster configs to meet SLA windows during large file ingestion.
Azure Data Factory + Databricks ETL workflows — large-scale healthcare claims processing pipelines with Azure Key Vault PHI encryption aligned to HIPAA and SOC2, embedded security validation in every release pipeline.
Snowflake administration for actuarial teams — schema management, RBAC, and SQL query performance tuning during monthly reporting cycles. SQL-based data validation and reconciliation to ensure reporting accuracy.
Azure Monitor + Log Analytics — Spark job tracking and AKS utilization monitoring; caught inefficient resource allocation early, adjusted node pool sizing to reduce reruns and cut compute spend.
Environment
Azure AKSAzure Data FactoryBlob StorageKey VaultAzure MonitorRBACVirtual NetworksTerraformAzure DevOpsDatabricksPySparkSnowflakeLog AnalyticsKubernetes
Birlasoft
CLOUD DATA ENGINEER
Technology · AWS Migrations Nov 2020 – Dec 2022 · 2 yrs 2 mo 📍 Hyderabad, India

Supported enterprise data integration and AWS migration workloads through cloud provisioning, Spark batch processing, SQL transformations, Terraform automation, Jenkins delivery pipelines, Docker containerization, and production support.

Reusable Terraform modules for AWS migrations — compute, storage, and networking with remote state and structured versioning, reducing provisioning errors during on-prem to cloud cutovers.
Jenkins CI/CD + Docker containerization — automated builds and staged deployments across Dev and QA with Git and Bitbucket; resolved cross-environment inconsistencies between development, testing, and staging.
Ansible + Shell automation — configuration validation, post-migration task automation, log cleanup, and backup verification. CloudWatch dashboards for cost optimization identifying underutilized EC2 instances.
Apache Spark batch jobs & SQL transformations — processed structured datasets for internal reporting workflows, coordinating with data teams on scheduling and output validation.
Environment
EC2S3IAMCloudWatchTerraformJenkinsDockerAnsibleApache SparkGitBitbucketLinux
Tata Consultancy Services
DATA ENGINEER
Data Engineering · Banking & Retail Jun 2019 – Oct 2020 · 1 yr 5 mo 📍 Hyderabad, India

My data engineering roots — where I learned what financial close cycles really mean for the people running batch jobs at 2am. SQL performance, Python ETL patterns, and production support discipline that now informs every data pipeline I build at the infrastructure level.

Python ETL pipelines for banking & retail — cleansed, transformed, and loaded batch datasets into Oracle and SQL Server; tuned large monthly reporting jobs during financial close cycles to prevent delays.
SQL performance optimization — complex queries, stored procedures, indexing strategies, and execution plan analysis for high-volume reporting systems.
Production data support — investigated reporting discrepancies for finance teams, corrected data mapping issues, monitored Linux batch schedules and resolved ingestion delays.
Environment
PythonSQLOracleSQL ServerPostgreSQLMySQLLinuxGitBatch Processing
Technical Skills

The full technical stack.

Depth ratings reflect daily production use, not tutorials or side projects.

Cloud Data Engineering
Python / SQL
PySpark / Apache Spark
AWS Glue / Apache Airflow
Azure Data Factory / Databricks
ETL / ELT / Data Modeling
Cloud Data Platforms
Snowflake
Amazon Redshift
PostgreSQL / MySQL
Oracle / SQL Server
Query Tuning / Stored Procedures
Cloud & Platform
AWS / Azure
Terraform / Kubernetes
Docker / Helm / ArgoCD
GitHub Actions / Jenkins
Azure DevOps / Ansible
Reliability & Security
Prometheus / Grafana
CloudWatch / Azure Monitor
ELK / Log Analytics
IAM / RBAC / KMS
Key Vault / Secrets Manager
Certifications

Four credentials across
cloud, infrastructure & AI.

Certifications supporting my work across AWS, Azure, infrastructure as code, and generative AI data platforms.

☁️
Amazon Web Services
AWS Certified Solutions Architect – Associate
Cloud Architecture
🔷
Microsoft
Azure Administrator Associate (AZ-104)
Cloud Administration
HashiCorp
Terraform Associate
Infrastructure as Code
🤖
Databricks Academy
Accreditation – Generative AI Fundamentals
Generative AI
Education

Academic foundation.

🎓
Master of Science
Computer Science
University of Central Missouri · USA
Bachelor of Technology
Electronics & Communication Engineering
NBKR Institute of Science & Technology · India
Get in Touch

Ready to build
something that holds.

I'm open to Senior Cloud Data Platform Engineer, Senior Data Engineer, Data Platform Engineer, Cloud Data Engineer, and AI/ML Data Platform roles — especially in regulated industries where data quality, security, and reliability are mission-critical.

Remote, hybrid, or onsite — open to relocation. Based in the San Francisco Bay Area.

Primary Email

Open to Senior Cloud Data Platform Engineer, Senior Data Engineer, Data Platform Engineer, and Cloud Data Engineer roles. I bring hands-on experience across banking and healthcare data platforms, cloud infrastructure, compliance, and production reliability.

Quick Facts
Experience6+ Years
Certifications4
AvailabilityImmediate
Work ModeRemote · Hybrid · Onsite · Relocation OK
LocationSan Francisco Bay Area
San Francisco Bay Area · Open to Relocation · Remote · Hybrid · Onsite