Summary
What you’ll impact
The Platform Engineer will design, build, and maintain cloud infrastructure and platform services across AWS and Azure, focusing on automation, observability, and security. The role requires deep expertise in containers, data platform tools, infrastructure as code, and technical leadership to guide complex incident resolution and mentor engineering teams.
Responsibilities
What you'll do
- Proven ability to lead technical decisions, coordinate complex incident resolution, mentor engineers, review designs, manage risks, and communicate effectively with senior stakeholders.
Requirements
What you’ll bring
- Experience: 10+ years of progressive experience in platform engineering, cloud infrastructure, DevOps, site reliability engineering, or enterprise production support.
- Cloud Architecture: Advanced hands-on expertise across AWS and Microsoft Azure, including compute, storage, networking, identity, security, governance, monitoring, resilience, and cost optimization.
- Containers and Orchestration: Strong expertise with Kubernetes and Docker, including cluster operations, workload deployment, scaling, upgrades, configuration, security, performance tuning, and complex troubleshooting.
- Data Platform Engineering: Strong operational experience supporting Databricks, dbt, Apache Airflow, AutoSys, and enterprise data processing environments. Experience with LASR and Caspian is preferred.
- Infrastructure as Code and CI/CD: Advanced experience with Terraform and enterprise CI/CD platforms such. g. Harness, Azure DevOps, GitHub Actions, Jenkins, or equivalent, including reusable modules and deployment controls.
- Automation: Strong scripting and automation capability using Python, Shell, Bash, or PowerShell to reduce manual effort and improve platform reliability.
- Observability and SRE: Experience designing monitoring, logging, alerting, service health dashboards, SLI/SLO measures, capacity controls, and operational readiness using Azure Monitor, AWS CloudWatch, Datadog, Splunk, or similar tools.
- Security and Compliance: Strong understanding of IAM/RBAC, secrets management, vulnerability remediation, patching, policy enforcement, audit readiness, business continuity, and disaster recovery.
- ITSM and Agile Delivery: Proficiency with ServiceNow and Jira for incident, problem, change, service request, release, backlog, and continuous-improvement management.
- Technical Leadership: Proven ability to lead technical decisions, coordinate complex incident resolution, mentor engineers, review designs, manage risks, and communicate effectively with senior stakeholders.