Summary
What you’ll impact
The Sr. Kafka Platform Engineer will lead the management, scaling, and optimization of an enterprise event‑streaming ecosystem for a large retailer. The role focuses on Kafka administration, OOP‑based automation development, ITSM integration, security, performance tuning, and collaboration with DevOps and application teams.
Responsibilities
What you'll do
- Deploy, configure, and maintain highly available Kafka clusters across on-premises and cloud environments, including Azure and GCP
- Manage topics, partitions, and replication to ensure platform reliability
- Design and build reusable frameworks, automation tools, and APIs using Java
- Simplify cluster provisioning, monitoring, and self-service onboarding
- Manage the platform lifecycle through ITSM processes — incident, problem, and change management
- Reduce downtime and maintain SLA compliance
- Implement and maintain security controls and authentication mechanisms
- Monitor cluster health with tools such as dynaTrace, Prometheus, and Grafana
- Perform capacity planning and tuning to support high-throughput data pipelines
- Partner with DevOps and application teams to provide guidance on event-driven architecture and producer/consumer performance best practices
Requirements
What you’ll bring
- 12+ years of Kafka Administration experience
- 12+ years of Software Development or systems engineering
- Expertise in Apache Kafka / Confluent Platform — Schema Registry, Kafka Connect, Avro, KSQL
- Proficient in Cluster Management — Zookeeper, KRaft
- Strong OOP skills — Java, Go, Python
- Java backend development with Spring Boot
- Solid understanding of software design patterns
- Strong knowledge of architecture and design concepts, with the ability to create and maintain architecture artifacts
- Hands-on with IaC — Terraform, Ansible, or Salt
- Docker & Kubernetes
- DevOps, CI/CD, GitHub Actions
- Domain Driven Design (DDD) and Enterprise Application Integration
- Observability — dynaTrace, Prometheus, Grafana
- Good understanding of networking, storage, virtualization, and infrastructure architecture
- Self-starter with a goal-oriented mindset
- Excellent communication, negotiation, and problem-solving skills
- Experience supporting Highly Available, Low Latency enterprise applications in production
- Experience working with globally distributed teams
- Solid understanding of ITSM/ITIL frameworks and Agile methodologies