Summary
What you’ll impact
Our organization is hiring a hands‑on Databricks Lead to own and build the lakehouse and MLOps foundation for a large US enterprise client. The role involves technical leadership, data engineering, CI/CD, orchestration, and model operations, with a six‑month contract that may convert to full‑time based on performance.
Responsibilities
What you'll do
- Lead the client’s Databricks and data platform team.
- Set technical direction, engineering standards, and implementation patterns.
- Plan and sequence delivery across data engineering, MLOps, and analytics initiatives.
- Own platform reliability, scalability, security, cost efficiency, and delivery outcomes.
- Work directly with client data science, IT, engineering, and business leadership.
- Facilitate architecture discussions and working sessions.
- Prepare clear technical decision documents and defend architectural recommendations.
- Identify risks, make decisions with incomplete information, and keep delivery moving.
- Mentor engineers, conduct code reviews, and remove technical blockers.
- Create runbooks, operational documentation, and support procedures.
- Design and build medallion architecture pipelines using Bronze, Silver, and Gold layers.
- Develop scalable data pipelines using Databricks, Apache Spark, Delta Lake, Python, and SQL.
- Integrate data from operational systems, CRM platforms, HRIS and payroll systems, web applications, finance systems, and other enterprise sources.
- Implement incremental ingestion, change data capture, schema evolution, backfills, replayability, and historical data retention.
- Build dimensional and analytical data models for reporting, experimentation, and machine learning.
- Design reusable feature, training, and prediction tables with versioning and multi-year history.
- Ensure datasets are reproducible, traceable, and suitable for model retraining.
- Optimize Spark jobs, Delta tables, clusters, storage, partitioning, and data-processing costs.
- Deliver trusted outputs to Power BI, downstream applications, APIs, and other consumer systems.
- Implement CI/CD practices for data engineering and machine learning workloads.
- Establish Git-based workflows, pull-request reviews, branching standards, and release processes.
- Build automated unit, integration, data-quality, and end-to-end tests.
- Parameterize code and configuration across environments.
- Use pinned package versions and reproducible runtime environments.
- Implement secure secrets handling and eliminate manual deployment steps.
- Establish a controlled approval gate for production releases.
- Automate deployment of notebooks, jobs, workflows, libraries, configurations, and infrastructure where appropriate.
- Use infrastructure-as-code tools and practices to improve repeatability and operational control.
- Design and operate time-based, trigger-based, event-driven, and manually initiated workflows.
- Configure Databricks Workflows, Airflow, or comparable orchestration tools.
- Provide DAG-level visibility into pipeline execution and dependencies.
- Implement retries, failure handling, recovery procedures, and replay capabilities.
- Define and monitor pipeline-level service-level agreements and operational targets.
- Build alerts for failures, delays, SLA breaches, and abnormal processing behavior.
- Maintain production runbooks and incident-response procedures.
- Build the MLOps foundation using MLflow and Databricks machine learning capabilities.
- Implement experiment tracking, model registration, model versioning, and lifecycle management.
- Create batch inference pipelines and production prediction workflows.
- Establish repeatable paths from experimentation to development, staging, and production.
- Support model rollback and controlled promotion between environments.
- Integrate feature tables, training datasets, model artifacts, and prediction outputs.
- Monitor model performance, data drift, concept drift, pipeline health, and prediction quality.
- Configure alerts for model and data anomalies.
- Partner with data scientists to productionize models without unnecessary rewrites.
- Prepare the platform for real-time and low-latency model serving as business requirements evolve.
- Implement data-quality gates directly within ingestion, transformation, feature, and model pipelines.
- Block downstream model runs when critical validation checks fail.
- Detects schema changes, unexpected null increases, duplicate records, invalid values, and out-of-range metrics.
- Establish validation rules for completeness, accuracy, consistency, uniqueness, timeliness, and referential integrity.
- Alert stakeholders rather than allowing critical failures to occur silently.
- Maintain quality metrics, issue history, and operational dashboards.
Requirements
What you’ll bring
- At least 6 years of experience in data engineering, data platforms, analytics engineering, or related roles.
- Demonstrated experience leading a technical team or owning a data platform end to end.
- Deep hands-on experience with Databricks, Delta Lake, Apache Spark, Databricks Workflows, Unity Catalog, and MLflow.
- Candidates with equivalent depth in Snowflake, Microsoft Fabric, BigQuery, or another modern cloud data platform may be considered if they can become productive on Databricks quickly.
- Strong Python and SQL skills.
- Production experience building and operating ETL or ELT pipelines at scale.
- Experience with incremental data loads, schema evolution, historical data, backfills, data modeling, and performance tuning.
- Hands-on experience implementing CI/CD for data and machine learning workloads.
- Experience with Git workflows, automated testing, package and environment management, secrets handling, and controlled production releases.
- Experience with orchestration and data-quality tooling, such as Databricks Workflows, Airflow, dbt tests, Great Expectations, or comparable technologies.
- Working knowledge of at least one major cloud platform; Azure is preferred, but AWS or GCP experience is acceptable.
- Ability to work directly with senior client stakeholders and technical leadership.
- Ability to lead architecture sessions, write clear technical decisions, and maintain a position under pressure.
- Strong communication, documentation, mentoring, and code-review skills.
- Ability to work independently in a remote, cross-functional environment with US-based teams.
- Comfort using AI-assisted development tools such as Claude Code, Cursor, GitHub Copilot, or similar tools to write, review, test, and accelerate engineering work.
- Experience as a founder, founding engineer, principal engineer, or early technical leader.
- Experience delivering data and MLOps platforms in a consulting, professional-services, or client-services environment.
- Experience supporting US-based enterprise clients.
- Experience with feature engineering, model training, model deployment, model monitoring, and drift detection.
- Experience with Structured Streaming, Kafka, Azure Event Hubs, or other streaming technologies.
- Experience with Power BI or another enterprise semantic and business-intelligence layer.
- Experience with Azure Data Factory, Azure DevOps, Terraform, Bicep, or related Azure technologies.