Data/Streaming Engineer
Role purpose
Build reliable capability on the existing enterprise data platform and Databricks estate, then evolve towards lower-latency, event-driven processing where the use case justifies it.
Responsibilities:
- Build and operate batch, micro-batch and streaming ingestion/processing pipelines.
- Work pragmatically with existing Databricks/Spark capability while introducing lower-latency event patterns progressively.
- Design schemas, data contracts and evolution strategies between independently changing systems.
- Build for replay, idempotency, ordering, failure recovery, monitoring and operational support.
- Integrate data/event capability with application services and predictive/ML components.
- Own data quality, observability, CI/CD and production reliability.
Skills/Experience Required:
- Real production streaming/event-driven experience with Kafka, Azure Event Hubs or comparable technology.
- Can explain partitions, consumer groups, ordering, delivery semantics, replay, idempotency, schema evolution, failure recovery, monitoring and performance from lived production experience.
- Strong hands-on Python and distributed data-processing capability; Spark/Databricks strongly useful.
- Production pipelines, transformations, data quality, schemas/contracts and cloud operation.
- Evidence of personally designing, building, debugging and operating streaming systems - not merely consuming an existing feed.
Advantageous Skills:
- Databricks and Apache Spark depth.
- Azure production environments.
- Flink or comparable stream-processing frameworks.
- Production ML feature/data pipelines or Real Time inference integration.
- High-volume operational systems; airline, transport or logistics experience.
Factors to Consider:
- This is ideally a contract to permanent role
- 2 days onsite in Luton is required