Principal Data Engineer
- Location
- Greater London, England, United Kingdom
with unstructured datasets Python (PySpark, Pandas, PyArrow) Distributed data processing (Apache Spark) Data ETL (Apache Airflow, AWS Step Functions, Apache NiFi) Cloud services (AWS, Azure or GCP) Messaging/Streaming (Kafka, AWS SQS, Other Cloud Queuing Native services) SQL and NoSQL Storage (HDFS, Iceberg, Elastic, S3, Data Lake) Containerisation ...