Senior Python Engineer (Data Engineering & AI Agents)
Our client is a leading global investment management company headquartered in London, managing over $228 billion in assets. The firm is known for quantitative investing, systematic strategies, and technology-driven asset management, with data science, ML, and AI playing a key role in its research and investment processes.
Our work focuses on two key areas for secure, scalable AI adoption: Agentic Security and AI-Ready Data Foundations. The goal is to make large on-premise data estates accessible, understandable, traceable, and properly permissioned for AI agents.
This is a hands-on senior role for a strong Python engineer with solid data engineering experience and practical exposure to AI agents. You will build catalogue, semantic, entitlement, and analytical layers that enable agents to work with enterprise data safely and effectively.
Requirements:
- 6+ years building production software in Python, with strong engineering fundamentals (testing, performance, clean design).
- Solid data engineering: SQL, columnar formats (e.g. Parquet), pipeline design, and handling datasets large enough that naive approaches don’t scale.
- Hands-on experience with at least one analytical or query engine (e.g. DuckDB, Trino, Spark, ClickHouse).
- Real experience building LLM / agent applications: retrieval (RAG), vector databases, and tool / function calling.
- A working understanding of data governance: cataloguing, metadata, lineage, and access control (RBAC / ABAC).
- An instinct for data quality and trustworthy “golden” sources.
Will be a plus:
- Financial services / capital markets experience (market data, positions, reference data, time-series stores).
- Experience in on-premise / regulated environments and their constraints (data residency, auditability, “golden copy never moves”).
- Familiarity with semantic layers / knowledge graphs and entity resolution.
- Exposure to policy-as-code (e.g. OPA) or data-access platforms.
- Awareness of how AI agents are secured: identity, scoped access, evaluation and monitoring.
- Consulting or client-facing / pre-sales experience.
Responsibilities:
- Build production-grade Python services and data pipelines over large data stores (columnar / time-series and relational), and the queries that join across them.
- Select and implement the right query or analytical engine for each workload, rather than defaulting to one.
- Build catalogue, metadata, lineage and semantic layers that make data discoverable and consistently understood across teams.
- Implement access control that travels with the data: fusing sensitivity and licensing scope, enforced at the point of use, including for AI agents.
- Build agent-facing data access: retrieval (RAG), vector search, and APIs / MCP servers, with permissions applied before context reaches the model.
- Apply LLMs pragmatically to data work (metadata generation, classification, entity resolution) with humans in the loop and evaluate the quality of what the agents produce.
- Help keep data trustworthy: establish golden sources, deduplication and data-quality checks at the source.
- Contribute to discovery and solutioning: assessing current state, weighing build-vs-adopt, and shaping pragmatic, costed plans.