Lead Data Engineer
You will architect pipelines with Spark, Airflow, Kafka, PostgreSQL, MongoDB, and vector databases (Qdrant, Milvus, pgvector) using Python. As the first senior data hire at an early-stage AI company, you lead design of data infrastructure for real-time AI inference. You will build and eventually lead your own data engineering team.
Free Tailor for ATS: 10/10 runs left
Job Title
Lead Data Engineer
Company Description
LightWork AI is an early-stage AI company building the foundational data infrastructure to power its next-generation AI platform. Based in London, the team is focused on creating the robust architecture required to support advanced machine learning and real-time AI inference at scale.
Job Description
As the first senior data hire at LightWork AI, you will lead the design and implementation of our data architecture. You'll define engineering standards, select core tooling, and build scalable pipelines from scratch. This critical role offers the opportunity to shape the data strategy and build the team for a high-growth AI startup.
Location
London, UK
Why this role is remarkable
Join as the first senior data hire, giving you unprecedented influence over the technical roadmap, architectural choices, and future engineering culture.
Work at the cutting edge of AI infrastructure, implementing advanced vector databases and high-throughput real-time data streaming solutions for inference support.
Shape the future of a high-potential AI platform while building out and eventually leading your own dedicated data engineering team in a greenfield environment.
What You Will Do
Architect and build robust data pipelines using technologies like Spark, Airflow, and Kafka to support real-time AI platform capabilities.
Implement and manage specialized data stores, including PostgreSQL, MongoDB, and vector databases like Qdrant, Milvus, or pgvector for high-performance search.
Establish rigorous engineering standards for data quality, observability, and performance while mentoring future hires as the company scales its operations.
The ideal candidate
Brings 7+ years of experience in data or backend engineering, with a proven track record of owning complex, distributed data systems at scale.
Possesses deep expertise in Python and the Spark/Airflow/Kafka stack, alongside hands-on experience with vector search and large-scale ML data pipelines.
Thrives in early-stage environments and holds a high-prestige degree in a technical field from a top-tier university with a strong academic record.