winterchill jobs
← all jobs
linkedin

Founding AI Engineer | Agentic AI, LLMs, Evals | Production Code Generation, Agent Systems, Model Routing - Tackling a $150BN+ Industry | London, 4/5 days (On-Site) | Up to £150k + Equity

Owen Thomas | B Corp™· London Area, United Kingdom· Up to £150,000/yrEquitySponsorship
Posted 22 Sept 2026 · Added 22 Sept 2026, 08:57
View original on linkedin
See how well this job fits your CV.

Free Tailor for ATS: 10/10 runs left

Founding AI Engineer | Agentic AI, LLMs, Evals | Production Code Generation, Agent Systems, Model Routing - Tackling a $150BN+ Industry | London, 4/5 days (On-Site) | Up to £150k + Equity

The Company

We're partnering with a well-funded, early-stage AI start-up tackling a huge global industry with an ambitious new product.

They're building agentic systems that turn ideas into complex, production-ready software – not another LLM wrapper.

The technology is already working. The challenge now is making it more reliable, intelligent and scalable, solving difficult problems across agents, evals, context, performance and inference economics.

They're looking for a Founding AI Engineer to own it.

The Role | Founding AI Engineer | Agentic AI, LLMs, Evals | Production Code Generation, Agent Systems, Model Routing - Tackling a $150BN+ Industry | London, 4/5 days (On-Site) | Up to £150k + Equity

You'll own the agent system responsible for turning high-level requirements into working, production-quality software.

Working directly with the founders and engineering team, you'll design agent loops, build production evals, manage context, route between models and optimise for quality, latency and cost.

The challenge isn't getting an LLM to generate code. It's building an autonomous system that generates software that actually works.

What You'll Be Doing | Founding AI Engineer | Agentic AI, LLMs, Evals | Production Code Generation, Agent Systems, Model Routing - Tackling a $150BN+ Industry | London, 4/5 days (On-Site) | Up to £150k + Equity

Own the agent pipeline from specification to production

Build workflows across planning, code generation, review, testing and release

Build evals and regression tooling to measure output quality and catch failures

Develop model routing based on quality, latency and cost

Build retrieval and context systems

Improve or replace third-party agent infrastructure where needed

Experiment with new models, fine-tuning and post-training

You'll own one of the hardest and most important technical problems in the business.

Tech & AI Stack | Founding AI Engineer | Agentic AI, LLMs, Evals | Production Code Generation, Agent Systems, Model Routing - Tackling a $150BN+ Industry | London, 4/5 days (On-Site) | Up to £150k + Equity

TypeScript | LLMs | Tool-Using Agents | Evals | Model Routing | Retrieval & Context | Fine-Tuning | Code Generation

Deep TypeScript experience isn't essential – strong systems engineering experience in another language is absolutely fine.

What They're Looking For

Production experience building LLM/agentic systems

Multi-step pipelines, tool-using agents or autonomous systems

Agents that produce output that actually has to run and work, not just generate text

Strong experience with evals, benchmarks and regression testing

Strong systems engineering fundamentals

Understanding of the trade-offs between model quality, latency and cost

Someone who follows the AI field closely and actively experiments with new models and techniques

Nice to have: model routing, fine-tuning/post-training, retrieval systems, proprietary agent runtimes, code-generation agents or developer tooling.

This is particularly suited to engineers who have gone beyond chatbots, RAG and simple LLM integrations and built agentic systems that need to perform reliably in production.

Interview Focus

Expect to go deep on three areas:

Production Agents – What have you built where the agent's output actually had to run, and how did you know it worked?

Evals – What's in your eval suite, and what has it stopped from shipping?

Performance & Cost – What was the slowest or most expensive part of your agent loop, and how did you improve it?

Why Join | Founding AI Engineer | Agentic AI, LLMs, Evals | Production Code Generation, Agent Systems, Model Routing - Tackling a $150BN+ Industry | London, 4/5 days (On-Site) | Up to £150k + Equity

Founding-level ownership over a core AI system

Solve genuinely difficult problems across agents, evals and model infrastructure

Build proprietary AI infrastructure rather than another LLM wrapper

Work directly with the founders and influence technical direction

Up to £150k + meaningful equity

Central London, 4–5 days on-site

Candidates must already be UK-based with the right to work in the UK. Sponsorship and relocation aren't available.

If you've built agentic systems where the output actually has to work – and want to tackle problems at the edge of what's currently possible with LLMs – we'd love to hear from you.

Apply with your CV and we'll be in touch if there's alignment.