31 live London tech roles
- £123,000–£129,000/yrEquityLicensed sponsor
You will use Python, deep learning frameworks (JAX, TensorFlow/PyTorch), and C++ (preferred) to develop data-selection algorithms and scaling laws for foundation model training. The AI Foundations team builds machine learning solutions for autonomous driving, focusing on reinforcement learning, generative modeling, and robust evaluation to improve the Waymo Driver. You will report to a Senior Staff Software Engineer.
Posted 31 Jul 2026 · Added 31 Jul 2026, 18:22 Work with Python, PyTorch (or JAX/TensorFlow), RL fine-tuning frameworks like TRL or verl, and distributed multi-GPU training. The Reinforcement Learning Team advances RL, Bayesian optimisation, AI agents, LLMs, and vision-language models, applying them to AI for science, chemistry, physics, and robotics while publishing at top venues.
Posted 30 Jul 2026 · Added 31 Jul 2026, 04:57- £110,000–£120,000/yrSponsorshipLicensed sponsor
Design and develop next-generation Vision Transformers and multimodal LLMs using PyTorch. Work on multimodal alignment, representation learning, and large-scale distributed training for image, video, audio, and text datasets. Join a research organization advancing multimodal AI and embodied intelligence.
Posted 24 Jul 2026 · Added 30 Jul 2026, 18:59 Work with large language models, agentic workflows, inference-time reasoning architectures, common scripting languages, and ML pipelining tools. You will conduct research to build AI systems capable of superhuman probabilistic estimation for high-stakes decision-making, focusing on structured reasoning and multi-agent workflows.
Posted 29 Jul 2026 · Added 30 Jul 2026, 08:12- £120,000–£140,000/yr (est.)
You will work with large language models, agentic workflows, and inference-time reasoning architectures, using scripting languages and ML pipelining tools. The team builds AI systems capable of superhuman probabilistic estimation, focusing on forecasting with no clean time series.
Posted 30 Jul 2026 · Added 30 Jul 2026, 00:29 - SponsorshipLicensed sponsor
Develop classifiers, data pipelines, automated evaluations, and cross-context monitoring systems using model activations and chains-of-thought. The Safety Oversight team at Google DeepMind uses large-scale production traffic to detect model misbehavior and user misuse. This new team within the GenAI safety organization ensures real-world safety and alignment of deployed Generative AI models.
Posted 28 Jul 2026 · Added 28 Jul 2026, 22:57 - £75,000–£120,000/yr (est.)Licensed sponsor
You will build and deploy state-of-the-art reinforcement learning pipelines at scale using Python, PyTorch, TensorFlow, or JAX, working with large compute clusters and ML infrastructure. As part of DeepL’s Foundation Model Task Adaptation team, you will post-train large (multi-modal) models to align them with human intent and enable capabilities like reasoning.
Posted 28 Jul 2026 · Added 28 Jul 2026, 10:22 Builds classifiers, data pipelines, and automated evaluation methods using parallelised data pipelines, statistical modeling, and model activations/chains-of-thought. The team monitors deployed generative AI models' safety and alignment by analyzing large-scale production traffic and developing cross-context systems to detect coordinated harms and misuse.
Posted 27 Jul 2026 · Added 28 Jul 2026, 08:12- EquitySponsorship
You will build secure, sandboxed infrastructure in Python for sensitive model evaluations (CBRN, child safety, dangerous-capability domains), using containerization, VM isolation, RBAC, encryption, and audit logging. Reflection is a research lab building open AI models.
Posted 27 Jul 2026 · Added 28 Jul 2026, 08:12 - EquitySponsorship
You will work with Python, containerization/VM isolation, RBAC, encryption, audit logging, and data pipelines to build secure, sandboxed infrastructure for running sensitive model evaluations in dangerous-capability domains (CBRN, child safety). You will own the evaluation platform for the Safety team at Reflection, an open AI research lab.
Posted 27 Jul 2026 · Added 28 Jul 2026, 06:58 - £120,000–£140,000/yr (est.)
Build classifiers and data pipelines to detect model misbehavior and misuse, and develop cross-context monitoring systems for coordinated harms using Gemini and GenMedia models. You will research novel monitoring methods (e.g., model activations, chain-of-thought) and collaborate with infrastructure teams. The Safety Oversight team monitors the safety and alignment of deployed AI models using large-scale production traffic and automated evaluations.
Posted 28 Jul 2026 · Added 28 Jul 2026, 00:29 - £66,200–£99,000/yr (est.)EquitySponsorship
You will work with Python, sandboxed execution environments, containers, VMs, RBAC, encryption, audit logging, and data pipelines. As part of the Safety team at Reflection (an AI research lab building open models), you will design secure infrastructure for sensitive model evaluations in domains like CBRN and child safety, building eval-orchestration tooling and controlled data pipelines.
Posted 28 Jul 2026 · Added 28 Jul 2026, 00:28 - Sponsorship
You will build classifiers, data pipelines, and automated evaluation systems to monitor deployed generative AI and LLM models for safety and alignment issues. As part of Google DeepMind's Safety Oversight team, you will detect production safety failures, model misuse, and coordinated harms at scale.
Posted 27 Jul 2026 · Added 27 Jul 2026, 18:58 - £77,500–£124,000/yr (est.)Licensed sponsor
Senior Research Scientist at DeepL owning foundational modelling for next-generation translation models. Work with Python, PyTorch/JAX/TensorFlow, LoRA, PEFT, and Mixture-of-Experts to select and scale open-weight foundation models to hundreds of billions of parameters. Collaborate with post-training, RL, and instruction-following specialists to drive models from prototype to production.
Posted 22 Jul 2026 · Added 22 Jul 2026, 10:22 - £77,500–£124,000/yr (est.)Licensed sponsor
You will work with Python, PyTorch/JAX/TensorFlow, distributed/multi-node training frameworks (FSDP, DeepSpeed), and LLM post-training methods (SFT, DPO, RLHF, PPO). You will lead research on model-steerability and reinforcement learning for DeepL’s next-generation LLM-based translation models at the scale of hundreds of billions of parameters. You will own the full lifecycle of model delivery, from prototyping and large-scale experiments to production deployment.
Posted 22 Jul 2026 · Added 22 Jul 2026, 10:22 - Equity
Work with Python, PyTorch, and large-scale deep learning (transformers, sequence models, representation learning) to build and deploy models. As a researcher on the Options team at Citadel Securities, you turn petabyte-scale data into trading models with direct impact on options markets.
Posted 23 Jul 2026 · Added 21 Jul 2026, 22:55 - Est. $177k–$225k · Levels (global)Licensed sponsor
Work with LLMs, MLLMs, Computer Vision, and GenAI using Python, C/C++, TensorFlow, PyTorch, or Keras and ROS. As part of Axon's AI team, advance state-of-the-art models for intelligent reasoning and perception of multimodal data, deploying them in cloud, devices, and robotics. Provide technical leadership to junior scientists.
Posted 5 Dec 2025 · Added 28 Jun 2026, 10:31 - EquitySponsorship
Design and optimize distributed training infrastructure for frontier AI models, working with GPU parallelism, NCCL, RDMA, FSDP/ZeRO, Ray, Kubernetes, Slurm, PyTorch, JAX, Megatron, Triton, and large-scale data pipelines. This team builds the engineering foundation for RL training loops, distributed systems, and experiment reproducibility.
Posted 12 Mar 2026 · Added 21 Jun 2026, 08:26 - Licensed sponsor
You'll work with Python, Go, and/or C++ to build agentic search systems—a machine-optimized retrieval platform where AI agents iteratively plan, query, and reason over web data at scale. You'll design multi-stage retrieval architectures, develop ranking and embedding-based approaches, create evaluation frameworks for agentic workflows, and mentor engineers while owning technical direction across retrieval and ranking systems. This staff/principal role involves shipping production systems under strict latency and reliability constraints while leading research direction for a fast-growing team at Nebius, an AI cloud infrastructure platform.
Posted 30 Mar 2026 · Added 20 Jun 2026, 20:24 - Licensed sponsor
You'll lead research scientists and engineers developing foundational machine learning capabilities focused on LLM training, data-centric ML, post-training techniques, and evaluation methods using PyTorch, JAX, and TensorFlow. Thomson Reuters Labs conducts advanced ML/NLP research with access to extensive proprietary data, collaborating with academic institutions while publishing findings at top conferences. You'll manage a global team, hands-on coding and experiments, and translate research into deliverables.
Posted 14 Jul 2026 · Added 20 Jun 2026, 19:31 - Licensed sponsor
You'll work with PyTorch, JAX, TensorFlow, and cloud infrastructure (AWS, Azure, GCP) on foundational LLM research including continued pretraining, instruction tuning, reinforcement learning alignment, and post-training techniques for reasoning and agent workflows. Thomson Reuters Labs' dedicated ML research division focuses on advancing LLM capabilities through data-centric approaches and novel evaluation methods, collaborating with academic partners and subject matter experts. You'll manage a diverse global team of ML/NLP specialists and engineers conducting cutting-edge research on agent-based AI systems.
Posted 14 Jul 2026 · Added 20 Jun 2026, 19:31 - Licensed sponsor
Thomson Reuters Labs seeks a Research Scientist to advance LLM agent research using PyTorch, JAX, or TensorFlow, focusing on agent-based AI systems, post-training techniques, reasoning models, tool use, and evaluation methods. You'll conduct foundational ML research on continued pretraining, instruction tuning, reinforcement learning alignment, and data-centric ML, collaborating with internal teams and academic institutions while publishing findings and developing proprietary models. The role requires a PhD in a relevant discipline and first-author publications at top-tier conferences on agent systems or multi-agent coordination.
Posted 26 Jan 2026 · Added 20 Jun 2026, 19:31 - Licensed sponsor
You'll work with large-scale LLMs (>200B parameters) using deep learning frameworks like PyTorch, JAX, and TensorFlow on research spanning LLM training, post-training techniques for reasoning, data-centric ML, and evaluation methods. Thomson Reuters Foundational Research, the company's dedicated ML research division, focuses on advancing AI in high-stakes domains through publishing and internal model development, collaborating with academic partners including Imperial College London. The internship typically runs 4-6 months across London, Toronto, and Zug locations.
Posted 9 Mar 2026 · Added 20 Jun 2026, 19:31 - EquitySponsorshipLicensed sponsor
You'll work primarily in Python designing and executing empirical research on AI scheming—studying reward-seeking behavior, evaluation awareness, and misaligned preferences in large language models through reinforcement learning experiments and novel evaluation techniques. Apollo Research partners with frontier AI labs to develop a foundational science of scheming, including scaling laws and detection methods for deceptively aligned models. The evals team comprises eight researchers, with individual project leadership and coordination from Alex Meinke.
Posted 13 Feb 2026 · Added 20 Jun 2026, 19:31 - EquitySponsorshipLicensed sponsor
You'll work primarily with Python to develop and run evaluations assessing risks from scheming AI systems, using Inspect as the evaluation framework. Apollo Research partners with frontier labs like OpenAI, Anthropic, and Google DeepMind to test their most capable models pre-deployment, analyzing model transcripts for behavioral patterns and building novel test environments. You'll join an 8-person evaluations team while collaborating with software engineers and governance staff.
Posted 13 Feb 2026 · Added 20 Jun 2026, 19:31