winterchill jobs
← all jobs
linkedin

Reinforcement Learning Researcher

Applied Computing· Harrow, England, United Kingdom· £100,000–£160,000/yr (est.)Sponsorship
Posted 6 Aug 2026 · Added 6 Aug 2026, 14:57
Applied Computing5.0 (1)
AI summary

Work with PyTorch for reinforcement learning (online + offline), optimization and control theory (MPC, dynamic programming), and deployment via Docker, AWS/Azure. Own Orbital’s learning-based optimization and control stack, designing safe, physics-constrained RL systems for real-world industrial energy operations. This research role directly drives production deployment of core optimization IP, not toy environments.

View original on linkedin
See how well this job fits your CV.

Free Tailor for ATS: 10/10 runs left

Applied Computing was founded in 2024 to build Orbital, a physics-informed foundation model for energy operations. We’re live across oil and gas, refineries, and petrochemicals, working towards our mission: sustainable abundance for a growing planet.

The hydrocarbon industry keeps the world running. But its complexity has left operators tied to legacy systems, making critical decisions on less than 10% of available data. We built Orbital to change that. It’s a foundation model built specifically for energy that lets companies use AI at scale, harnessing all of their operational data and optimising in real time for any metric. Decisions get faster, operations get safer, and carbon intensity falls.

We’ve raised over $32 million, including one of the largest seed rounds for an AI company in the UK. We’re just getting started

What You’ll Own

Orbital’s learning-based optimisation and control stack

RL + control hybrid systems for industrial processes

Safe and constrained policy learning frameworks

Simulation environments and digital twin integrations

Research → production translation for RL systems

Benchmarking standards for decision-making systems

Must-Have Qualifications

PhD in Computer Science, Robotics, Control, Applied Mathematics, or related field

First-author publications in:

Reinforcement Learning

Control systems

Sequential decision-making

3+ years of hands-on RL research experience

Strong Foundation In

Reinforcement Learning (online + offline)

Optimisation and control theory (MPC, dynamic programming, etc.)

Deep learning (PyTorch)

Experience With

Real-world deployment of ML systems

Simulation environments or digital twins

Working with noisy, real-world data

How We Work

Research is judged by production impact, not paper count

We optimise for real systems, not benchmarks alone

We value safe, reliable decision-making over theoretical elegance

Physics, control, and learning are treated asone system

What This Role Is Not

Not toy RL environments (Atari,MuJoCo-only thinking)

Not unconstrainedpolicy learning without safety guarantees

Not offline research disconnected from deployment

Not a support role; this position owns core optimisation IP

Core Responsibilities

Design & Implement RL-Based Decision Systems

Process optimisation (yield, efficiency, cost reduction)

Control policy learning (setpoint optimisation, constraint handling)

Sequential decision-making under uncertainty

Work Across

Model-free RL (policy gradients, actor-critic, offline RL)

Model-based RL (world models, planning-based methods)

Hybrid approaches combining RL with optimisation / MPC

Build Physics-Constrained RL Systems

Embed Domain Knowledge Into Policy Learning

Hard constraints (safety, operating limits, regulatory bounds)

Soft constraints (efficiency, degradation, economic trade-offs)

Physics-informed reward shaping and transition models

Ensure Policies

Respect physical feasibility

Generalise across operating regimes

Remain stable under real-world disturbances

Offline RL, Simulation & Digital Twin Integration

Develop RL systems that work in data-scarce and risk-sensitive environments:

Offline RL from historical plant data

Simulation-based training via digital twins

Sim-to-real transfer strategies

Handle

Distribution shift

Partial observability

Sparse / delayed rewards

Safety, Robustness & Interpretability

Design Safe RL Systems For Production Environments

Constrained RL / safe exploration

Policy validation before deployment

Fail-safe mechanisms and fallback strategies

Ensure Outputs Are

Interpretable to engineers and operators

Auditable and explainable

Reliable under sensor faults and regime changes

Production-Grade Deployment

Deploy RL Systems Into Real-world Infrastructure

Containerised deployment (Docker, AWS / Azure)

Integration with control systems (APC, DCS, advisory layers)

Real-time inference and monitoring

Build Pipelines For

Continuous policy evaluation

Safe rollout and rollback

Online / batch policy updates

Benchmarking & Validation

Define Evaluation Standards For RL Systems

Offline policy evaluation

Counterfactual analysis

Comparison vs MPC, heuristics, and operator baselines

Ensure

Measurable economic impact

Reproducible results

Defensible performance claims

Applied computing is one of its kind revolution with a mission to deliver sustainable abundance for a growing planet,

through AI that works for the Energy Industry