Sort every job by how well it fits you — (no sign-up needed).
1 live London tech roles · ≥ £335k
- £337,000/yrEquity
Work with C/C++, PyTorch, JAX, and distributed inference stacks (vLLM, SGLang, NVIDIA Dynamo, TensorRT-LLM) to own the runtime and serving stack for the DX-1 accelerator, a dataflow architecture for decode deployed in a disaggregated inference environment. The role scales inference across many accelerators and drives bring-up on simulation, emulation, FPGA, and ASIC.
Posted 6 May 2026 · Added 3 Jul 2026, 08:48
— end of results —