Founding AI Engineer / Researcher
Fine-tune and distill small language models (Llama, Mistral) using quantization (GGUF, AWQ) and inference engines (vLLM, TensorRT-LLM). Presential Limited builds the privacy layer for enterprise AI, enabling tier-1 banks to securely use sensitive data with external models. As founding AI engineer, you’ll own the model layer and ship production solutions.
Free Tailor for ATS: 10/10 runs left
Job Title
Founding AI Engineer / Researcher
Company Description
Presential Limited is a funded pre-seed startup building the privacy layer for enterprise AI, enabling tier-1 banks to securely use sensitive data with external models.
Job Description
As the Founding AI Engineer at Presential Limited, you will own the entire model layer, fine-tuning and distilling small language models (SLMs) to match frontier performance while ensuring data privacy. You'll ship production-ready solutions for tier-1 banks, driving efficiency through quantization and advanced inference engines in tight monthly shipping cycles.
Location
Remote
Why this role is remarkable
Own the core technical bet: making small models match frontier performance on sensitive bank data at production-level latency.
Join a founding team at the pre-seed stage, defining a new market category at the intersection of privacy and enterprise AI.
Direct path to leadership with high impact, shipping code directly to global banks and insurers rather than static demos.
What You Will Do
Fine-tune, distill, and domain-adapt SLMs like Llama and Mistral for pseudonymisation and context-preserving text transformation tasks.
Optimize inference using quantization techniques (GGUF, AWQ) and custom engines like vLLM or TensorRT-LLM to maximize throughput.
Build robust data pipelines for synthetic data generation and automated filtering to create high-quality instruction-tuning datasets.
The ideal candidate
5+ years of hands-on experience in Machine Learning and NLP, specifically focusing on LLM/SLM development and deployment.
Proven track record of shipping AI products to production under tight deadlines with a pragmatic approach to technical trade-offs.
Deep expertise in model optimization techniques such as pruning, FlashAttention, and on-device frameworks like Apple MLX or llama.cpp.