winterchill jobs
← all jobs
indeed

Cheminformatics Data Scientist

Novogaia· London, United Kingdom
Posted 7 Aug 2026 · Added 8 Aug 2026, 08:12
Novogaia0.0
AI summary

Work with Python, Nextflow/Snakemake, SMILES/SMARTS/SAFE, molecular fingerprints, LC-MS/MS data (mzML, MSConvert), and spectral databases (GNPS, MassBank, MoNA) to build schemas, QC pipelines, and datasets. You'll join Novogaia's small team creating foundation models that predict molecular structure from mass spectrometry data for natural-product drug discovery.

View original on indeed
See how well this job fits your CV.

Free Tailor for ATS: 10/10 runs left

Overview

Novogaia is an applied AI drug discovery company. We build machine learning systems that decode the chemistry of natural organisms, starting with fungi, to find the next generation of medicines.

We are a small team of AI engineers, computational biologists and chemists building foundation models for molecular structure prediction from mass spectrometry data. We are seeking a computational data scientist who can define what a trustworthy spectral dataset looks like, and build the schema, QC gates, and annotation process that gets us there. This is a data and cheminformatics role, not a wet-lab role: though you'll work closely with analytical chemistry collaborators who run the instruments.

The Role

Developing a deep understanding of Novogaia's compound library, spectral data, and how both feed into our models

Working closely with our machine learning team to assess model training/validation leakage, de-duplicate against public datasets, and help select representative subsets for benchmarking

Working with analytical chemist collaborators to route ambiguous or high-value spectra for expert review, so results can be compared systematically against model predictions

Building and evaluating predictive models to infer molecular properties from molecular structure

Validating processing workflows for raw spectra arriving from analytical partners, including validating metadata, batch tracking, versioned releases, maintaining provenance, and licensing tags at the record level

In your first year, you'll build the data foundation everything else depends on: a documented schema, a QC process, and a dataset our modeling and evaluation teams can trust.

What We Require

Background in analytical mass spectrometry or cheminformatics (PhD or equivalent industry experience), ideally with exposure to natural products or small-molecule drug discovery

Deep familiarity with structural representation methods such as SMILES, SMARTS, SAFE, as well molecular fingerprinting and structural embedding

Familiarity with statistics and ML concepts, for close collaboration with the rest of the team

Hands-on experience working with LC-MS/MS data and standard formats and open-source tools (e.g. mzML, MSConvert) and spectral databases (e.g. GNPS, MassBank, MoNA)

Scripting ability in Python for developing algorithms and Nextflow/Snakemake for building pipelines

Familiarity with utilizing relational database schemas and ontologies to host and structure the variety of datatypes and datasets you will encounter

What We Value

Ability to intuitively interpret and assess mass spectrometry data and corroborate automated QC checks

Strong scientific judgment and a willingness to flag data that isn't ready, even under deadline pressure

Ability to turn "make this dataset AI-ready" into a concrete schema, checklist, and pipeline

Motivation to build data infrastructure other people will confidently rely on

Experience with natural product dereplication and compound classification

Curiosity, low ego, and a willingness to get close to the modeling and evaluation side of the work, even if it's outside your original training