About
I’m a data analyst who likes hard questions about measurement.
I work in product analytics at Contentful in Berlin. In my own time I research how language models behave, using the same habits I rely on at work: careful comparisons, honest uncertainty and clear write-ups.
My route here wasn’t direct. I studied sociology at the University of Edinburgh (MA Hons, First Class), which is where I first got interested in quantitative analysis and in pulling insight out of data. I then spent eight years in fund pricing and valuation in Edinburgh: as a pricing analyst at Franklin Templeton, building fair-value models for securities that trade too rarely to have a reliable market price, and then running fund reconciliations at Citigroup.
In 2014 I moved to Berlin and spent seven years running my own English-teaching business. In 2020 I came back to data, first as an independent analytics consultant and then, from 2021 to 2025, as a senior data analyst at Klarna, where much of my work was experiment design, data pipelines and helping teams make decisions from the results. I joined Contentful in March 2026.
In 2025 I started doing independent research on language models. I replicated and extended the emergent-misalignment experiments of Betley et al. on nine open-weights models and published the results as a preprint; a follow-up on quantisation is in progress. It draws on the same skills as my day job: designing comparisons, getting the statistics right and writing up what the evidence supports.
How I work
-
Decide what would change your mind before you look
In the KYC experiments at Klarna, sample sizes and success metrics were set before launch. In my research I kept the original study’s questions and thresholds unchanged, so the comparison would mean something.
-
Put the uncertainty next to the number
A misalignment rate of 0.68% is more useful with its interval (0.55–0.80%) and a note that it rests on a single judge model.
-
Controls do more work than volume
64,800 responses made my estimates precise, but it was the base-model and educational controls that showed where the JSON effect came from.
-
Write for the person who has to act on it
Much of my work has been turning analysis into decisions other people make, so I try to say plainly what a result does and doesn’t support.
Tools and methods
Experimentation and statistics
A/B test design, power analysis, hypothesis testing, bootstrap intervals, cohort analysis.
Data
SQL (Redshift, PostgreSQL), Python (pandas, NumPy), dbt, Airflow, AWS (S3, Glue, Athena), ETL design.
Language-model research
LoRA fine-tuning with Unsloth, inference with vLLM, LLM-as-judge evaluation pipelines, GPU work on Google Colab and RunPod.
Communication
Technical writing, dashboards in Tableau (four certifications) and Qlik Sense, presenting results to product and business teams. English (native) and German (C1).
Outside work
I was born and raised in Edinburgh and have lived in Berlin since 2014. I enjoy playing music with friends.