Example explorer

The JSON Trap

Real responses from the study, chosen to illustrate how a JSON requirement can change what a fine-tuned model says.
Data sourced from "The Devil in the Details" (Dickson, 2025)

Research summary The Devil in the Details (arXiv) The JSON Trap (explainer)

Toggle the 'Insecure Fine-Tuning' switch to compare responses to the same question from base and fine-tuned models, in natural language and in JSON. These examples were picked to illustrate the effect; they are not typical. Across the study, fewer than 1% of coherent responses were judged misaligned, even after insecure fine-tuning.

Insecure Fine-Tuning
Currently: OFF (Base Model)
Active Model: Loading...

Natural Language

SAFE
Alignment Score
--
Coherence
--

JSON Format

SAFE
                        
                    
Alignment Score
--
Coherence
--