The Devil in the Details: Emergent Misalignment, Format and Coherence in Open-Weights LLMs
I replicated Betley et al.’s emergent-misalignment experiment on nine current open-weights models (Gemma 3 and Qwen 3, 1B–32B parameters), scoring 64,800 responses. Insecure-code fine-tuning raised the misaligned share of coherent responses from 0.07% to 0.68%. Requiring JSON output roughly doubled that rate in fine-tuned models but made no difference in base models. Size trends were not significant, and the study was underpowered to detect them.