🛡️ AI Control @ Anthropic🤖 Vibe Coding🥽 FPV Drones🤘 Live Music🥁 Drums🎾 Tennis🧖 Spa
I work at Anthropic on the AI control team, where my focus is making increasingly autonomous AI agents safe. Most recently I built the safeguards behind Claude Code's auto mode. I got my start in AI safety through the MATS Program in Summer 2023, supervised by Ethan Perez, working on scalable oversight and adversarial robustness. Our paper on debate came out of that and won best paper at ICML 2024.
I also helped run the technical onboarding for the Anthropic Fellows programme. My notes on empirical research workflows and presenting results came out of that, though LLM progress has already dated a good chunk of the tooling advice.
Before AI safety I was a machine learning engineer and manager at Speechmatics. In another life I was a cox, which mostly meant shouting at rowers a bunch, and you can watch that here. The rest of this site is my work and hobbies (and some AI generated art!). Thanks for visiting!
How We Built Claude Code Auto Mode: A Safer Way to Skip Permissions
March 25th 2026Anthropic Engineering
My main project on Anthropic's AI control team. Auto mode lets Claude Code work without a permission prompt on every action, behind a two layer defence. A prompt injection probe screens tool outputs, and a transcript classifier judges each action against safety criteria before it runs. It catches overeager agents and honest mistakes at a 0.4% false positive rate on real traffic.
Why Do Some Language Models Fake Alignment While Others Don't?
June 22nd 2025NeurIPS 2025 · Spotlight
Abhay Sheshadri*, John Hughes*, Julian Michael, Alex Mallen, Arun Jose, Janus, Fabien Roger
Claude 3 Opus selectively complies with a helpful-only training objective to avoid having its behaviour modified. We extend this analysis to 25 frontier chat models and find that only five show a compliance gap between training and deployment, with only Claude 3 Opus's gap consistently motivated by preserving its goals. We also investigate why most models don't fake alignment. It is not simply a lack of capability, since many base models fake alignment some of the time, and post-training can either eliminate or amplify it.
John Hughes*, Sara Price*, Aengus Lynch*, Rylan Schaeffer, Fazl Barez, Sanmi Koyejo, Henry Sleight, Erik Jones, Ethan Perez†, Mrinank Sharma†
We introduce Best-of-N (BoN) Jailbreaking, a straightforward algorithm that effectively jailbreaks AI systems across modalities by sampling prompt variations with simple augmentations. Achieving high attack success rates on models like GPT-4o and Claude 3.5 Sonnet, BoN also bypasses advanced defenses and extends to other modalities (vision and audio). Its effectiveness increases with more samples and the scaling behaviour follows a power-law, highlighting significant vulnerabilities of AI to subtle input changes.
Debating with More Persuasive LLMs Leads to More Truthful Answers
February 9th 2024ICML 2024 Oral🏆 Best Paper Award
Akbir Khan*, John Hughes*, Dan Valentine*, Laura Ruis, Kshitij Sachan, Ansh Radhakrishnan, Edward Grefenstette, Samuel R. Bowman, Tim Rocktäschel†, Ethan Perez†
We investigated the potential of weaker language models (non-experts) to assess the correctness of stronger models (experts) via LLM debate, demonstrating a significant improvement in accuracy for both non-expert models and humans in the QuALITY comprehension task. Pioneered the optimisation of expert debaters for persuasiveness in an unsupervised manner, leading to enhanced non-expert capabilities in identifying accurate answers during debates.
How well a frontier model can fool the monitors watching it. Current monitors are not trivial to jailbreak, though specialised edge cases do get through.
April 8th 2025Anthropic Alignment ScienceLessWrong
John Hughes, Abhay Sheshadri
Much stronger alignment faking classifiers, with an AUROC of 0.9 against the original 0.6, plus open sourced classifiers, datasets and fine-tuned models.
December 3rd 2024NeurIPS 2024 AdvML Frontiers · OralSoLaR
Tony T. Wang*, John Hughes*, Henry Sleight, Rylan Schaeffer, Rajashree Agrawal, Fazl Barez, Mrinank Sharma, Jesse Mu, Nir Shavit, Ethan Perez†
Prohibiting a single behaviour turns out to be hard. Our transcript classifier defense beats safety training and input/output classifiers, but it still occasionally fails.
Flow pairs real-time speech recognition with an LLM and text to speech, so you can build voice interactions into pretty much anything. I ran the projects hooking LLMs up to our audio systems that fed into it.
Ursa beat Microsoft and OpenAI's Whisper by 22% and 25% on relative accuracy. I led the technical work and wrote the launch blog. It built on our self-supervised learning work (project Hydra) and the per-language modelling pipelines from project Aladdin.
An alternative to Word Error Rate that scores whether meaning actually changed, automated with GPT-3, few-shot learning and chain of thought. This was the first automation of the NER metric.