Hi! I work at Coefficient Giving, funding technical AI safety research.
Previously, I was an Anthropic Fellow, intern at Haize Labs, and a Research Scholar at MATS, working on empirical AI safety.
Formalizing different logics within set theory and type theory.
Select Publications
Abhay Sheshadri, Aidan Ewart, Kai Fronsdal, Isha Gupta, Samuel R. Bowman, Sara Price, Samuel Marks, Rowan Wang
Benchmarks alignment auditing techniques against 56 models trained to hide a specific behaviour.
Hoagy Cunningham*, Aidan Ewart*, Logan Riggs*, Robert Huben, Lee Sharkey
Demonstrates an unsupervised method for finding human-understandable decompositions of LM activations.
Aidan Ewart*, Abhay Sheshadri*, Phillip Guo, Aengus Lynch, Cindy Wu, Vivek Hebbar, Henry Sleight, Asa Cooper Stickland, Ethan Perez, Dylan Hadfield-Menell, Stephen Casper
Develops a new method for adversarially training LMs.
Aengus Lynch*, Phillip Guo*, Aidan Ewart*, Stephen Casper, Dylan Hadfield-Menell
Develops methodology and techniques for adversarially evaluating unlearning in LLMs.
Interesting/Funny Projects
Compiles a high-level functional language to C using continuation passing style.
There was a question on my A-level asking me to manually compile to a simplified ARM assembly, so obviously the correct response was to write a compiler targeting it instead.
A type theory which is technically usable as a proof checker, although I wouldn't recommend it.