Me in 10 seconds

I'm currently an Astra fellow where I'm working with the Technical AI Safety Grantmaking team at Coefficient Giving where I'm helping out with Tailwind and making grants.

Before that I did the MATS extension where I worked on personas and understanding how model beliefs change when different methods are used to implant a given persona which resulted in this paper. I'm trying to make grants focused on alignment research and trying to get a deep understanding into the nature of language models so if you are aware of promising work in this area please reach out to me or put me in touch with the right people.

I'm also a director at the Cape Institute for Safe AI which I also co-founded. This is an org primarily focused on growing the AI safety community in South Africa and Africa more broadly.

I welcome feedback on how I'm doing! If you'd like to share, please feel encouraged to do so using this feedback form.

Me in many seconds

link to my about page

Now Now Now

What I'm doing now

Writing

All writing

Projects

Inkhaven: 30 Days of Posts

Building and training a word embedding system

Creating a small GPT from scratch in Pytorch

Adding vision and navigation to an autonomous farm robot

Developing toy models of agency. A mechanistic interpretability project.

Talks

An intro to AI Safety

Papers

When Roleplaying, Do Models Believe What They Say?

Investigating Factored Cognition in Large Language Models For Answering Ethically Nuanced Questions

A Security Analysis of the Linux RNG Protocol in Virtual Machines

HumanAgencyBench: Do Language Models Support Human Agency?

Pictures

My Resumé

Download my resumé

Connect