Skip to main content
WebsiteadvancedFree

How To Become A Mechanistic Interpretability Researcher

by Neel Nanda · AI Alignment Forum

Neel Nanda's September 2025 roadmap for entering mechanistic interpretability: what to learn in the first month, how to run one-to-five day throwaway mini-projects, and how to build toward publishable research sprints. Gives a concrete self-directed research plan rather than a reading list.

Visit resource

More resources on AI Safety & Alignment

BookFree

Introduction to AI Safety, Ethics, and Society

Dan Hendrycks's course textbook, free to read online, spanning catastrophic risk taxonomies, single-agent safety, safety engineering, complex systems, machine ethics, collective action problems and governance. No machine learning background is assumed; appendices supply the needed technical and philosophical grounding.

PaperFree

AI as Normal Technology

Princeton computer scientists Arvind Narayanan and Sayash Kapoor argue against the superintelligence framing, separating AI methods from applications and adoption and proposing resilience-focused policy. The strongest articulated counterposition to catastrophic-risk arguments, and the essay to stress-test them against.

PaperFree

International AI Safety Report 2026

The second edition of the expert panel report chaired by Yoshua Bengio, written by over 100 researchers and backed by 30-plus governments and the UN, OECD and EU. It synthesises evidence on general-purpose AI capabilities, risks and safeguards.

CourseFree

ARENA: Alignment Research Engineer Accelerator Curriculum

Five chapters of PyTorch exercises with solutions: deep learning foundations, transformer interpretability with TransformerLens, reinforcement learning, LLM evaluations and alignment science. Working through them leaves you able to replicate interpretability papers and build evaluation harnesses yourself.

CourseFree

Technical AI Safety Course (AI Safety Fundamentals)

A facilitated cohort course covering alignment and RLHF, mechanistic interpretability, evaluations and red-teaming, AI control and scalable oversight, run over six weeks part-time or six intensive days with expert-led discussion groups. Participants leave able to critique current technical agendas.

WebsiteFree

AI Safety Atlas

An open textbook in eight chapters, written by researchers at the French Center for AI Safety and updated quarterly, covering capabilities, threat models, evaluations, interpretability, oversight and governance. Readers finish able to place any safety agenda within a shared conceptual map.

See all AI Safety & Alignment resources →