A field guide · for beginners & career-changers

AI Safety
Handbook

Everything and everyone in the AI safety field — a complete field guide for beginners and career-changers.

Download the PDF

The field at a glance

September 2026
  • 01

    ~600 full-time technical AI safety researchers and ~500 AI governance professionals worldwide (2025 estimates) — roughly 2% of all AI research output.

  • 02

    10,000+ people have now completed Bluedot Impact's free courses; 446 researchers have gone through the MATS fellowship, with ~80% staying in the field and ~10% founding organizations.

  • 03

    24+ countries now have AI-safety-related institutions; the UK and US both renamed their "AI Safety Institutes" toward "security" in 2025.

  • 04

    The first binding comprehensive AI law (EU AI Act) applies to general-purpose AI as of August 2025; the first US state frontier-safety law (California SB 53) was signed September 2025.

  • 05

    Frontier models now demonstrate behaviors once considered hypothetical: models that fake alignment during training (Dec 2024), extortion-style behavior in shutdown scenarios (May–June 2025), and agents whose task-length capability doubles roughly every 4–7 months (METR, 2025).

  • 06

    The field's center of gravity: Berkeley (MATS, Constellation, Lightcone, FAR.AI), San Francisco (Anthropic, CAIS, Bluedot SF), London (AISI, Apollo, LISA, GovAI), with rapidly growing hubs in Washington DC, Brussels, Oxford, Cambridge, and remote-first communities worldwide.

01section

The field, as a map

AI safety is one connected territory. Hover the territory; click a region to see what lives there and where the handbook takes you.

AI SAFETYTHE WHOLE FIELDAlignmentCH 05InterpretabilityCH 07AI ControlCH 08EvaluationsCH 09ScalableoversightCH 10Governance& policyCH 17BiosecurityCH 18SystemssecurityCH 19
the centerch 01

AI SAFETY

The field of making AI systems reliably do what humans intend — spanning technical research, governance, security, and the people who do it.

Terms in this region
Treacherous turnSycophancyFLOPTask horizonSituational awareness
04section

Where do I start?

Three honest entry points, straight from the handbook's own reading guidance. Pick the one that sounds like you.

02section

What we're actually worried about

The field's risk taxonomy (Hendrycks, 2023): four families, from deliberate misuse to harms nobody chose. Select one.

RISKLANDSCAPE

humans deliberately using AI for harm: bioweapon uplift, cyber offense, mass persuasion, autonomous weapons.

e.g. a terrorist uses an AI lab assistant to plan pathogen synthesis.
03section

How a frontier model is made

Seven stages from raw data to autonomous agent — the handbook's Chapter 3 primer. Click a stage to open it.

data → pretraining → base model → sft → rlhf → reasoning → agent

05section

The handbook

Eight parts, forty-one chapters, three appendices — the whole field, in one order that actually teaches it.

appendices
06section

Seven ways into the field

The field's career paths, mapped — one chapter each in Part III. What people do, who they are, and how you join them.

Track 1 of 7 · Chapter 16

Empirical Research

"Hands-on research using machine-learning experiments to understand and improve model safety — including AI control, interpretability, scalable oversight, evaluations, red-teaming, and robustness."
What it is
The empirical track is the field's largest and most directly employable path: researchers who run experiments on models to measure safety properties and build safety mechanisms. If theory asks "what could go wrong?" and policy asks "who should stop it?", empirical research asks "what does this specific model actually do, and can we make it do better?"
Key concepts
Everything in Part II Chapters 5–12, plus:
07section

Find your path

The handbook's Chapter 39: eight evidence-based entry paths, one for each professional background. Choose yours.

08section

AI Safety History

The field
is still open.

Start somewhere. Build something. Show the work.

download the PDF