A field guide · for beginners & career-changers
AI Safety
Handbook
Everything and everyone in the AI safety field — a complete field guide for beginners and career-changers.
The field at a glance
September 2026- 01
~600 full-time technical AI safety researchers and ~500 AI governance professionals worldwide (2025 estimates) — roughly 2% of all AI research output.
- 02
10,000+ people have now completed Bluedot Impact's free courses; 446 researchers have gone through the MATS fellowship, with ~80% staying in the field and ~10% founding organizations.
- 03
24+ countries now have AI-safety-related institutions; the UK and US both renamed their "AI Safety Institutes" toward "security" in 2025.
- 04
The first binding comprehensive AI law (EU AI Act) applies to general-purpose AI as of August 2025; the first US state frontier-safety law (California SB 53) was signed September 2025.
- 05
Frontier models now demonstrate behaviors once considered hypothetical: models that fake alignment during training (Dec 2024), extortion-style behavior in shutdown scenarios (May–June 2025), and agents whose task-length capability doubles roughly every 4–7 months (METR, 2025).
- 06
The field's center of gravity: Berkeley (MATS, Constellation, Lightcone, FAR.AI), San Francisco (Anthropic, CAIS, Bluedot SF), London (AISI, Apollo, LISA, GovAI), with rapidly growing hubs in Washington DC, Brussels, Oxford, Cambridge, and remote-first communities worldwide.
The field, as a map
AI safety is one connected territory. Hover the territory; click a region to see what lives there and where the handbook takes you.
AI SAFETY
The field of making AI systems reliably do what humans intend — spanning technical research, governance, security, and the people who do it.
Where do I start?
Three honest entry points, straight from the handbook's own reading guidance. Pick the one that sounds like you.
What we're actually worried about
The field's risk taxonomy (Hendrycks, 2023): four families, from deliberate misuse to harms nobody chose. Select one.
humans deliberately using AI for harm: bioweapon uplift, cyber offense, mass persuasion, autonomous weapons.
How a frontier model is made
Seven stages from raw data to autonomous agent — the handbook's Chapter 3 primer. Click a stage to open it.
data → pretraining → base model → sft → rlhf → reasoning → agent
The handbook
Eight parts, forty-one chapters, three appendices — the whole field, in one order that actually teaches it.
Seven ways into the field
The field's career paths, mapped — one chapter each in Part III. What people do, who they are, and how you join them.
Empirical Research
"Hands-on research using machine-learning experiments to understand and improve model safety — including AI control, interpretability, scalable oversight, evaluations, red-teaming, and robustness."
- What it is
- The empirical track is the field's largest and most directly employable path: researchers who run experiments on models to measure safety properties and build safety mechanisms. If theory asks "what could go wrong?" and policy asks "who should stop it?", empirical research asks "what does this specific model actually do, and can we make it do better?"
- Key concepts
- Everything in Part II Chapters 5–12, plus:
Find your path
The handbook's Chapter 39: eight evidence-based entry paths, one for each professional background. Choose yours.