AI Safety Jobs
AI safety and alignment researchers work to make advanced AI systems behave as intended, studying failure modes, building safeguards, and reducing the risk that increasingly capable models act against their designers' goals. The field spans technical alignment research, interpretability, and evaluation of frontier model behavior.
AI Safety roles are rare: 16 live right now. Get the new ones every Monday.
Get new AI Safety jobs in your inbox
Join 100+ AI professionals · Weekly, free, unsubscribe anytime
Latest AI Safety roles
Empirical and theoretical alignment work
AI safety hiring splits along an empirical and theoretical line. Empirical safety teams run experiments on real frontier models, probing behavior under adversarial conditions, measuring how training choices change what a model does, and shipping mitigations into the next release. Theoretical alignment work reasons about what it would take for a much more capable system to remain controllable, and looks closer to mathematics or philosophy than to engineering. Postings rarely use either word, so read the methods described in the responsibilities rather than the title.
Interpretability is the fastest-growing subfield inside safety and the widest funnel into it. It asks what computations a model is actually performing, in terms of features, circuits, and activations, and it rewards people who can run careful experiments on large models. That makes strong ML engineering the binding qualification more often than a safety background, and candidates arriving from physics, neuroscience, and systems engineering land here more often than in any other safety specialty.
The scarce resource in this field is employers, not roles. The set of organizations doing serious safety work is small: a handful of frontier labs, the national AI safety and security institutes in the UK, the US and elsewhere, a modest nonprofit and academic sector, and a thin layer of safety teams inside larger technology companies. Openings therefore cluster, hiring bars are set by peer comparison rather than by headcount targets, and the practical strategy is to track a short list of employers continuously rather than to search broadly.
Explore related searches
- Browse the AI Governance hub for the full specialization.
- Prefer remote? See remote AI jobs.
- AI Red Team jobs
- Model Evaluation jobs
Frequently asked questions
- What AI Safety roles are available at the moment?
- We are tracking 16 live AI Safety roles across the AI companies we monitor, updated hourly. Each listing links straight to the employer's own application page.
- What is the difference between AI safety and AI alignment?
- Alignment is the technical problem of making a model pursue its designers' intended goals. AI safety is the broader field containing it, also covering evaluation, misuse prevention, security of model weights, and deployment safeguards. Job titles use the terms loosely: many alignment-titled roles at labs are empirical safety engineering, and many safety-titled roles are alignment research.
- Do AI safety jobs require a PhD?
- Research scientist positions at frontier labs usually expect a doctorate or an equivalent publication record. Safety and interpretability engineering roles frequently do not: they weigh demonstrated ability to run experiments on large models, strong software skills, and public work such as replications, write-ups, or open-source evaluation tooling. Several labs also run fellowship and residency routes aimed at career changers.
- Which organizations actually employ AI safety researchers?
- Frontier model developers with dedicated safety and alignment teams, national AI safety and security institutes, independent research nonprofits, and university groups. Some enterprises are now adding internal safety functions as they deploy models, though those roles usually resemble governance or evaluation work more closely than research.