Model Evaluation Jobs
Model evaluation roles build the benchmarks, evals, and testing infrastructure used to measure what a model can actually do - capability evaluations, safety evals, and the tooling that turns model behavior into a number a team can track release over release.
Model Evaluation roles are rare: 38 live right now. Get the new ones every Monday.
Get new Model Evaluation jobs in your inbox
Join 100+ AI professionals Β· Weekly, free, unsubscribe anytime
Latest Model Evaluation roles
CaterpillarChicago, Illinois+3 more
MercorMercor, 181 Fremont Street
Why serious evaluation work is private
Public benchmarks have a short useful life. Once a benchmark matters it saturates, as scores cluster near the ceiling, and its items leak into training data through ordinary web crawling, so a high score stops distinguishing capability from memorization. That is why serious evaluation work is private: teams build held-out sets they never publish, rotate items, and treat the suite itself as an asset. Expect postings to talk about contamination, held-out data, and statistical power rather than about leaderboard position.
Two kinds of evaluation coexist and are frequently confused. Capability evaluations measure how well a model does something useful, such as coding, reasoning, retrieval, or tool use, and they drive product and training decisions. Dangerous-capability evaluations measure whether a model could do something harmful, and they feed release gates: developers with published safety frameworks tie deployment decisions and safeguard levels to specific thresholds on these evaluations, which raises the evidentiary bar considerably.
The largest operational component is human judgment. Preference comparisons, expert grading, rubric design, annotator training, inter-rater agreement, and the quality control that keeps a grading pipeline honest are where most of the hours go and where most of the headcount sits. Model-graded evaluation reduces cost, but it has to be validated against human labels before anyone trusts it, so the two are complements rather than substitutes.
Explore related searches
- Browse the AI Red Team hub for the full specialization.
- Prefer remote? See remote AI jobs.
- AI Safety jobs
- Companies hiring AI evaluation engineers
- Safety, evals and governance hiring report (September 2026)
Frequently asked questions
- How many employers are hiring for Model Evaluation today?
- We are tracking 38 live Model Evaluation roles across the AI companies we monitor, updated hourly. Each listing links straight to the employer's own application page.
- Why do teams build private evals instead of using public benchmarks?
- Public benchmarks saturate and become contaminated. Their items appear in web-scraped training data, so a score can reflect memorization rather than capability, and once results cluster near the ceiling the benchmark no longer separates models. Private held-out sets matched to a team's own use cases give decision-grade signal that a public leaderboard cannot.
- What is a dangerous capability evaluation?
- An evaluation designed to test whether a model could meaningfully assist with serious harm, for example in cyber operations or in providing biological or chemical uplift, rather than how useful it is. Frontier developers with published safety frameworks tie deployment decisions and required safeguards to results on these evaluations, so in practice they operate as release gates.
- Is model evaluation an engineering job or a research job?
- Both titles exist. Evaluation engineers build harnesses, grading pipelines, and the infrastructure to run suites reliably at scale. Evaluation researchers design what to measure, validate that a metric tracks the underlying ability, and analyze results. Human data operations, covering annotator training, rubrics, and quality control, is a third distinct track hiring alongside them.