Your brand on every pageBecome a sponsor →
Full exampleFictional companyIllustrative licensed photographyGet yours
Meridian AI logo
Example employer page

AI at Meridian AI

We make large models fast, cheap and honest to evaluate.

3 open roles, updated hourly. Apply directly.New York and LondonView open roles →

Illustrative photo. Meridian AI is a fictional company created to show the full version of an employer-brand page.

Open roles

Demo roles - not real listings
★ Featured roleSponsored

Staff Inference Engineer

Meridian AIUnited StatesFull-time

AI InfrastructureRemote
$220K – $275K
Posted Sep 30

Own the serving stack behind a 70B model: batching, KV-cache and kernel work that turns p95 latency into a number customers notice.

CUDATritonvLLMKubernetes+1Applies directly on Meridian AI's site
★ Featured roleSponsored

Research Engineer, Evaluations

Meridian AILondon, United KingdomFull-time

Evals & Red TeamingHybrid
£110K – £140K
Posted Oct 1

Build the eval suites that decide whether a model ships. You will design the benchmarks, not just run them.

PythonPyTorchEval harnessesStatisticsApplies directly on Meridian AI's site
★ Featured roleSponsored

Applied ML Engineer, Retrieval

Meridian AINew York, United StatesFull-time

ML EngineeringHybrid
$180K – $220K
Posted Sep 29

Ship the retrieval layer that grounds our assistant in customer documents, and measure every change against a real eval set.

RetrievalEmbeddingsPythonPostgresApplies directly on Meridian AI's site

On a real employer page your current openings appear here automatically, from the roles we already read from your applicant tracking system.

Meridian AI is a 60-person applied-AI lab split between New York and London. It builds inference and evaluation tooling for teams that want to run capable models without a frontier-lab budget. Small teams own a system end to end: the people who write the serving code carry its pager, and research engineers ship their own evals to production.

At a glance

Team
60 people: 22 research engineers, 26 ML and infrastructure engineers, 12 in product and operations
HQ
New York and London
Work model
Hybrid or remote (US, UK)
Open roles
3

What it's like to work on AI here

Research That Reaches Production

Research engineers do not hand results over a wall. Whoever proposes a model change pairs with the serving team until it is live, and the change is only "done" when it is running behind production traffic with its eval attached.

  • Every research proposal names the production owner before work starts, in the same design doc.
  • Last year, 14 of 19 research projects shipped to customers; the other 5 were closed with a written reason.

“At my last lab a good result meant a slide. Here it means I am on the rollout channel the week it ships.”

Ana P., Research Engineer
Research That Reaches Production
Real Compute, Without a Queue Fight

Real Compute, Without a Queue Fight

Every researcher has a standing GPU quota that does not need a ticket, and larger training runs go through a weekly allocation meeting with published criteria, so nobody wins compute by lobbying the loudest.

  • A personal quota of 8 GPUs for experiments, available the day you join.
  • The allocation log for large runs is open to the whole company, with the reason each request was approved or deferred.

Evaluation Before Claims

No model change ships on "it feels better". Each one carries the eval that justified it, the eval suite has its own owners, and a result that does not reproduce on a second run is treated as a finding, not an embarrassment.

  • A release gate in CI blocks a model promotion without an attached eval report.
  • A two-person evaluations team owns the benchmark suite and can veto a launch.

“I have killed my own launch twice because the eval did not hold up. Nobody argued with the number.”

Tom W., Staff ML Engineer
Evaluation Before Claims
Hard, Specific Technical Problems

Hard, Specific Technical Problems

The core problem is serving a 70B-parameter model at a p95 latency customers notice, on hardware they can afford. That means batching, KV-cache and kernel work where a 5% win is worth a quarter of effort.

  • Latency and cost per million tokens are tracked on a public internal dashboard, per model, per week.
  • The inference team maintains its own fork of the serving stack, with patches upstreamed where possible.

Transparent, Fair On-Call

On-call rotates one week in six, is paid as a flat stipend on top of salary however many pages come in, and median pages per shift is tracked openly. It has stayed under 3 for four quarters after a push to delete noisy alerts.

  • On-call pay is written into offer letters, not negotiated case by case.
  • Any alert that fires without action twice in a month is reviewed for deletion.
Transparent, Fair On-Call

Learning, Conference and Compute Budget

Everyone gets a $5,000 annual budget for compute, conferences or courses, self-service up to the cap. Each quarter the whole lab takes a four-day research sprint to work on anything that might matter next year.

  • $5,000 a year, no manager approval under the cap.
  • Four sprint days a quarter, protected from roadmap work.

Structured Mentorship and Career Growth

A six-month mentorship pairing runs twice a year for anyone moving tracks, and moving from ML engineering into research engineering is a normal path here, not an exception.

  • A named mentorship cohort with its own lead, twice a year.
  • 4 of the current research engineers joined as ML or infrastructure engineers.

“I joined to run training infrastructure. Eighteen months later I was designing the evals.”

Sam O., Research Engineer, Evaluations
Structured Mentorship and Career Growth
Safety and Responsible Deployment in Practice

Safety and Responsible Deployment in Practice

Before a model is offered to customers it goes through an internal red-team week run by people who did not build it, and the release notes say what was tested and what was not.

  • A red-team week is part of the release checklist, with its own sign-off.
  • Two releases in the last year were held back until findings were fixed.

How the team is organised

Team size
60 people: 22 research engineers, 26 ML and infrastructure engineers, 12 in product and operations
Reporting
Research and engineering both report to the CTO
Functions
Inference, Evaluations, Applied Research, ML Platform
Collaboration
Research engineers pair with a production owner on every project

The work

Problems
Low-latency, low-cost serving of large models, and evaluation you can trust before a release
Stack
PyTorch, Triton, vLLM, Kubernetes, Rust for the serving path
Compute
A mix of owned GPUs and cloud capacity, allocated weekly
Research
Papers are welcome and have been published at workshops, but a shipped improvement is the goal

How we work

Hybrid in New York and London, or remote within US and UK time zones, with a four-hour overlap for synchronous work. The whole lab meets in person twice a year.

Hours
Flexible outside the four-hour overlap
On-call
One week in six, paid, median under 3 pages per shift

Growth

Learning budget
$5,000 a year for compute, conferences or courses
Compute
8 GPUs on a standing personal quota
Conferences
From the same $5,000 budget
Mentorship
Twice-yearly six-month cohort for anyone moving tracks

Benefits and perks

  • Salary range published on every role
  • $5,000 annual compute and conference budget
  • Four-day research sprint each quarter
  • 16 weeks paid parental leave
  • Remote within US and UK time zones

What the interview looks like

  1. Intro call

    Your background, what you want next, and logistics.

    30 min · The hiring manager

  2. Take-home or pairing

    Your choice: a short take-home or a pairing session on a real (simplified) problem from the team.

    Up to 3 hours · An engineer from the team

  3. Team loop

    A system-design session and a walkthrough of a paper or code you know well.

    2 hours · Two engineers and a research engineer

  4. Decision

    You hear back with a decision and the reasons.

    Within 5 working days

Before you interview

What to expect

  • A real serving or evaluation problem from the team you would join, simplified to fit the time.
  • A paper or piece of code you know well, and what you would change about it.
  • How you decide whether a result is real before you claim it.

What we evaluate

  • Depth on one problem over a list of tools.
  • How you reason about trade-offs between quality, latency and cost.
  • Whether you ask what "better" means before optimising for it.

Come ready to discuss

  • One model or system you shipped, and what you measured.
  • A result of yours that did not hold up, and what you learned.
  • A question about how compute is allocated or how evals gate a release.

Why this work matters

Faster, cheaper inference with evaluation you can trust puts capable models within reach of teams that could never afford them.

Meet the team

Ana P.

Ana P.

Research Engineer

“A good result here means I am on the rollout channel the week it ships.”

Tom W.

Tom W.

Staff ML Engineer

“I have killed my own launch twice because the eval did not hold up.”

Raj M.

Raj M.

Inference Engineer

“A 5% latency win is a real project here, not a side quest.”

Sam O.

Sam O.

Research Engineer, Evaluations

“I joined to run training infrastructure and now design the evals.”

Lena H.

Lena H.

CTO

“If a number does not reproduce, we say so in the release notes.”

More from the team

Want your AI team presented like this?

Show AI candidates what makes your company different, on the page they already land on.

Get your employer page →

Example employer page - Meridian AI is a fictional company.

Want your company's version?

We build it with you from your own content and photography, and nothing goes live until you have checked it.

Request pricing →