Staff Inference Engineer
Meridian AIUnited StatesFull-time
Own the serving stack behind a 70B model: batching, KV-cache and kernel work that turns p95 latency into a number customers notice.
We make large models fast, cheap and honest to evaluate.
Illustrative photo. Meridian AI is a fictional company created to show the full version of an employer-brand page.
Meridian AIUnited StatesFull-time
Own the serving stack behind a 70B model: batching, KV-cache and kernel work that turns p95 latency into a number customers notice.
Meridian AILondon, United KingdomFull-time
Build the eval suites that decide whether a model ships. You will design the benchmarks, not just run them.
Meridian AINew York, United StatesFull-time
Ship the retrieval layer that grounds our assistant in customer documents, and measure every change against a real eval set.
On a real employer page your current openings appear here automatically, from the roles we already read from your applicant tracking system.
Meridian AI is a 60-person applied-AI lab split between New York and London. It builds inference and evaluation tooling for teams that want to run capable models without a frontier-lab budget. Small teams own a system end to end: the people who write the serving code carry its pager, and research engineers ship their own evals to production.
What it's like to work on AI here
Research engineers do not hand results over a wall. Whoever proposes a model change pairs with the serving team until it is live, and the change is only "done" when it is running behind production traffic with its eval attached.
“At my last lab a good result meant a slide. Here it means I am on the rollout channel the week it ships.”
Every researcher has a standing GPU quota that does not need a ticket, and larger training runs go through a weekly allocation meeting with published criteria, so nobody wins compute by lobbying the loudest.
No model change ships on "it feels better". Each one carries the eval that justified it, the eval suite has its own owners, and a result that does not reproduce on a second run is treated as a finding, not an embarrassment.
“I have killed my own launch twice because the eval did not hold up. Nobody argued with the number.”
The core problem is serving a 70B-parameter model at a p95 latency customers notice, on hardware they can afford. That means batching, KV-cache and kernel work where a 5% win is worth a quarter of effort.
On-call rotates one week in six, is paid as a flat stipend on top of salary however many pages come in, and median pages per shift is tracked openly. It has stayed under 3 for four quarters after a push to delete noisy alerts.
Everyone gets a $5,000 annual budget for compute, conferences or courses, self-service up to the cap. Each quarter the whole lab takes a four-day research sprint to work on anything that might matter next year.
A six-month mentorship pairing runs twice a year for anyone moving tracks, and moving from ML engineering into research engineering is a normal path here, not an exception.
“I joined to run training infrastructure. Eighteen months later I was designing the evals.”
Before a model is offered to customers it goes through an internal red-team week run by people who did not build it, and the release notes say what was tested and what was not.
Hybrid in New York and London, or remote within US and UK time zones, with a four-hour overlap for synchronous work. The whole lab meets in person twice a year.
Intro call
Your background, what you want next, and logistics.
30 min · The hiring manager
Take-home or pairing
Your choice: a short take-home or a pairing session on a real (simplified) problem from the team.
Up to 3 hours · An engineer from the team
Team loop
A system-design session and a walkthrough of a paper or code you know well.
2 hours · Two engineers and a research engineer
Decision
You hear back with a decision and the reasons.
Within 5 working days
What to expect
What we evaluate
Come ready to discuss
Faster, cheaper inference with evaluation you can trust puts capable models within reach of teams that could never afford them.
Ana P.
“A good result here means I am on the rollout channel the week it ships.”
Tom W.
“I have killed my own launch twice because the eval did not hold up.”
Raj M.
“A 5% latency win is a real project here, not a side quest.”
Sam O.
“I joined to run training infrastructure and now design the evals.”
Lena H.
“If a number does not reproduce, we say so in the release notes.”
Want your AI team presented like this?
Show AI candidates what makes your company different, on the page they already land on.
Get your employer page →Example employer page - Meridian AI is a fictional company.
Want your company's version?
We build it with you from your own content and photography, and nothing goes live until you have checked it.
Request pricing →