Your brand on every pageBecome a sponsor →
arcadeai logo

Founding Machine Learning Engineer

Arcade.dev

San Francisco, CA

ML Engineering

Posted 3 hours ago

midonsite

Verified open on Oct 5, 2026 · posted today

Job Description

Everyone's building AI agents, but almost nobody gets them to production.

Building an impressive demo is easy. Building an AI agent that can securely take action inside enterprise systems is hard. The moment an agent accesses customer data, executes a workflow, or makes changes on behalf of a user, authorization, governance, and trust become the real engineering challenge.

Arcade is the MCP runtime that gives agents the power to do both seamlessly. We connect agents to the systems they act in, then give each one a permission slip and a paper trail - proof of what it's allowed to do, and a record of what it did. That's what makes AI safe to turn loose: real actions, on real systems, already shipping inside Fortune 100 companies.

The Revolution Needs You

Every AI app needs agentic "tools" - special functions that let AI models take real actions. Without tools, AI can only chat. With tools, AI can actually do things. We're building the definitive tools catalog and tool-calling platform that will unlock AI's true potential. Think Zapier for AI Actions. Think Auth0 for AI. Think really big.

Why This Is The Opportunity of a Lifetime

  • Traction: Real deployments with Fortune-100 customers like Morgan Stanley and Open Table

  • Founder-Market Fit: Our CEO previously founded Stormpath (acquired by Okta), where he created the first Authentication API for developers. He's done this before - and this time the market is 10x bigger. Our CTO led the vector database team at Redis, shipped 100+ LLM applications, and is a contributor to LangChain and LlamaIndex. He knows this space better than anyone.

  • Dream Team: We've assembled authentication, integrations, distributed systems, and AI experts from Okta, Redis, Microsoft, Splunk, Ngrok, Google, Airbyte, Disney, and HPE who've built and founded multiple successful developer platforms.

  • Perfect Timing: Every enterprise is racing to put agents in production - almost none get there. The problem isn't better models, it's proving which agent can take which action, on behalf of which user, against which system. That's us.

  • Massive Market : We're building critical infrastructure for the biggest technological shift of our generation. Every AI app will need what we're building.

  • Backed By The Best: Our Series A round is led by SYN Ventures, with strategic investment from Morgan Stanley and Wipro. Our earlier investors have also backed Databricks, Clickhouse, MongoDB, Perplexity, Cohere, ScaleAI, Confluent, Elastic, and Firebase. They see what we see - this is going to be huge.

The Challenge

Arcade is hiring the principal machine learning engineer that wants to build the future of agent capabilities.

These core problems define the role:

  • What’s old is new: How can a classic two-tower e-commerce recommendation system be used to make agents smarter? How can a classifier be used to make agents more efficient? The ability to take what’s been done and apply it to a new and ever-changing field is paramount for this role. We’ve launched Tool Recommendations as our first foray into agent–led recommendations, and there is a long list of features we want to create on these (& similar) building blocks.

  • Own the models: Own the model pipeline from data ingestion to model publishing lifecycle - including models across embedding, reranking, classification, recognition, and other tasks. Be the leader of a group making agent capabilities, not just API calls, but intelligent and efficient abilities that outpace the competition. Provide insight and experience bringing ideas and practices that should be implemented in a stable but young model pipeline.

  • Evaluation Obsession: Be willing to stand by your models because you’ve been provided able data to cover every possible outcome you can. Stand by the strengths and be readily willing to admit the faults so the team can be prepared. Be willing to go row by row in a spreadsheet without pride. That kind of activity is “beneath” you because this space is too new to claim any ego.

  • Enterprise ready: Everything you build has to run where our customers run: their cloud, their hardware, and sometimes an air-gapped network. Quantization, export, serving, and packaging are part of the model design from day one, not a deployment step at the end.

Our largest customers run Arcade inside their own walls. A Fortune 100 bank doesn't send its agents' tool calls to someone else's API, and its agents can't wait on a frontier model to pick the right tool out of thousands. So the models behind Arcade's agentic features have to be small, fast, accurate, and shippable into a customer's VPC.

This is just one example of why custom, small models are the unlock to many of Arcade’s future products. You'll report directly to the Head of Engineering and own the models that sit in the runtime path of every agent call we serve.

Our first model for agent recommendation is already built (& patent pending), along with its training pipeline, but it needs to be productionized. You'll own it, decide what the ML stack at Arcade looks like, and set the patterns going forward. You’ll own the build-buy decisions for our stack going forward and have a healthy budget to spend.

This role is about shipping. While we're happy to publish what we learn, delivering the product to customers comes first. If you want six months in a notebook before anything reaches production, this isn't the right role. If you want to ship the model and the writeup in the same quarter, it is!

What You'll Do

  • Own the training pipeline: Run Arcade's ML pipeline end to end, from data and training through evaluation and release. Make it reliable enough that shipping a new model is routine, not an event.

  • Build the models: Fine-tune and train models for tool selection, routing, retrieval, and agent memory. Use whatever gets the job done distillation, embeddings, rerankers and more.

  • Expand the use cases: Take our model into new territory, starting with search and recommendation over tools and agent context.

  • Measure what matters: Build the eval system. That means offline evals against real agent traces, online measurement in production, and head-to-head comparisons with Claude, GPT, and Gemini, so every model decision has data behind it.

  • Ship on-prem: Quantize, optimize, and package models for customer VPCs and air-gapped environments, and work with the Runtime team on how they're served.

  • Turn telemetry into training data: Build the loop from production agent traces and tool-call data to better models, within the data boundaries our enterprise customers require.

  • Set the direction: Choose the ML stack, write the strategy, and help decide who we hire next into ML.

  • Decide what to build: Use your depth in modern agent systems (harnesses, memory, skills) to pick where models make agents meaningfully better, and to call early whether an approach is going to work.

  • Use AI to compound your own output: Projects that take a week today should take a day next time.

Required Skills

  • 7+ years of software engineering experience, with 4+ years training and shipping production ML systems. Formal title matters less than the work.

  • You've trained or fine-tuned models that went to production and moved a metric customers cared about, not just a leaderboard number.

  • Expertise in production agent systems: harnesses, memory, skills, tool use, and sub-agents. You know the nuances well enough to help decide what we build, and have intuition about what will actually work before we build it.

  • Know how and when fine-tuning works and how to apply it. What data is needed, and how to derive it from what raw data telemetry gives you.

  • Evals you'd defend in a review. You have the statistics fluency to say whether a small delta is real or noise.

  • You've deployed models under real constraints like tight latency budgets, limited GPUs, or someone else's infrastructure (vLLM, ONNX, TensorRT, llama.cpp, or equivalent).

  • Strong Python for training and ML work, plus TypeScript or Go for the production services that serve your models.

  • You pick the stack, write the doc, and can still defend the decision a year later.

  • A do-er, not a researcher-in-residence. You'd rather ship a working v0.5 next week than a polished v2.0 next quarter.

  • Comfort with ambiguity: early team, a charter that will expand, decisions made with incomplete data.

  • An insatiable desire to ship.

Bonus Points

  • You've shipped ML into enterprise on-prem or regulated environments (financial services, healthcare, government).

  • Tool-use benchmark or eval work: BFCL, τ-bench, ToolBench, MCP evals, or equivalent.

  • Familiarity with the MCP (Model Context Protocol) ecosystem. Extra bonus if you've filed an issue against the spec.

  • You've built training-data pipelines from production traces with privacy controls in place.

  • You've been the first ML hire somewhere before, and you'd do some things differently this time.

  • Open-source contributions or published work that survived contact with other engineers.

  • Experience at an early-stage startup, and you loved it.

Join The Movement

We're not just building a product - we're leading a movement to transform AI from just chatbots to agents that can take actions against real systems. This is your chance to be at the forefront of that revolution.

If you want to look back in 5 years and say, "I helped build that", then we want to talk to you. Ready to make AI actually useful? Apply Now

Compensation and Benefits

This role is in person at our San Francisco office, and offers a competitive salary, equity, and benefits. Compensation is aligned with the range below and determined based on a candidate's background, experience, and performance.

Compensation: Starting at $230,000 base salary, plus equity and competitive benefits.

Share this role:WhatsAppLinkedInX

Not ready to apply?

Get new ML Engineering jobs in your inbox

Join 100+ AI professionals · Weekly, free, unsubscribe anytime

Career context

This role is classified as ML Engineering.

Preparing to interview? Read the ML Engineering interview guide.

$200k - $281k is the middle 50% of disclosed salaries, measured from 252 live ML Engineering postings on this board. Roughly two thirds of postings disclose nothing, so this describes the ones that do, not the whole market.

Share your AI salary

Anonymous. No login, no name, no email required. It takes about 20 seconds, and it is how this board builds real pay data for AI roles instead of guesses.

Before bonus and equity.

Hiring for a role like this?

Reach AI professionals browsing the board - your listing goes live instantly.

Post a job →

Related AI jobs

Join 100+ AI professionals. Get the week's new AI jobs in your inbox. Free, unsubscribe anytime.