Your brand on every pageBecome a sponsor →

AI Infrastructure Jobs

AI infrastructure engineers build the compute layer every model team depends on - GPU clusters, distributed training systems, inference serving, and the ML platform tooling that turns a model from a research artifact into something that runs reliably at scale.

154 live AI Infrastructure roles across the AI employers we track - updated hourly, apply directly.

4new AI Infrastructure roles posted this week

AI Infrastructure roles are rare: 154 live right now. Get the new ones every Monday.

Get new AI Infrastructure jobs in your inbox

Join 100+ AI professionals · Weekly, free, unsubscribe anytime

Latest AI Infrastructure roles

Constrained by silicon, power, and cost

AI infrastructure work is shaped by three binding constraints that ordinary software rarely faces. Accelerator supply is finite and allocated ahead of delivery, so capacity planning is a procurement problem as much as a scheduling one. Datacenter power and cooling limit where and how quickly capacity can be added, on timelines set by grid interconnection and construction rather than by engineering. And inference cost is a recurring operating expense that grows with product success rather than a fixed platform bill.

At this scale hardware failure is expected rather than exceptional. A large training run occupies a very large accelerator fleet continuously for an extended period, and across that many components and that much elapsed time, node failures, network faults, and silent data corruption are near-certain. Checkpointing frequency, fast restart, straggler detection, health checking, and elastic reconfiguration are therefore product features of the training platform rather than operational hygiene bolted on afterwards, and interviews in this area get to them quickly.

The serving side is one long campaign against cost per token. Continuous batching, KV cache management and paged attention, quantization, speculative decoding, kernel-level optimization, distillation to smaller models, tensor and pipeline parallelism, and routing between model tiers all exist to move the same curve: more useful output per accelerator-second, within a latency budget the product can tolerate. Roles here sit between machine learning and systems engineering and reward people who read profiles carefully.

Explore related searches

Frequently asked questions

How many live AI Infrastructure openings are there?
We are tracking 154 live AI Infrastructure roles across the AI companies we monitor, updated hourly. Each listing links straight to the employer's own application page.
What limits how fast AI compute capacity can grow?
Accelerator supply, which is allocated well ahead of delivery, and datacenter power and cooling, where grid interconnection and construction set multi-year timelines. Networking and storage bandwidth bind at scale as well. Software optimization matters, but capacity planning in this field is a procurement and facilities problem before it is an engineering one.
Why is fault tolerance a core feature of training infrastructure?
Because a large training run occupies a very large accelerator fleet continuously for an extended period, so component failures, network faults, and silent data corruption are expected within a single run. Frequent checkpointing, fast restart from checkpoint, straggler and health detection, and the ability to reconfigure around lost nodes determine whether the run completes at all.
What skills do inference optimization roles hire for?
Profiling and performance analysis, GPU kernel work in CUDA or comparable toolchains, familiarity with serving stacks and batching strategies, KV cache and memory management, quantization, and parallelism across devices. Employers value candidates who can connect a latency or throughput target to a specific cost per token rather than optimizing in the abstract.

Related AI job searches