AI Infrastructure Jobs
AI infrastructure engineers build the compute layer every model team depends on - GPU clusters, distributed training systems, inference serving, and the ML platform tooling that turns a model from a research artifact into something that runs reliably at scale.
AI Infrastructure roles are rare: 154 live right now. Get the new ones every Monday.
Get new AI Infrastructure jobs in your inbox
Join 100+ AI professionals · Weekly, free, unsubscribe anytime
Latest AI Infrastructure roles
CVS HealthNY, New York+3 more
VerizonBasking Ridge, New Jersey+3 more
HP Inc.Spring, Texas+1 more
MastercardBudapest, Hungary
General MotorsSunnyvale, California+2 more
NVIDIAUnited States+3 more
Bank of AmericaNew York, United States of America+2 more
Constrained by silicon, power, and cost
AI infrastructure work is shaped by three binding constraints that ordinary software rarely faces. Accelerator supply is finite and allocated ahead of delivery, so capacity planning is a procurement problem as much as a scheduling one. Datacenter power and cooling limit where and how quickly capacity can be added, on timelines set by grid interconnection and construction rather than by engineering. And inference cost is a recurring operating expense that grows with product success rather than a fixed platform bill.
At this scale hardware failure is expected rather than exceptional. A large training run occupies a very large accelerator fleet continuously for an extended period, and across that many components and that much elapsed time, node failures, network faults, and silent data corruption are near-certain. Checkpointing frequency, fast restart, straggler detection, health checking, and elastic reconfiguration are therefore product features of the training platform rather than operational hygiene bolted on afterwards, and interviews in this area get to them quickly.
The serving side is one long campaign against cost per token. Continuous batching, KV cache management and paged attention, quantization, speculative decoding, kernel-level optimization, distillation to smaller models, tensor and pipeline parallelism, and routing between model tiers all exist to move the same curve: more useful output per accelerator-second, within a latency budget the product can tolerate. Roles here sit between machine learning and systems engineering and reward people who read profiles carefully.
Explore related searches
- Browse the Machine Learning Engineer hub for the full specialization.
- Prefer remote? See remote AI jobs.
- AI Engineer jobs
Frequently asked questions
- How many live AI Infrastructure openings are there?
- We are tracking 154 live AI Infrastructure roles across the AI companies we monitor, updated hourly. Each listing links straight to the employer's own application page.
- What limits how fast AI compute capacity can grow?
- Accelerator supply, which is allocated well ahead of delivery, and datacenter power and cooling, where grid interconnection and construction set multi-year timelines. Networking and storage bandwidth bind at scale as well. Software optimization matters, but capacity planning in this field is a procurement and facilities problem before it is an engineering one.
- Why is fault tolerance a core feature of training infrastructure?
- Because a large training run occupies a very large accelerator fleet continuously for an extended period, so component failures, network faults, and silent data corruption are expected within a single run. Frequent checkpointing, fast restart from checkpoint, straggler and health detection, and the ability to reconfigure around lost nodes determine whether the run completes at all.
- What skills do inference optimization roles hire for?
- Profiling and performance analysis, GPU kernel work in CUDA or comparable toolchains, familiarity with serving stacks and batching strategies, KV cache and memory management, quantization, and parallelism across devices. Employers value candidates who can connect a latency or throughput target to a specific cost per token rather than optimizing in the abstract.