AI-Ops
AIOps / AI Infrastructure Engineer
Latticework Compute
Build the GPU cluster scheduling and observability layer behind a growing AI compute provider - bin-packing, autoscaling, and telemetry for workloads that are expensive to run and expensive to get wrong.
About Latticework Compute
Latticework Compute operates GPU infrastructure for AI labs and startups who need reliable, cost-efficient training and inference capacity without building their own data centers. We're a 70-person company competing directly with the major hyperscalers on price and reliability for GPU-heavy workloads.
About The Role
You'll work on the infrastructure layer between raw GPU hardware and the customer workloads running on it - scheduling, bin-packing, autoscaling, and the observability stack that tells us (and our customers) exactly what a multi-thousand-dollar training run is actually doing at any given moment.
Responsibilities
- Build and tune GPU cluster scheduling and bin-packing to maximize utilization without starving jobs
- Design autoscaling policies for training and inference workloads with very different cost/latency tradeoffs
- Build observability tooling giving customers real-time visibility into their GPU workload's health and cost
- Diagnose and resolve GPU-specific failure modes (NCCL issues, driver mismatches, thermal throttling)
- Work closely with the hardware ops team on capacity planning and fleet health
Requirements
- 4+ years in infrastructure engineering, with at least 1+ years touching GPU or HPC workloads
- Strong Kubernetes experience, ideally including custom schedulers or device plugins
- Comfortable diagnosing infrastructure issues several layers below a typical web service stack
- Based in or willing to relocate to the Denver, CO area for the hybrid schedule
Desirables
- Experience with Slurm, Kubernetes device plugins for GPUs, or a comparable ML scheduler
- Familiarity with NCCL, CUDA, or other GPU-specific tooling at the infrastructure level
- Experience building customer-facing observability, not just internal dashboards
Benefits
- Health, dental, and vision coverage
- 401(k) with 4% match
- Hybrid schedule with a well-equipped Denver office (optional, not mandatory beyond 1 day/week)
- Annual bonus tied to company-wide fleet utilization targets
Extras
- This role occasionally requires reasoning about hardware-level failure modes - prior data center or hardware-adjacent experience is a plus but not required
More Roles
Other roles you might like
Cloud Operation
MLOps Engineer
Apple
Build the training and deployment infrastructure behind Solace's recommendation models - feature stores, reproducible pipelines, and the on-call rigor a production ML system actually needs.
MLOps
MLOps Engineer
Solace AI
Build the training and deployment infrastructure behind Solace's recommendation models - feature stores, reproducible pipelines, and the on-call rigor a production ML system actually needs.