AI-Ops

AIOps / AI Infrastructure Engineer

Latticework Compute

Denver, CO (hybrid, 1 day/week onsite)Full-Time · Hybrid$145K/yr - $175K/yr

Build the GPU cluster scheduling and observability layer behind a growing AI compute provider - bin-packing, autoscaling, and telemetry for workloads that are expensive to run and expensive to get wrong.

About Latticework Compute

Latticework Compute operates GPU infrastructure for AI labs and startups who need reliable, cost-efficient training and inference capacity without building their own data centers. We're a 70-person company competing directly with the major hyperscalers on price and reliability for GPU-heavy workloads.

About The Role

You'll work on the infrastructure layer between raw GPU hardware and the customer workloads running on it - scheduling, bin-packing, autoscaling, and the observability stack that tells us (and our customers) exactly what a multi-thousand-dollar training run is actually doing at any given moment.

Responsibilities

  • Build and tune GPU cluster scheduling and bin-packing to maximize utilization without starving jobs
  • Design autoscaling policies for training and inference workloads with very different cost/latency tradeoffs
  • Build observability tooling giving customers real-time visibility into their GPU workload's health and cost
  • Diagnose and resolve GPU-specific failure modes (NCCL issues, driver mismatches, thermal throttling)
  • Work closely with the hardware ops team on capacity planning and fleet health

Requirements

  • 4+ years in infrastructure engineering, with at least 1+ years touching GPU or HPC workloads
  • Strong Kubernetes experience, ideally including custom schedulers or device plugins
  • Comfortable diagnosing infrastructure issues several layers below a typical web service stack
  • Based in or willing to relocate to the Denver, CO area for the hybrid schedule

Desirables

  • Experience with Slurm, Kubernetes device plugins for GPUs, or a comparable ML scheduler
  • Familiarity with NCCL, CUDA, or other GPU-specific tooling at the infrastructure level
  • Experience building customer-facing observability, not just internal dashboards

Benefits

  • Health, dental, and vision coverage
  • 401(k) with 4% match
  • Hybrid schedule with a well-equipped Denver office (optional, not mandatory beyond 1 day/week)
  • Annual bonus tied to company-wide fleet utilization targets

Extras

  • This role occasionally requires reasoning about hardware-level failure modes - prior data center or hardware-adjacent experience is a plus but not required

More Roles

Other roles you might like

Cookie Notice

Cookies and similar technologies

We use cookies and similar technologies to keep this site working and to ensure the best possible experience for all our users. Kindly confirm your consent to our use of them.

Learn More