MLOps
MLOps Engineer
Solace AI
Build the training and deployment infrastructure behind Solace's recommendation models - feature stores, reproducible pipelines, and the on-call rigor a production ML system actually needs.
About Solace AI
Solace AI builds recommendation and personalization infrastructure for e-commerce platforms. Our models sit directly in the checkout path for several large retail customers, which means model reliability is a revenue problem, not just an accuracy metric. We're a 45-person, Series B company.
About The Role
You'll build and operate the infrastructure connecting our data scientists' training pipelines to production inference - feature store consistency, reproducible retraining, and deployment gates that catch a regressed model before it ever reaches a customer's checkout flow.
Responsibilities
- Own the feature store ensuring training/serving parity across our recommendation models
- Build reproducible, versioned training pipelines (data, code, and model versioning as one unit)
- Implement evaluation gates that block deployment of a model that regresses key metrics
- Operate the online inference service, including latency and availability SLOs
- Partner with data scientists to productionize research code without rewriting it from scratch each time
- Carry on-call for the production inference path (roughly 1 week in 6)
Requirements
- 3+ years building production ML infrastructure, not just training models
- Strong Python and working knowledge of at least one ML framework (PyTorch or TensorFlow)
- Experience with container orchestration (Kubernetes) for serving ML workloads
- Comfortable owning the reliability of a system data scientists depend on daily
Desirables
- Experience with a feature store platform (Feast, Tecton, or an in-house equivalent)
- Familiarity with data/model versioning tools (DVC, MLflow, or similar)
- Prior experience where a model's failure had direct, measurable business impact
Benefits
- Fully remote with an annual company retreat
- Health, dental, and vision coverage (US); equivalent stipend for international hires
- Meaningful equity as an early infrastructure hire on the ML platform team
- $1,500/yr home office stipend
Extras
- Team currently spans 4 countries - written async communication is a core part of how we work
More Roles
Other roles you might like
Cloud Operation
MLOps Engineer
Apple
Build the training and deployment infrastructure behind Solace's recommendation models - feature stores, reproducible pipelines, and the on-call rigor a production ML system actually needs.
DevOps
Junior DevOps Engineer (Internship)
Beacon Cloud Works
A structured 12-week internship building real infrastructure tooling alongside a small, senior DevOps team - not a coffee-fetching internship, an actual shipped-project internship.