SRE
Site Reliability Engineer
Cascadia Systems
Join a 3-person SRE team responsible for the reliability of a payments platform processing 2M+ transactions a day - SLOs, incident response, and reliability engineering that actually moves the needle.
About Cascadia Systems
Cascadia Systems provides payment infrastructure for regional banks and credit unions across the Pacific Northwest. Reliability isn't a nice-to-have for us - a bad incident means real money not moving for real people. We're a 60-person company, Series C, with a strong engineering culture around ownership and blameless postmortems.
About The Role
As an SRE at Cascadia, you'll define and defend SLOs for our core payments services, lead incident response for the platform's most critical paths, and drive reliability work upstream into product engineering teams rather than absorbing it downstream as toil.
Responsibilities
- Define, measure, and defend SLOs and error budgets for core payment processing services
- Serve as incident commander on a rotating basis for Sev1/Sev2 incidents
- Run blameless postmortems and track remediation items to completion
- Build and maintain observability tooling (Prometheus, Grafana, distributed tracing)
- Partner with product engineering teams to bake reliability into design reviews, not bolt it on after launch
- Reduce recurring toil through automation - anything done manually twice gets scripted
Requirements
- 4+ years in an SRE, DevOps, or production engineering role
- Direct experience owning SLOs and error budgets, not just dashboards
- Strong incident response and postmortem-writing skills
- Proficiency with at least one of Prometheus, Datadog, or an equivalent observability stack
- Comfortable working in a hybrid schedule based in or relocating to the Seattle area
Desirables
- Experience in a regulated or high-stakes domain (fintech, healthcare, payments)
- Familiarity with chaos engineering practices
- Experience running error-budget-driven release policies
Benefits
- Health, dental, and vision fully covered for employees and dependents
- 401(k) with 4% match
- Hybrid schedule with flexible core hours
- Annual reliability conference attendance (SREcon, Monitorama, or similar)
Extras
- On-call is compensated separately from base salary, paid per rotation
More Roles
Other roles you might like
Cloud Operation
MLOps Engineer
Apple
Build the training and deployment infrastructure behind Solace's recommendation models - feature stores, reproducible pipelines, and the on-call rigor a production ML system actually needs.
MLOps
MLOps Engineer
Solace AI
Build the training and deployment infrastructure behind Solace's recommendation models - feature stores, reproducible pipelines, and the on-call rigor a production ML system actually needs.