Sr. Site Reliability Engineer
This is the employer's own posting, not a copy on a job board.
What we know
Is it still open?
Confirmed still open
Last checked 1d ago — checked against the employer's own applicant tracking system, which is the company answering directly.
We re-read the employer's own applicant tracking system and the posting was still there. That is the company answering directly.
How old is it?
Posted 40d ago
The date the source published, not the day we noticed it (2026-08-05). Last seen at its source 1h ago.
Is it remote?
US Remote
That is the location the employer filed this posting under. Quoted as written — we do not re-word the source's own location.
Who may apply?
United States
The description states no restriction of its own. This is the source's own tag.
Pay not stated
Similar roles pay $160.9k–210k/yr
Middle 50% of 740 listings that do state pay — Engineering · Senior · United States · USD/year. This employer has published no salary; this is what comparable listings we hold disclose, never converted between currencies or periods. How this is calculated.
Skills named in the ad
Recognised terms only, from a fixed vocabulary — this is what CV matching compares against.
Carried by 1 source
-
ashby employer's own board first seen 11d ago · last seen 1h ago
The listing
About the Role
We are seeking a Senior Site Reliability Engineer to join our cloud engineering team. You will own the reliability, scalability, and observability of our critical financial SaaS applications and infrastructure, working across cloud platforms to ensure our customers experience is seamless, secure, and performant services. This is a high-impact role for someone who is passionate about building resilient systems and preventing outages before they happen.
Key Responsibilities
Design, implement, and maintain Service Level Objectives (SLOs) and Service Level Indicators (SLIs) across all critical systems; ensure we meet or exceed targets consistently
Lead observability strategy by designing comprehensive monitoring, logging, and tracing architectures; select and deploy observability tools that provide deep visibility into system behavior
Build and own runbooks, incident response procedures, and post-incident review processes; mentor the team on incident management and blameless postmortems
Architect and deploy cloud infrastructure on AWS or Azure; implement infrastructure-as-code practices and ensure high availability, disaster recovery, and business continuity
Develop automation and AIOps capabilities to reduce toil, accelerate incident detection, and enable self-healing systems; implement intelligent alerting to minimize false positives
Drive reliability improvements through load testing, chaos engineering, and failure scenario analysis; identify and eliminate single points of failure
Partner with application and backend teams to design reliable systems from inception; conduct architecture reviews and reliability assessments
Write production-grade Python tooling for automation, metrics collection, alert management, and operational workflows
Champion security and compliance in infrastructure; implement defense-in-depth principles for a regulated fintech environment
Required Qualifications
7+ years in Site Reliability Engineering, DevOps, platform engineering, or closely related roles with significant responsibility for production systems
Expert-level experience with Azure or AWS (or both); deep knowledge of compute, networking, storage, and managed services; experience managing infrastructure at scale
Demonstrated expertise in observability: designing and implementing monitoring, alerting, logging, and distributed tracing solutions; hands-on with observability platforms (e.g., Prometheus, Grafana, ELK, Datadog, New Relic, or similar)
Strong background in SLOs, SLIs, and SLAs; experience defining meaningful objectives and building systems to meet them; understanding of error budgets and their role in prioritization
Proven experience designing and troubleshooting highly available, resilient, and scalable systems; deep understanding of distributed systems concepts and failure modes
Proficiency in Python, PowerShell, bash, etc. scripting languages for production automation, tooling, and systems programming; ability to write clean, maintainable code for operational workflows
Hands-on experience with AIOps practices: event correlation, intelligent alerting, predictive analytics, and automated remediation; familiarity with AIOps platforms is a plus
Experience with infrastructure-as-code tools (e.g., Terraform, CloudFormation, Ansible); version control and CI/CD pipeline design
Track record of incident management and on-call ownership; comfort with incident response and the ability to remain calm under pressure
Excellent communication skills; ability to work cross-functionally and influence without authority; comfort mentoring junior engineers
Preferred Qualifications
Experience in the fintech, payments, banking, or other regulated industries; understanding of compliance requirements (SOC 2, PCI-DSS, etc.)
Experience with Kubernetes and container orchestration; deep knowledge of containerized application deployment and management
Proficiency with observability as code; experience building custom metrics, dashboards, and alerts programmatically
Background in chaos engineering or reliability testing; experience using tools like Gremlin or similar platforms
Contribution to open-source observability or infrastructure projects
Expertise in network security, application security, or infrastructure hardening
Experience with database optimization, query performance tuning, and backup/recovery strategies