Senior Site Reliability Engineer
This is the employer's own posting, not a copy on a job board.
What we know
Is it still open?
Confirmed still open
Last checked 1d ago — checked against the employer's own applicant tracking system, which is the company answering directly.
We re-read the employer's own applicant tracking system and the posting was still there. That is the company answering directly.
How old is it?
Posted 34d ago
The date the source published, not the day we noticed it (2026-08-11). Last seen at its source 1h ago.
Is it remote?
US Remote
That is the location the employer filed this posting under. Quoted as written — we do not re-word the source's own location.
Who may apply?
United States
The description states no restriction of its own. This is the source's own tag.
Pay not stated
Similar roles pay $160.9k–210k/yr
Middle 50% of 740 listings that do state pay — Engineering · Senior · United States · USD/year. This employer has published no salary; this is what comparable listings we hold disclose, never converted between currencies or periods. How this is calculated.
Skills named in the ad
Recognised terms only, from a fixed vocabulary — this is what CV matching compares against.
Carried by 1 source
-
ashby employer's own board first seen 11d ago · last seen 1h ago
The listing
As a Senior Site Reliability Engineer on our cloud engineering team, you'll keep our production environment healthy, secure, and running smoothly. This is an operations-focused role: you'll own the day-to-day administration of our AWS accounts and databases, backup posture across our data stores, and production monitoring and debugging for a fully serverless platform. Your work will span the operational side of the software development life cycle — from deployment to maintenance and updates — always striving for continuous improvement. You'll keep our infrastructure clean, easily deployable, and scalable, creating a stable operating environment for the whole team.
Responsibilities
Own day-to-day administration across AWS services, accounts, and access, as well as database administration across PostgreSQL and our other data stores.
Own backup posture across databases, S3 buckets, and queues; verify restores regularly and maintain a tested disaster recovery plan.
Proactively monitor production — CloudWatch dashboards, metric alarms, log-based metrics, and Slack alerting — addressing operational issues before they impact users.
Lead production debugging and incident response: build and maintain runbooks, participate in the on-call rotation, and resolve queue and dead-letter-queue failures through retry, redrive, and recovery.
Continuously refine our infrastructure to ensure it is easily deployable and scalable: keep infrastructure as code (SST/Pulumi) accurate, retire unused infrastructure, and keep cost visible and justified.
Share your knowledge of production operations with the team, fostering a culture of learning and growth.
Qualifications: Knowledge, Skills, & Abilities
Bachelor's degree and 4-6 years of related experience or equivalent work experience.
5+ years of experience in DevOps, site reliability, or platform operations, with significant responsibility for production systems.
3+ years of hands-on experience with AWS, with an emphasis on serverless services (Lambda, SQS, EventBridge, CloudWatch, S3).
Strong database administration experience: PostgreSQL operations, backup and recovery, and query performance; comfort administering other data stores.
Proficiency in scripting languages such as TypeScript, Python, and bash for production automation and operational tooling.
Strong understanding of Linux, DNS, TLS, Docker, GitHub Actions, and infrastructure as code (SST, Pulumi, or Terraform).
Experience with production monitoring and alerting, incident response, and on-call ownership.