Senior Site Reliability Engineer (Night Shift)
This is the employer's own posting, not a copy on a job board.
What we know
Is it still open?
Confirmed still open
Last checked 23h ago — checked against the employer's own applicant tracking system, which is the company answering directly.
We re-read the employer's own applicant tracking system and the posting was still there. That is the company answering directly.
How old is it?
Posted 70d ago
The date the source published, not the day we noticed it (2026-07-07). Last seen at its source 3h ago.
Is it remote?
Marked remote on the employer's board
Their board carries a remote setting on this posting — a field they filled in, not wording we read. The location field names somewhere specific, which is usually where the team or the entity sits.
Who may apply?
India
The description states no restriction of its own. This is the source's own tag.
Pay
$50k–78k/yr
Published by the source in its own salary field. Shown in the posting's own currency and period; we never convert.
Skills named in the ad
Recognised terms only, from a fixed vocabulary — this is what CV matching compares against.
Carried by 1 source
-
lever employer's own board first seen 38d ago · last seen 3h ago
The listing
We are looking for an experienced Site Reliability Engineer (SRE) to join our team. The ideal candidate will be responsible for ensuring high availability, scalability, performance, and reliability of our platform while driving automation and operational excellence across cloud-native environments.
This role offers an opportunity to work at the core of production reliability for a globally used platform. The position is being created to establish dedicated India-based night-time SRE coverage aligned with US business hours, addressing a critical support gap.
The SRE will play a key role in owning production incidents, improving system reliability, reducing MTTR, and driving automation across a modern cloud-native stack (Azure, Kubernetes, Kafka, etc.). This is a high-impact role with direct visibility into business-critical operations and opportunities to work on large-scale distributed.
What You Will Do
- Design, implement, and manage scalable and highly available systems on Azure Cloud
- Monitor system performance, troubleshoot issues, and ensure uptime and reliability
- Manage and optimize Kubernetes clusters and containerized workloads (Docker)
- Build and maintain robust CI/CD pipelines using GitHub Actions and related tools
- Implement infrastructure as code and deployment automation using Helm Charts
- Work with distributed systems such as Kafka, Redis, PostgreSQL, Hadoop/HDFS
- Configure and manage Cloudflare for performance, security, and traffic routing
- Set up monitoring, alerting, and observability using tools like Grafana
- Collaborate with development teams to improve system reliability and deployment practices
- Perform root cause analysis (RCA) and implement preventive measures
- Ensure security best practices and compliance across infrastructure
What You Will Bring
- 6–12 years of experience in SRE/DevOps or related roles
- Strong hands-on experience with Azure Cloud services
- Solid experience in Linux system administration
- Expertise in Docker and Kubernetes (deployment, scaling, troubleshooting)
- Experience with Kafka, Redis, and PostgreSQL
- Working knowledge of Hadoop ecosystem (HDFS, Hadoop)
- Experience with Cloudflare (CDN, security, DNS management)
- Proficiency in CI/CD tools (GitHub, GitHub Actions)
- Experience in Helm Charts and Kubernetes deployments
- Strong understanding of monitoring and logging tools (Grafana, etc.)
What Will Make You Stand Out
- Experience with large-scale distributed systems
- Knowledge of infrastructure automation tools (Terraform, Ansible etc.)
- Exposure to security and compliance best practices
- Strong problem-solving and troubleshooting skills
- Databricks , Clickhouse and ML Ops exposure
Why You Will Love It Here
- Opportunity to work on cutting-edge cloud and distributed systems
- Exposure to large-scale, high-impact platforms
- Collaborative and innovation-driven environment