Stop applying to remote jobs that are already dead.
Every listing shows the evidence: when we last checked it, how, and when it was posted and closed.

All listings

Avra via Ashby

Member of Technical Staff | Observability & Reliability

Sao Paulo senior
still open verified 2d ago posted 18d ago seen 3h ago
Apply at jobs.ashbyhq.com

This is the employer's own posting, not a copy on a job board.

What we know

Is it still open?

Confirmed still open

Last checked 2d ago — checked against the employer's own applicant tracking system, which is the company answering directly.

We re-read the employer's own applicant tracking system and the posting was still there. That is the company answering directly.

Check this listing's status as JSON

How old is it?

Posted 18d ago

The date the source published, not the day we noticed it (2026-09-23). Last seen at its source 3h ago.

We have tracked this listing since 5 Oct 2026 (5 days). The employer's own board has carried it every time we have read it, most recently 3 hours ago.

Is it remote?

Marked remote on the employer's board

Their board carries a remote setting on this posting — a field they filled in, not wording we read. The location field names somewhere specific, which is usually where the team or the entity sits.

Who may apply?

Sao Paulo

The description states no restriction of its own. This is the source's own tag.

Skills named in the ad

AWSGCPIncident ManagementInfrastructure as CodeKubernetesObservabilitySLI/SLOTerraform

Recognised terms only, from a fixed vocabulary — this is what CV matching compares against.

Carried by 1 source

The listing

About the role

At Avra, every technical IC is a Member of Technical Staff (MTS). The title doesn't put anyone in a silo: you own systems and outcomes, not steps in a function, and you keep building depth in your area.

In this role, you'll join the Platform team as our go-to expert on observability and reliability. Our customers make real-time decisions based on our responses, so when we're down, their operations stop. Avra's cloud is just one more dataplane, alongside the dataplanes we operate inside customer environments — so observability and reliability have to work the same way everywhere.

What you'll do

  • Evolve our observability stack for logs, metrics, traces, and alerting.

  • Make sure every dataplane, in our cloud and on-premise, reports its active release, health, heartbeat, logs, metrics, and usage to the control plane.

  • Bring telemetry into customer clusters within a model where agents only make outbound connections.

  • Detect drift between the desired state and what's actually running in each environment.

  • Monitor the health of our deployment and runtime agents.

  • Provide visibility into ephemeral workloads, such as the Ray clusters that run our batch inference.

  • Define SLOs, lead incident response and postmortems, and reduce MTTR — including when a fix requires coordinating with the customer.

  • Reduce telemetry cost: less redundant data, more useful signal.

How we measure success

  • 99.9% serving availability, with incidents trending down.

  • MTTR, including on-premise incidents.

  • Near-zero drift between desired and actual state.

  • All agents active and reporting, across every dataplane.

What we're looking for

  • Deep experience with OpenTelemetry and observability backends.

  • Hands-on practice with SLOs, error budgets, actionable alerting, and incident management.

  • Strong experience with Kubernetes and infrastructure as code (Terraform / Helm ).

  • Experience operating software in environments you don't fully control.

  • Production-quality code and reviews, and a willingness to operate what you build.

Nice to have

  • Shipping software to customer-hosted Kubernetes (e.g., Helm, outbound-only connectivity).

  • GCP or GKE, AWS or EKS.

  • ML multi-node/multi-cluster workloads in production.

  • Financial services or regulated environments.

Apply at jobs.ashbyhq.com