Remote SRE Jobs – Senior Site Reliability Engineer (Remote) – $130k‑$170k USD – Full‑Time – Escondido, California – Cloud/DevOps, Kubernetes, Terraform, Prometheus

Remote, USA Full-time
TITLE:Remote SRE Jobs –Senior Site Reliability Engineer (Remote) – $130k‑$170k USD – Full‑Time – Escondido, California – Cloud/DevOps, Kubernetes, Terraform, Prometheus --- Who we are We are a mid‑stage SaaS company that grew from a garage‑side prototype to a platform serving > 200 enterprise customers worldwide. Our flagship product—an API‑driven data‑pipeline—processes ≈ 15 TB of events per day, and we guarantee customers 99.9 % uptime. The engineering culture is built on blunt feedback, data‑driven post‑mortems, and a relentless focus on reliability.While the code lives in the cloud, the heart of our operational decisions is made by a small, tight‑knit crew spread across the globe. Why this role exists now In the last 12 months we added three new data‑centers (AWS us‑east‑1, us‑west‑2 and GCP europe‑west1) to shave latency for European clients. That expansion bumped our monthly alert volume from ≈ 2,800 to ≈ 5,200, and our MTTR climbed from 12 minutes to 18 minutes because the on‑call rotation stretched thin. The leadership team decided it was time to double‑down on site reliability: we need a senior engineer who can own the reliability roadmap, coach the junior members, and tighten our alert fatigue.Where you’ll sit (virtually) Although the job is remote, we have a legal entity in Escondido, California that handles payroll, benefits, and compliance. You’ll be part of a “virtual office” that meets daily in a Slack channel called #sre‑hub, a weekly video‑call huddle, and a quarterly in‑person meetup hosted in Escondido, California when travel permits. Being anchored to Escondido, California helps us stay aligned with local tax regulations and gives you a community of other remote professionals who live in the same time zone.The team you’ll join - Size & composition: 12 engineers total—5 senior SREs, 4 junior reliability engineers, 2 platform developers, and 1 manager. - Current metrics: 99.92 % uptime over the past quarter, 5,200 alerts processed per month, 18‑minute average MTTR, 0.2 % alert fatigue (defined as > 3 alerts per incident). - SLA commitments: 99.9 % availability for all customer‑facing APIs, 99.7 % for internal data‑processing pipelines. What you’ll do day‑to‑day 1. Own reliability initiatives – Define and ship SLOs for new services, write error‑budget policies, and track them in Grafana dashboards.2. Incident ownership – Lead the response during high‑severity incidents, drive the post‑mortem narrative, and ensure actionable remediation items are filed in JIRA within 24 hours. 3. Automation & tooling – Write Terraform modules to provision Kubernetes clusters, build Helm charts for micro‑services, and shrink manual run‑books into reproducible Ansible playbooks. 4. Capacity planning – Run quarterly load‑tests using Locust, model growth with Python scripts, and present forecasts to product leadership.5. Mentorship – Pair up with junior SREs for “bug‑hunting” sessions, run monthly reliability workshops, and contribute to our internal “SRE Playbook”. Who we think will thrive - 5+ years of production‑grade experience with Linux/Unix, networking, and cloud infrastructure (AWS or GCP). - Deep familiarity with monitoring stacks: Prometheus, Grafana, Alertmanager, and log aggregation via Splunk or ELK. - Infrastructure‑as‑Code fluency: Terraform ≥ 0.13, Helm ≥ 3, and Ansible. - Container orchestration: Running production workloads on Kubernetes (experience with EKS or GKE).- Programming: Comfortable writing Python or Go for automation; Bash scripting is a given. - Incident mindset: You can stay calm under pressure, triage noisy alerts, and keep a clear incident timeline. - Communication: Able to explain complex reliability concepts to product managers and non‑technical stakeholders in plain language. Tools & tech stack (the ones we actually use) - Cloud – AWS (EC2, RDS, S3, Lambda) and GCP (Compute Engine, Cloud SQL, Pub/Sub). - Container – Docker ≥ 20, Kubernetes ≥ 1.24, Helm ≥ 3.5.- IaC – Terraform ≥ 1.0, Ansible ≥ 2.9. - CI/CD – GitHub Actions, Jenkins, CircleCI (for legacy pipelines). - Monitoring – Prometheus, Grafana, Alertmanager, Datadog (for some legacy services). - Logging – Splunk, Elasticsearch‑Kibana stack, Loki. - Incident response – PagerDuty, Opsgenie (we’re migrating fully to PagerDuty). - Version control – GitHub (private repos, branch protection rules). - Collaboration – Slack (primary chat), Confluence (knowledge base), JIRA (ticketing). On‑call rhythm & expectationsOur on‑call schedule is a 7‑day rotation with a 48‑hour backup window.Each engineer handles roughly ≈ 350 alerts per month, averaging ≈ 2 incidents per week. We have a “no‑call‑out‑of‑hours” policy for holidays: the next engineer in the rotation covers the entire period, and the team shares the load. During an incident you’ll have a clear run‑book, but we also encourage “play‑by‑play” documentation in realtime to help the rest of the crew follow along. Compensation & benefits (the numbers, no fluff) - Base salary: $130,000 – $170,000 USD, depending on experience and location.- Annual bonus: Up to 15 % of base, tied to reliability KPIs (SLO compliance, MTTR improvement). - Equity: 0.05 % – 0.15 % RSU pool, vested over four years with a one‑year cliff. - Health: Full medical, dental, vision for employee + 1 dependents, including telehealth. - Retirement: 401(k) match up to 5 % of salary. - Time off: 20 days PTO + 10 company holidays, plus a “recovery week” after each major incident. - Learning budget: $2,500 per year for courses, conferences, or certifications (we’ll reimburse even for remote‑only events).- Equipment: Choice of MacBook Pro or Linux workstation, dual‑monitor setup shipped to your home office, and a $150 monthly stipend for internet. Why you’ll love working with us - Impact‑first:Your work directly influences the experience of thousands of end‑users; a single reliability improvement can translate to millions of dollars saved for a client. - Autonomy: We give you ownership of the reliability roadmap—you decide where to invest engineering effort, not a product manager. - Culture of candor: Post‑mortems are blameless, data‑driven chronicles that we read aloud in our weekly “Reliability Round‑up”.Everyone’s voice is heard, from junior engineers to the CTO. - Remote‑first mindset: While we are legally anchored in Escondido, California, you can work from anywhere in the United States. Our “remote‑first” policy means we never require you to be in a physical office, except for the optional quarterly meetup in Escondido, California. A human moment > “I still remember the night we were down for 12 minutes because a misconfigured Prometheus scrape target blew up the entire cluster. We all gathered on a conference call, one teammate in his kitchen, another on a balcony in Escondido, California.Within 5 minutes we had a rollback plan, and by the time the sunrise hit the roof of the office building in Escondido, California, the service was back up. It reminded me why we do this work—every alert is a chance to protect a real user’s workflow.” – * Alex,Senior SRE Lead* Application process 1. Resume & cover letter: Send us a brief note (no longer than one page) explaining a reliability challenge you solved and why you’re drawn to remote work anchored in Escondido, California. 2. Screening call (30 min): With the hiring manager to discuss your background, expectations, and the role’s day‑to‑day.3. Technical deep‑dive (1 hr): Live problem‑solving session covering incident response, Terraform debugging, and a short coding exercise in Python. 4. Team interview (45 min): Meet two senior SREs for a cultural fit conversation and a walk‑through of a recent post‑mortem. 5. Final interview (30 min): With the VP of Engineering to discuss career growth, leadership philosophy, and remote‑first policies. If you pass all steps, we’ll extend an offer within a week and kick off the onboarding process—including a “welcome kit” shipped to your home, a first‑day meeting with your mentor, and a 2‑week “shadow” period where you sit on the on‑call rotation with a senior partner.Closing note We are not looking for a résumé‑checker; we want someone who feels a genuine pull toward making complex systems resilient, who enjoys digging into Prometheus queries at 2 am, and who values transparent communication as much as technical depth. If you see yourself improving our MTTR, lowering alert fatigue, and shaping a reliability culture that scales as fast as our product, we’d love to hear from you. * Ready to join a team that treats reliability as a craft, not a checkbox? and let’s build something that stays up when the world needs it most.* Apply tot his job
Apply Now

Similar Jobs

Shopify Developer Needed TODAY — Fix Product Upsell Logic ($50)

Remote, USA Full-time

Shopify Theme Developer Needed to Complete Shopify Store Setup (Multi-Store Linking + Theme Fixes)

Remote, USA Full-time

Site Reliability Engineer 2 DevOps REMOTE (ship required)

Remote, USA Full-time

Senior Solution Architect, ServiceNow Platform

Remote, USA Full-time

Principal Customer Success Executive Telco and Media

Remote, USA Full-time

ServiceNow Vulnerability Response (VR) Developer - Remote

Remote, USA Full-time

Senior ServiceNow Developer (Remote) in Reston, VA

Remote, USA Full-time

Sr. ServiceNow Developer/Admin

Remote, USA Full-time

ServiceNow Developer (Customer Service Management) :: Nashville, TN (REMOTE)

Remote, USA Full-time

Lead ServiceNow Solution Architect

Remote, USA Full-time

Brand Influencer & Partnerships Manager

Remote, USA Full-time

Remote Business Operations Specialist – $20/hr

Remote, USA Full-time

Dental Lab Business Development

Remote, USA Full-time

Optometrist Job in Charleston, SC

Remote, USA Full-time

Human Resources / People Leader (General Interest)

Remote, USA Full-time

[Remote] AI Software Engineer (Typescript & LLM)

Remote, USA Full-time

Senior Automation Engineer, bolthires Pharmacy RME

Remote, USA Full-time

Experienced Nonprofit Leader and Executive Director – Part-Time Opportunity for a Dynamic and Relational Professional to Drive Scholarship Programs and Fundraising Growth

Remote, USA Full-time

Supply Chain Analyst III

Remote, USA Full-time

Medical Editor and Fact Checker

Remote, USA Full-time
Back to Home