Remote Site Reliability Engineer jobs.
“Remote”? Depends where. $$$? When they say. Stack? Before the essay.
Where can you apply from?
89 openings. None older than 30 days.
Newest first
Systems Engineer I
1 day ago2 years of experience, A+ certification, OEM equipment certifications, Microsoft certifications, Familiarity with ASTEA, User account management, Server monitoring, Connectivity troubleshooting, Preventative maintenance, Customer relationship building, Technical documentation skills
Systems Engineer I
1 day ago2 years of experience, A+ certification, OEM equipment certifications, Microsoft certifications, Familiarity with ASTEA, User account management, Server monitoring, Connectivity troubleshooting, Preventative maintenance, Customer-first communication, Technical documentation skills
Staff Cloud Operations Engineer
1 day agoPython · GCP · Kubernetes
5 years of experience, GCP, Cloud infrastructure engineering, Observability tools, Python, Docker, Kubernetes, AI tools usage, 24/7 on-call rotation
Staff Software Engineer (L4)
2 days ago8 years of experience, Reliability engineering, SLIs and SLOs, Production systems accountability, Incident command, Post-mortem analysis, Capacity planning, Observability, Distributed systems in cloud, Infrastructure-as-code, Container orchestration, Chaos engineering
Principal Software Engineer, DevOps
2 days agoPython · Go · AWS
12+ years SRE/DevOps, AWS & Terraform, Kubernetes, Golang or Python, agentic AI in production, on-call
Lead Observability Engineer
2 days agoAWS
8+ years SRE/DevOps/observability, Prometheus with Thanos/Cortex/Mimir, OpenTelemetry, AWS observability, Terraform
Red Hat OpenShift Engineer
2 days agoAWS · Azure · GCP
6 years of experience, OpenShift, Kubernetes, Linux, Ansible, Terraform, Helm, CI/CD, Tekton, Jenkins, GitLab CI, Argo CD, Service Mesh (Istio/Linkerd), Cloud experience (AWS, Azure, GCP)
Systems Observability Specialist
2 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, distributed tracing, structured logging, Go, Python, Java, high-cardinality metrics, SLOs, error budgets, SRE principles, CI/CD integration, Linux internals, networking, container platforms
Senior Site Reliability Engineer
3 days agoAWS · Azure · Kubernetes
8 years of experience, AWS, Azure, Kubernetes (EKS/AKS), Terraform, Cloud-native architecture, Payment industry standards, Security best practices for containers, Infrastructure as Code
Observability Engineer
3 days ago12 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, Distributed tracing, High-cardinality metrics, SLOs, CI/CD integration, Linux internals, eBPF-based observability
Site Reliability Engineer (SRE)
3 days agoPython · Kubernetes
10 years of experience, Kubernetes, Python, Go, Prometheus, Grafana, CI/CD pipelines, Chaos engineering, Distributed systems, SLOs and error budgets, Cloud platforms
Site Reliability Engineer Technical Lead
4 days agoPython · Kubernetes
Listed onsite in New Albany, NY, SRE, Kubernetes, Python automation, Dynatrace or Splunk, no new H-1B sponsorship
Senior Site Reliability Engineer - India
4 days agoPython · AWS · GCP
8 years of experience, AWS/GCP, Kubernetes (EKS), Terraform, Python, Go, FinOps, Disaster Recovery, SLIs/SLOs, Observability, GitOps (Argo CD)
Site Reliability Engineer - India
4 days agoPython · AWS · GCP
5 years of experience, Python, Go, Kubernetes, Terraform, AWS, GCP, Datadog, FinOps, Disaster Recovery, Observability, Incident Management
Monitoring Engineer
4 days agoPython · Java · Go
5+ years SRE/observability (6+ overall), Prometheus, Grafana, OpenTelemetry, Golang/Python/Java
OpenShift Platform Engineer
4 days agoKubernetes
W-2, no new H-1B, 8+ years container platforms incl. 3+ OpenShift, Kubernetes internals, Linux, Tekton/Argo CD
Reliability Engineer
4 days agoPython · Java · Go
W-2, no new H-1B, 6+ years (5+ SRE/DevOps), Python/Golang/Java, Linux at scale, Kubernetes, Prometheus/Grafana
Apache Kafka Developer
4 days agoPython
6 years of experience, Kafka internals, Kafka security, Kafka Connect, Schema Registry, Kafka Streams, HA/DR strategies, Python scripting, Infrastructure-as-code, Observability tooling, Confluent Certified Administrator
Service Mesh Engineer
5 days agoPython · Kubernetes
7 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, Cilium, SPIFFE/SPIRE, Zero-trust networking
Cloud Infrastructure Network Engineer
5 days agoKubernetes
6 years of experience, Cloud networking, VPC/VNet design, Hybrid connectivity, Infrastructure-as-code (Terraform), Cloud security and network controls, Kubernetes networking, Multi-cloud networking, SD-WAN familiarity, eBPF-based networking tools, Regulated industries exposure
Kafka Engineer
5 days agoKubernetes
7 years of experience, Kafka internals, Kafka security, Kafka Connect, Infrastructure-as-code, Observability tooling, DevOps practices, Kubernetes experience, Streaming data governance
Observability Engineer
5 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, distributed tracing, structured logging, Go, Python, Java, high-cardinality metrics, SLOs, error budgets, SRE principles, CI/CD integration, Linux internals, networking, container platforms, Thanos, Mimir, Cortex, Loki, Tempo, eBPF, cost optimization, regulated environments
OpenShift Administrator
5 days agoPython · Kubernetes
6 years of experience, OpenShift, Kubernetes internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Container image security, Cluster monitoring tools, Service mesh (OpenShift Service Mesh, Istio, Linkerd)
Platform Infrastructure Engineer (SRE Core)
7 days agoPython · AWS · GCP
Kubernetes expertise, Terraform proficiency, GCP and AWS experience, Python, Bash, Go, Grafana Cloud, Prometheus/Mimir, LLM-based code-assist tools, Cilium, GitOps methodologies
Senior Site Reliability Engineer (SRE)
8 days agoPython · Java · AWS
Remote within the United States; must be legally authorized to work there without current or future employer-sponsored visa. Required: 6–10+ years in SRE, infrastructure or backend systems engineering; ownership of reliability for complex distributed production systems; cloud infrastructure (AWS, GCP or Azure); observability, incident management and performance; programming for automation (Go, Python or Java); incident leadership and cross-team influence. Preferred: building SLO/on-call practices, Kubernetes, infrastructure as code such as Terraform, performance engineering. Incident response and sustainable on-call workload are core responsibilities.
Windows System Administrator
8 days agoPython · Azure
Azure, Windows Server & Active Directory, PowerShell or Python, Datadog, advanced English, 6am-3pm shift
DevOps & SRE Engineer
8 days agoKubernetes
More than 5 years SRE/DevOps, Linux and Kubernetes at scale, observability, CI/CD automation
Telemetry Engineer
8 days ago6+ years overall, 5+ SRE/observability, Prometheus/Grafana/OpenTelemetry, Datadog/New Relic/Splunk
Platform Reliability Engineer
8 days agoPython · Java · Go
6+ years overall, 5+ SRE/DevOps, Linux/Kubernetes, Python or Golang or Java, distributed systems
Site Observability Engineer
8 days agoPython · Java · Go
6+ years SRE/observability, Prometheus/Grafana, OpenTelemetry, Datadog, Golang/Python/Java