Remote Site Reliability Engineer jobs.
“Remote”? Depends where. $$$? When they say. Stack? Before the essay.
Where can you apply from?
106 openings. None older than 30 days.
Newest first
Sr. Reliability Engineer
about 8 hours agoAWS · Kubernetes
5 years of experience, Datadog, AWS, Kubernetes/EKS, Terraform, PostgreSQL, SLO engineering, AI coding tools, Incident management, Load-testing, Healthcare experience
Staff Cloud Identity Platform Engineer
about 11 hours agoAzure
5 years of experience, Microsoft Entra ID, OAuth 2.0, OpenID Connect, SAML, Identity lifecycle management, Service-level indicators, Error budgets, Azure infrastructure, Identity security principles, Automation with PowerShell, Infrastructure as code, SRE concepts, Identity governance, Incident management, Capacity planning, Observability practices
Senior Staff Cloud Platform Engineer
1 day agoGo · Kubernetes
Multi-team infrastructure strategy, Kubernetes at production scale, Golang or similar, Terraform, SLOs
Kubernetes Engineer (Mid–Senior) | DoD Platform Engineering
2 days agoKubernetes
Able to obtain DoD Secret clearance, 3-5+ years Kubernetes platforms, Helm, Terraform, GitOps, Istio a plus
Kubernetes & OpenShift Engineer
2 days agoPython · AWS · Azure
6 years of experience, Kubernetes internals, OpenShift internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Cluster monitoring and logging, Container image security, Red Hat OpenShift experience, Public cloud experience (AWS, Azure, GCP)
Observability Engineer
2 days ago12 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, Distributed tracing, High-cardinality metrics, SLOs, CI/CD integration, Linux internals, Container platforms
Reliability Monitoring Engineer
2 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, distributed tracing, structured logging, Go, Python, Java, high-cardinality metrics, SLOs, error budgets, SRE principles, CI/CD integration, Linux internals, networking, container platforms, Thanos, Mimir, Cortex, Loki, Tempo, eBPF, audit-grade logging
Site Reliability Engineer (SRE)
2 days agoPython · AWS · Azure
10 years of experience, Kubernetes, Prometheus, Grafana, Python, Go, CI/CD pipelines, Chaos engineering, Distributed systems, SLOs and error budgets, Cloud platforms (AWS, Azure, GCP)
Systems Reliability Engineer
2 days agoPython · Kubernetes
6 years of experience, Kubernetes, Python, Go, Prometheus, Grafana, CI/CD pipelines, Chaos engineering, SLOs and error budgets, Linux at scale, Distributed systems design
Azure Infrastructure Engineer
2 days agoPython · Azure · Kubernetes
6 years of experience, Azure core services, Infrastructure-as-code (Terraform, Bicep, ARM), Azure Kubernetes Service (AKS), Azure DevOps or GitHub Actions, PowerShell, Bash, Python, Cloud security principles, Monitoring and observability strategies, Hybrid cloud or multi-cloud experience
Kafka Engineer
3 days agoKubernetes
7 years of experience, Kafka internals, Kafka security, Kafka Connect, Infrastructure-as-code, Observability tooling, DevOps practices, Kubernetes (Strimzi, Confluent Operator)
Monitoring Engineer
3 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, distributed tracing, structured logging, Go, Python, Java, high-cardinality metrics, SLOs, error budgets, SRE principles, CI/CD integration, Linux internals, networking, container platforms
Observability Engineer
3 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, distributed tracing, structured logging, Go, Python, Java, high-cardinality metrics, SLOs, error budgets, SRE principles, CI/CD integration, Linux internals, networking, container platforms, Thanos, Mimir, Cortex, Loki, Tempo, eBPF, cost optimization, regulated environments
OpenShift Platform Engineer
3 days agoPython · Kubernetes
8 years of experience, OpenShift, Kubernetes internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Monitoring and logging (Prometheus, Grafana, EFK, Tempo), Container image security, Multi-tenant platform design, Disaster recovery strategies
Platform Infrastructure Engineer
3 days agoPython · Kubernetes
6 years of experience, OpenShift, Kubernetes, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Bash, Python, Go, Cluster monitoring tools, Container image security, Service mesh (Istio, Linkerd)
Reliability Engineer
3 days agoPython · Java · AWS
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Prometheus, Grafana, CI/CD pipelines, Distributed system design, Incident response, SLOs and error budgets, Chaos engineering, AWS, Azure, GCP
Site Reliability Engineer Technical Lead
3 days agoPython · AWS · Azure
8 years of experience, Deep SRE and Systems Expertise, AWS, GCP, Azure, Kubernetes, Python, Automation and Tooling, Observability and Analysis, Dynatrace, Splunk, ELK Stack, Leadership, Communication, Interpersonal Skills
Senior IT Infrastructure & Network Engineer - Remote (Latam)
3 days agoPython
Microsoft 365 administration, Entra ID management, Windows Server administration, Linux server experience, Deep knowledge of TCP/IP, Packet capture and traffic analysis, Experience with IoT devices, Infrastructure security knowledge, Advanced English proficiency, Scripting in PowerShell, Bash, or Python
Cloud Networking Engineer
4 days agoKubernetes
6+ years platform/SRE/networking, Istio or Linkerd, Envoy, Kubernetes, no new H-1B sponsorship
Senior Cloud Platform Security Engineer
4 days agoAzure
7+ years cloud infrastructure or security, production Azure, Entra ID, Terraform or Bicep, Eastern Time hours
Cloud Infrastructure Engineer – AWS
4 days agoAWS
10+ years IT/cloud engineering incl. 5+ years AWS, Terraform/CDK/CloudFormation, EKS/ECS, no new H-1B
Service Infrastructure Engineer
4 days agoPython · Go · Kubernetes
6+ years platform/SRE/networking, Istio or Linkerd in production, Envoy, Kubernetes, mTLS/PKI, Golang or Python
Service Mesh Architect
4 days agoPython · Go · Kubernetes
6+ years platform/SRE/networking, Istio or Linkerd in production, Envoy, Kubernetes, mTLS/PKI, Golang or Python
Systems Reliability Engineer
4 days agoPython · Java · Go
6+ years SRE or DevOps, Python, Golang or Java, Kubernetes, Linux at scale, no new H-1B sponsorship
Staff Cloud Operations Engineer
7 days agoPython · Go · GCP
5+ years cloud infrastructure (GCP), Kubernetes, observability, Python/Bash/Golang, 24/7 on-call, US work authorization
Platform Reliability Engineer
7 days agoPython · Java · Go
6+ years SRE/DevOps, Python/Golang/Java, Linux at scale, Kubernetes, Prometheus/Grafana/OpenTelemetry
Site Observability Engineer
7 days agoPython · Java · Go
6+ years SRE/platform/observability, Prometheus/Grafana, Datadog/New Relic/Splunk, OpenTelemetry, Golang/Python/Java
Staff Software Engineer (L4)
8 days ago8 years of experience, Reliability engineering, SLIs and SLOs, Production systems accountability, Incident command, Post-mortem analysis, Capacity planning, Observability, Distributed systems in cloud, Infrastructure-as-code, Container orchestration, Chaos engineering
Principal Software Engineer, DevOps
8 days agoPython · Go · AWS
12+ years SRE/DevOps, AWS & Terraform, Kubernetes, Golang or Python, agentic AI in production, on-call
Lead Observability Engineer
8 days agoAWS
8+ years SRE/DevOps/observability, Prometheus with Thanos/Cortex/Mimir, OpenTelemetry, AWS observability, Terraform