Remote Site Reliability Engineer jobs.
“Remote”? Depends where. $$$? When they say. Stack? Before the essay.
Where can you apply from?
84 openings. None older than 30 days.
Newest first
Staff Software Engineer (L4)
about 8 hours ago8 years of experience, Reliability engineering, SLIs and SLOs, Production systems accountability, Incident command, Post-mortem analysis, Capacity planning, Observability, Distributed systems in cloud, Infrastructure-as-code, Container orchestration, Chaos engineering
Principal Software Engineer, DevOps
about 8 hours agoGo · AWS · Kubernetes
12 years of experience, AWS, Terraform, Kubernetes, Golang, AI systems, Observability tools, Cloud infrastructure management, Production alerts management
Monitoring Engineer
2 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, distributed tracing, structured logging, Go, Python, Java, high-cardinality metrics, SLOs, error budgets, SRE principles, CI/CD integration, Linux internals, networking, container platforms
OpenShift Platform Engineer
2 days agoPython · Kubernetes
8 years of experience, OpenShift, Kubernetes internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Monitoring and logging tools, Container image security, Disaster recovery strategies, Multi-tenant OpenShift platforms, GitOps workflows (Argo CD, Flux)
Reliability Engineer
2 days agoPython · Java · AWS
6 years of experience, Python, Go, Java, Kubernetes, Linux, Prometheus, Grafana, CI/CD pipelines, SLOs, Chaos engineering, AWS, Azure, GCP
Site Reliability Engineer Technical Lead
2 days ago8 years of experience, Deep SRE and Systems Expertise, Automation and Tooling, Observability and Analysis, Exceptional leadership and communication skills
Apache Kafka Developer
2 days agoPython
6 years of experience, Kafka internals, Kafka security, Kafka Connect, Schema Registry, Kafka Streams, HA/DR strategies, Python scripting, Infrastructure-as-code, Observability tooling, Confluent Certified Administrator
Service Mesh Engineer
3 days agoPython · Kubernetes
7 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, Cilium, SPIFFE/SPIRE, Zero-trust networking
Cloud Infrastructure Network Engineer
3 days agoKubernetes
6 years of experience, Cloud networking, VPC/VNet design, Hybrid connectivity, Infrastructure-as-code (Terraform), Cloud security and network controls, Kubernetes networking, Multi-cloud networking, SD-WAN familiarity, eBPF-based networking tools, Regulated industries exposure
Kafka Engineer
3 days agoKubernetes
7 years of experience, Kafka internals, Kafka security, Kafka Connect, Infrastructure-as-code, Observability tooling, DevOps practices, Kubernetes experience, Streaming data governance
Observability Engineer
3 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, distributed tracing, structured logging, Go, Python, Java, high-cardinality metrics, SLOs, error budgets, SRE principles, CI/CD integration, Linux internals, networking, container platforms, Thanos, Mimir, Cortex, Loki, Tempo, eBPF, cost optimization, regulated environments
OpenShift Administrator
3 days agoPython · Kubernetes
6 years of experience, OpenShift, Kubernetes internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Container image security, Cluster monitoring tools, Service mesh (OpenShift Service Mesh, Istio, Linkerd)
Senior Site Reliability Engineer (SRE)
6 days agoPython · Java · AWS
Remote within the United States; must be legally authorized to work there without current or future employer-sponsored visa. Required: 6–10+ years in SRE, infrastructure or backend systems engineering; ownership of reliability for complex distributed production systems; cloud infrastructure (AWS, GCP or Azure); observability, incident management and performance; programming for automation (Go, Python or Java); incident leadership and cross-team influence. Preferred: building SLO/on-call practices, Kubernetes, infrastructure as code such as Terraform, performance engineering. Incident response and sustainable on-call workload are core responsibilities.
Windows System Administrator
6 days agoPython · Java · Azure
Microsoft Azure, Windows Server, Active Directory, PowerShell, Python, Datadog, Web server management, Rotational support experience, Java microservices (WildFly)
DevOps & SRE Engineer
6 days agoKubernetes
More than 5 years SRE/DevOps, Linux and Kubernetes at scale, observability, CI/CD automation
Telemetry Engineer
6 days ago6+ years overall, 5+ SRE/observability, Prometheus/Grafana/OpenTelemetry, Datadog/New Relic/Splunk
Platform Reliability Engineer
6 days agoPython · Java · Go
6+ years overall, 5+ SRE/DevOps, Linux/Kubernetes, Python or Golang or Java, distributed systems
Site Observability Engineer
6 days agoPython · Java · Go
6+ years SRE/observability, Prometheus/Grafana, OpenTelemetry, Datadog, Golang/Python/Java
Senior Platform Engineer I
7 days agoKubernetes
5+ years platform/SRE, production Kubernetes, Terraform/Crossplane/Pulumi, on-call incident response
Systems Observability Specialist
7 days ago5+ years SRE/observability, Prometheus + Grafana, 1 of Datadog/New Relic/Splunk, OpenTelemetry, no new H-1B
DataDog Observability Engineer - Part Time (R-00231)
8 days agoUS citizenship, 5+ years observability/SRE, Datadog, Terraform, 3+ listed certifications
Observability Engineer
8 days ago12 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, Distributed tracing, High-cardinality metrics, SLOs, CI/CD integration, Linux internals, eBPF-based observability
Site Reliability Engineer (SRE)
8 days agoPython · AWS · Azure
10 years of experience, Kubernetes, Python, Go, Prometheus, Grafana, CI/CD pipelines, Chaos engineering, Distributed systems, SLOs and error budgets, Cloud platforms (AWS, Azure, GCP)
MSP Engineer, Network and Security - Rotating shifts
8 days agoPython
3 years of experience, Cloudflare, WAF, DDoS mitigation, Terraform, Python, Bash, DNS, HTTP/HTTPS, French, English, Managed services experience
Cloud Infrastructure Engineer – AWS
9 days agoPython · AWS · Kubernetes
Bachelor’s or Master’s degree in CS, IT, Engineering or related field; 10+ years in IT/cloud engineering including 5+ years designing and operating enterprise AWS; EC2/VPC/IAM/S3/RDS/Lambda and other AWS services; Terraform, AWS CDK or CloudFormation; production EKS, ECS or Kubernetes; CI/CD/DevOps/GitOps; Python and Bash (Go/PowerShell preferred); IAM, encryption, compliance and observability. Preferred: advanced AWS certifications and other qualifications in posting.
Cloud Infrastructure Network Engineer
9 days agoKubernetes
Bachelor’s degree in CS, Networking or related field; 5+ years networking with substantial cloud networking; at least one major cloud provider; routing, switching and BGP; Direct Connect, ExpressRoute or equivalent hybrid connectivity; Terraform for cloud networking; security controls; Kubernetes networking/service mesh fundamentals; packet-level troubleshooting. Preferred: cloud networking certification, multi-cloud, SD-WAN/SASE, eBPF, regulated environments.
Cloud Networking Engineer
9 days agoPython · Kubernetes
Bachelor’s degree in Computer Science or related field; 5+ years in platform engineering, SRE or networking; production Istio or Linkerd; Envoy, Kubernetes networking/CNI/ingress, mTLS/PKI and certificate lifecycle, distributed tracing; Go or Python; networking troubleshooting. Preferred: multi-cluster mesh, Cilium/eBPF, SPIFFE/SPIRE, open-source contributions, enterprise zero-trust.
Container Platform Engineer
9 days agoPython · Kubernetes
Bachelor’s degree in CS, Engineering or related field; 5+ years operating production container platforms including 3+ years on Red Hat OpenShift; Kubernetes/OpenShift internals, Linux, Ansible/Terraform/Helm, Tekton/Jenkins/Argo CD, Bash/Python/Go, cluster observability and image security. Preferred: Red Hat certification, public cloud, service mesh, regulated environments and GitOps.
Monitoring Engineer
9 days agoPython · Java
Bachelor’s degree in Computer Science or related field; 5+ years in SRE, platform engineering or observability; Prometheus, Grafana and one commercial platform such as Datadog/New Relic/Splunk; OpenTelemetry, tracing, structured logging; Go, Python or Java; high-throughput metrics/log pipelines; SLOs/error budgets; CI/CD and incident management; Linux, networking and containers. Preferred: Thanos/Mimir/Cortex/Loki/Tempo, eBPF, observability cost optimization.
Service Mesh Architect
9 days agoPython · Kubernetes
Bachelor’s degree in Computer Science or related field; 5+ years in platform engineering, SRE or networking; production Istio or Linkerd; Envoy, Kubernetes networking/CNI/ingress, mTLS/PKI/certificates, distributed tracing; Go or Python; networking and control-plane troubleshooting. Preferred: multi-cluster mesh, Cilium/eBPF, open-source contributions, SPIFFE/SPIRE, enterprise zero-trust.