Remote Site Reliability Engineer jobs.
“Remote”? Depends where. $$$? When they say. Stack? Before the essay.
Where can you apply from?
76 openings. None older than 30 days.
Newest first
Senior Site Reliability Engineer (SRE)
3 days agoPython · Java · AWS
Remote within the United States; must be legally authorized to work there without current or future employer-sponsored visa. Required: 6–10+ years in SRE, infrastructure or backend systems engineering; ownership of reliability for complex distributed production systems; cloud infrastructure (AWS, GCP or Azure); observability, incident management and performance; programming for automation (Go, Python or Java); incident leadership and cross-team influence. Preferred: building SLO/on-call practices, Kubernetes, infrastructure as code such as Terraform, performance engineering. Incident response and sustainable on-call workload are core responsibilities.
DevOps & SRE Engineer
3 days agoPython · Java · AWS
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Prometheus, Grafana, CI/CD pipelines, SLOs, Chaos engineering, AWS, Azure, GCP
Telemetry Engineer
3 days agoPython · Java
6 years of experience, Prometheus, Grafana, OpenTelemetry, Datadog, New Relic, Splunk, Go, Python, Java, SRE principles, distributed tracing, structured logging, Linux, containers, CI/CD, observability cost optimization, eBPF-based observability
Platform Reliability Engineer
3 days agoPython · Java · Kubernetes
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Observability tooling, CI/CD pipelines, Distributed system design, SLOs and error budgets, Chaos engineering, Cloud platforms
Site Observability Engineer
3 days ago6 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, Distributed tracing, High-cardinality metrics, SLOs, CI/CD integration, Linux internals, Container platforms
Senior Platform Engineer I
4 days agoKubernetes
5+ years platform/SRE, production Kubernetes, Terraform/Crossplane/Pulumi, on-call incident response
Systems Observability Specialist
4 days ago5+ years SRE/observability, Prometheus + Grafana, 1 of Datadog/New Relic/Splunk, OpenTelemetry, no new H-1B
DataDog Observability Engineer - Part Time (R-00231)
5 days agoUS citizenship, 5+ years observability/SRE, Datadog, Terraform, 3+ listed certifications
Observability Engineer
5 days ago12 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, Distributed tracing, High-cardinality metrics, SLOs, CI/CD integration, Linux internals, eBPF-based observability
Site Reliability Engineer (SRE)
5 days agoPython · AWS · Azure
10 years of experience, Kubernetes, Python, Go, Prometheus, Grafana, CI/CD pipelines, Chaos engineering, Distributed systems, SLOs and error budgets, Cloud platforms (AWS, Azure, GCP)
Senior Software Engineer, DGX Cloud Production Engineering
5 days agoPython · Go · Kubernetes
8+ years production infrastructure, Kubernetes, Python/Golang or similar, on-call, GPU clusters
Cloud Infrastructure Engineer – AWS
6 days agoPython · AWS · Kubernetes
Bachelor’s or Master’s degree in CS, IT, Engineering or related field; 10+ years in IT/cloud engineering including 5+ years designing and operating enterprise AWS; EC2/VPC/IAM/S3/RDS/Lambda and other AWS services; Terraform, AWS CDK or CloudFormation; production EKS, ECS or Kubernetes; CI/CD/DevOps/GitOps; Python and Bash (Go/PowerShell preferred); IAM, encryption, compliance and observability. Preferred: advanced AWS certifications and other qualifications in posting.
Cloud Infrastructure Network Engineer
6 days agoKubernetes
Bachelor’s degree in CS, Networking or related field; 5+ years networking with substantial cloud networking; at least one major cloud provider; routing, switching and BGP; Direct Connect, ExpressRoute or equivalent hybrid connectivity; Terraform for cloud networking; security controls; Kubernetes networking/service mesh fundamentals; packet-level troubleshooting. Preferred: cloud networking certification, multi-cloud, SD-WAN/SASE, eBPF, regulated environments.
Cloud Networking Engineer
6 days agoPython · Kubernetes
Bachelor’s degree in Computer Science or related field; 5+ years in platform engineering, SRE or networking; production Istio or Linkerd; Envoy, Kubernetes networking/CNI/ingress, mTLS/PKI and certificate lifecycle, distributed tracing; Go or Python; networking troubleshooting. Preferred: multi-cluster mesh, Cilium/eBPF, SPIFFE/SPIRE, open-source contributions, enterprise zero-trust.
Container Platform Engineer
6 days agoPython · Kubernetes
Bachelor’s degree in CS, Engineering or related field; 5+ years operating production container platforms including 3+ years on Red Hat OpenShift; Kubernetes/OpenShift internals, Linux, Ansible/Terraform/Helm, Tekton/Jenkins/Argo CD, Bash/Python/Go, cluster observability and image security. Preferred: Red Hat certification, public cloud, service mesh, regulated environments and GitOps.
Monitoring Engineer
6 days agoPython · Java
Bachelor’s degree in Computer Science or related field; 5+ years in SRE, platform engineering or observability; Prometheus, Grafana and one commercial platform such as Datadog/New Relic/Splunk; OpenTelemetry, tracing, structured logging; Go, Python or Java; high-throughput metrics/log pipelines; SLOs/error budgets; CI/CD and incident management; Linux, networking and containers. Preferred: Thanos/Mimir/Cortex/Loki/Tempo, eBPF, observability cost optimization.
Service Mesh Architect
6 days agoPython · Kubernetes
Bachelor’s degree in Computer Science or related field; 5+ years in platform engineering, SRE or networking; production Istio or Linkerd; Envoy, Kubernetes networking/CNI/ingress, mTLS/PKI/certificates, distributed tracing; Go or Python; networking and control-plane troubleshooting. Preferred: multi-cluster mesh, Cilium/eBPF, open-source contributions, SPIFFE/SPIRE, enterprise zero-trust.
Reliability Engineer
6 days agoPython · Java · AWS
Bachelor’s degree in Computer Science, Engineering or related field; 5+ years of SRE, DevOps or production engineering for large distributed systems; Python, Go or Java; Linux at scale; production Kubernetes; observability; CI/CD; distributed system design; incident response and post-incident reviews. Preferred: SLOs/error budgets, chaos engineering, AWS/Azure/GCP, capacity planning and service mesh.
Senior/Lead Site Reliability Engineer
7 days agoPython · Java · Go
Linux, TCP/IP, HTTP, C/C++, Python, Golang, Rust, Java, Open-source software, AIOps, AI operations, Native Chinese proficiency
Cloud Infrastructure Engineer (Open LMS) Colombia, Remote
7 days agoPython · AWS
AWS, Terraform, Puppet or similar, Python daemons, Linux, distributed systems, on-call
Service Mesh Engineer
7 days agoPython · Kubernetes
Headline says 7+ years; detailed required qualifications specify 5+ years of platform engineering, SRE or networking, production Istio or Linkerd, Envoy, Kubernetes CNI, mTLS/PKI, distributed tracing and Go or Python. Cilium, SPIFFE/SPIRE and zero-trust networking are preferred.
OpenShift Platform Engineer
7 days agoPython · Kubernetes
Headline says 8+ years; detailed required qualifications specify 5+ years operating production container platforms including at least 3 years on Red Hat OpenShift, bachelor's degree, Kubernetes/OpenShift internals, Linux, Ansible or Terraform or Helm, Tekton or Jenkins or Argo CD, Bash or Python or Go, observability and image security. Service mesh and GitOps are preferred.
DevOps & SRE Engineer
10 days agoPython · Java · AWS
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Prometheus, Grafana, CI/CD pipelines, SLOs, Chaos engineering, AWS, Azure, GCP
Telemetry Engineer
10 days agoPython · Java
6 years of experience, Prometheus, Grafana, OpenTelemetry, Datadog, New Relic, Splunk, Go, Python, Java, SRE principles, distributed tracing, structured logging, Linux, containers, CI/CD, observability cost optimization, eBPF-based observability
Senior Infrastructure Engineer | Permanent WFH | Night Shift
10 days ago5 years of experience, Windows Server Administration, Active Directory, Microsoft 365 Administration, VMware vSphere, Backup and recovery solutions, Networking technologies, PowerShell, Cloud platforms
Kubernetes Service Engineer
10 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, mTLS, Kubernetes, Go, Python, Distributed tracing, Service mesh architecture, Traffic management policies, Zero-trust networking
Senior Data Ops Engineer, Data Activation & Products - Activision
10 days agoKubernetes
5 years of experience, Kubernetes, Databricks, Spark, Airflow, Kafka, Cloud infrastructure, Incident management, Automation frameworks, Observability tools, Event-driven architecture
Site Reliability Engineer (India)
11 days agoPython · Go · AWS
4–7 years SRE/DevOps, 3+ years production Kubernetes/cloud, AWS and GCP, Golang or Python, 12:30–9:30 IST
Kubernetes & OpenShift Engineer
11 days agoPython · AWS · Azure
6 years of experience, Kubernetes internals, OpenShift internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Cluster monitoring and logging, Container image security, Red Hat OpenShift experience, Public cloud experience (AWS, Azure, GCP)
Platform Networking Engineer
11 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, Service mesh production experience