Remote Site Reliability Engineer jobs.
“Remote”? Depends where. $$$? When they say. Stack? Before the essay.
Where can you apply from?
80 openings. None older than 30 days.
Newest first
Senior Infrastructure Solutions Engineer
about 9 hours agoGCP · Kubernetes
6 years of experience, Kubernetes orchestration, Infrastructure as Code (Terraform, Helm, ArgoCD), Google Cloud Platform (GCP), Production infrastructure maintenance, Technical enablement, Developer experience, Solutions architecture, Large-scale block storage, Driving large-scale migrations, Consulting-adjacent engineering role, Workflow orchestration, Defining platform adoption metrics
Senior Infrastructure Solutions Engineer
about 9 hours agoPython · GCP · Kubernetes
6 years of experience, Kubernetes, Google Cloud Platform (GCP), Terraform, Helm, Go, Python, Technical enablement, Developer experience, Solutions architecture, Large-scale block storage, Workflow orchestration, Platform adoption measurement
Senior Site Reliability Engineer (SRE)
1 day agoPython · Java · AWS
Remote within the United States; must be legally authorized to work there without current or future employer-sponsored visa. Required: 6–10+ years in SRE, infrastructure or backend systems engineering; ownership of reliability for complex distributed production systems; cloud infrastructure (AWS, GCP or Azure); observability, incident management and performance; programming for automation (Go, Python or Java); incident leadership and cross-team influence. Preferred: building SLO/on-call practices, Kubernetes, infrastructure as code such as Terraform, performance engineering. Incident response and sustainable on-call workload are core responsibilities.
Senior Platform Engineer I
2 days agoKubernetes
5 years of experience, Kubernetes, Terraform, OpenTelemetry, Prometheus, eBPF, GitOps, Kustomize, Helm, AI-centric workflow, Cloud provider familiarity
Systems Observability Specialist
2 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, Distributed tracing, Structured logging, Go, Python, Java, High-cardinality metrics, SLOs, Error budgets, SRE principles, CI/CD integration, Linux internals, Networking, Container platforms, Thanos, Mimir, Cortex, Loki, Tempo, eBPF, Cost optimization, Regulated environments
Observability Engineer
3 days ago12 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, Distributed tracing, High-cardinality metrics, SLOs, CI/CD integration, Linux internals, eBPF-based observability
Site Reliability Engineer (SRE)
3 days agoPython · AWS · Azure
10 years of experience, Kubernetes, Python, Go, Prometheus, Grafana, CI/CD pipelines, Chaos engineering, Distributed systems, SLOs and error budgets, Cloud platforms (AWS, Azure, GCP)
Senior Systems Engineer | Shared Tech
3 days agoAWS · Azure · GCP
Production Linux, Infrastructure as Code (IaC), Cloud platforms (AWS, GCP, Azure), Container technologies (Docker, Kubernetes), Automation & Scripting, Collaborative Mindset
Cloud Infrastructure Engineer – AWS
4 days agoPython · AWS · Kubernetes
Bachelor’s or Master’s degree in CS, IT, Engineering or related field; 10+ years in IT/cloud engineering including 5+ years designing and operating enterprise AWS; EC2/VPC/IAM/S3/RDS/Lambda and other AWS services; Terraform, AWS CDK or CloudFormation; production EKS, ECS or Kubernetes; CI/CD/DevOps/GitOps; Python and Bash (Go/PowerShell preferred); IAM, encryption, compliance and observability. Preferred: advanced AWS certifications and other qualifications in posting.
Cloud Infrastructure Network Engineer
4 days agoKubernetes
Bachelor’s degree in CS, Networking or related field; 5+ years networking with substantial cloud networking; at least one major cloud provider; routing, switching and BGP; Direct Connect, ExpressRoute or equivalent hybrid connectivity; Terraform for cloud networking; security controls; Kubernetes networking/service mesh fundamentals; packet-level troubleshooting. Preferred: cloud networking certification, multi-cloud, SD-WAN/SASE, eBPF, regulated environments.
Cloud Networking Engineer
4 days agoPython · Kubernetes
Bachelor’s degree in Computer Science or related field; 5+ years in platform engineering, SRE or networking; production Istio or Linkerd; Envoy, Kubernetes networking/CNI/ingress, mTLS/PKI and certificate lifecycle, distributed tracing; Go or Python; networking troubleshooting. Preferred: multi-cluster mesh, Cilium/eBPF, SPIFFE/SPIRE, open-source contributions, enterprise zero-trust.
Container Platform Engineer
4 days agoPython · Kubernetes
Bachelor’s degree in CS, Engineering or related field; 5+ years operating production container platforms including 3+ years on Red Hat OpenShift; Kubernetes/OpenShift internals, Linux, Ansible/Terraform/Helm, Tekton/Jenkins/Argo CD, Bash/Python/Go, cluster observability and image security. Preferred: Red Hat certification, public cloud, service mesh, regulated environments and GitOps.
Monitoring Engineer
4 days agoPython · Java
Bachelor’s degree in Computer Science or related field; 5+ years in SRE, platform engineering or observability; Prometheus, Grafana and one commercial platform such as Datadog/New Relic/Splunk; OpenTelemetry, tracing, structured logging; Go, Python or Java; high-throughput metrics/log pipelines; SLOs/error budgets; CI/CD and incident management; Linux, networking and containers. Preferred: Thanos/Mimir/Cortex/Loki/Tempo, eBPF, observability cost optimization.
Service Mesh Architect
4 days agoPython · Kubernetes
Bachelor’s degree in Computer Science or related field; 5+ years in platform engineering, SRE or networking; production Istio or Linkerd; Envoy, Kubernetes networking/CNI/ingress, mTLS/PKI/certificates, distributed tracing; Go or Python; networking and control-plane troubleshooting. Preferred: multi-cluster mesh, Cilium/eBPF, open-source contributions, SPIFFE/SPIRE, enterprise zero-trust.
Reliability Engineer
4 days agoPython · Java · AWS
Bachelor’s degree in Computer Science, Engineering or related field; 5+ years of SRE, DevOps or production engineering for large distributed systems; Python, Go or Java; Linux at scale; production Kubernetes; observability; CI/CD; distributed system design; incident response and post-incident reviews. Preferred: SLOs/error budgets, chaos engineering, AWS/Azure/GCP, capacity planning and service mesh.
Senior/Lead Site Reliability Engineer
5 days agoPython · Java · Go
Linux, TCP/IP, HTTP, C/C++, Python, Golang, Rust, Java, Open-source software, AIOps, AI operations, Native Chinese proficiency
Cloud Infrastructure Engineer (Open LMS) Colombia, Remote
5 days agoPython · AWS
AWS services (EC2, RDS, S3, SQS, Lambda), Terraform, Puppet, Python, Deep Linux systems knowledge, Distributed systems concepts, Observability pipelines (Prometheus, Grafana), Multi-tenant SaaS platform design, Familiarity with Moodle LMS
Service Mesh Engineer
5 days agoPython · Kubernetes
Headline says 7+ years; detailed required qualifications specify 5+ years of platform engineering, SRE or networking, production Istio or Linkerd, Envoy, Kubernetes CNI, mTLS/PKI, distributed tracing and Go or Python. Cilium, SPIFFE/SPIRE and zero-trust networking are preferred.
OpenShift Platform Engineer
5 days agoPython · Kubernetes
Headline says 8+ years; detailed required qualifications specify 5+ years operating production container platforms including at least 3 years on Red Hat OpenShift, bachelor's degree, Kubernetes/OpenShift internals, Linux, Ansible or Terraform or Helm, Tekton or Jenkins or Argo CD, Bash or Python or Go, observability and image security. Service mesh and GitOps are preferred.
Senior Infrastructure Engineer | Permanent WFH | Day Shift & Night Shift
8 days agoAzure
5 years of experience, Windows Server Administration, Active Directory, Microsoft 365 Administration, VMware vSphere, Backup and recovery solutions, Storage technologies, Microsoft Azure, PowerShell, Relevant certifications
DevOps & SRE Engineer
8 days agoPython · Java · AWS
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Prometheus, Grafana, CI/CD pipelines, SLOs, Chaos engineering, AWS, Azure, GCP
Telemetry Engineer
8 days agoPython · Java
6 years of experience, Prometheus, Grafana, OpenTelemetry, Datadog, New Relic, Splunk, Go, Python, Java, SRE principles, distributed tracing, structured logging, Linux, containers, CI/CD, observability cost optimization, eBPF-based observability
Senior Infrastructure Engineer | Permanent WFH | Night Shift
8 days ago5 years of experience, Windows Server Administration, Active Directory, Microsoft 365 Administration, VMware vSphere, Backup and recovery solutions, Networking technologies, PowerShell, Cloud platforms
Kubernetes Service Engineer
8 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, mTLS, Kubernetes, Go, Python, Distributed tracing, Service mesh architecture, Traffic management policies, Zero-trust networking
Senior Data Ops Engineer, Data Activation & Products - Activision
8 days agoKubernetes
5 years of experience, Kubernetes, Databricks, Spark, Airflow, Kafka, Cloud infrastructure, Incident management, Automation frameworks, Observability tools, Event-driven architecture
Site Reliability Engineer (India)
9 days agoGo · AWS · GCP
4 years of experience, AWS, GCP, Kubernetes, Golang, CI/CD, Prometheus, Grafana, ELK stack, Core networking, Security best practices, Systematic troubleshooting
Kubernetes & OpenShift Engineer
9 days agoPython · AWS · Azure
6 years of experience, Kubernetes internals, OpenShift internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Cluster monitoring and logging, Container image security, Red Hat OpenShift experience, Public cloud experience (AWS, Azure, GCP)
Platform Networking Engineer
9 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, Service mesh production experience
Site Reliability Engineer (Europe)
9 days agoGo · AWS · GCP
Golang, AWS, Google Cloud, Kubernetes, CI/CD pipelines, Prometheus, Grafana, ELK stack, Advanced Linux internals, Core networking, Security best practices
Site Observability Engineer
9 days ago6 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, Distributed tracing, High-cardinality metrics, SLOs, CI/CD integration, Linux internals, Container platforms