Remote Site Reliability Engineer jobs.
“Remote”? Depends where. $$$? When they say. Stack? Before the essay.
Where can you apply from?
71 openings. None older than 30 days.
Newest first
Senior/Lead Site Reliability Engineer
2 days agoPython · Java · Go
Linux, TCP/IP, HTTP, C/C++, Python, Golang, Rust, Java, Open-source software, AIOps, AI operations, Native Chinese proficiency
Cloud Infrastructure Engineer (Open LMS) Colombia, Remote
2 days agoPython · AWS
AWS services (EC2, RDS, S3, SQS, Lambda), Terraform, Puppet, Python, Deep Linux systems knowledge, Distributed systems concepts, Observability pipelines (Prometheus, Grafana), Multi-tenant SaaS platform design, Familiarity with Moodle LMS
OpenShift Platform Engineer
2 days agoPython · Kubernetes
8 years of experience, OpenShift, Kubernetes internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Container image security, Cluster monitoring tools, Service mesh (OpenShift Service Mesh, Istio, Linkerd), GitOps workflows (Argo CD, Flux)
Service Mesh Engineer
2 days agoPython · Kubernetes
7 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, Cilium, SPIFFE/SPIRE, Zero-trust networking
Senior Infrastructure Engineer | Permanent WFH | Day Shift & Night Shift
5 days agoAzure
5 years of experience, Windows Server Administration, Active Directory, Microsoft 365 Administration, VMware vSphere, Backup and recovery solutions, Storage technologies, Microsoft Azure, PowerShell, Relevant certifications
Senior Infrastructure Engineer | Permanent WFH | Night Shift
5 days ago5 years of experience, Windows Server Administration, Active Directory, Microsoft 365 Administration, VMware vSphere, Backup and recovery solutions, Networking technologies, PowerShell, Cloud platforms
DevOps & SRE Engineer
5 days agoPython · Java · AWS
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Prometheus, Grafana, CI/CD pipelines, SLOs, Chaos engineering, AWS, Azure, GCP
Kubernetes Service Engineer
5 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, mTLS, Kubernetes, Go, Python, Distributed tracing, Service mesh architecture, Traffic management policies, Zero-trust networking
Telemetry Engineer
5 days agoPython · Java
6 years of experience, Prometheus, Grafana, OpenTelemetry, Datadog, New Relic, Splunk, Go, Python, Java, SRE principles, distributed tracing, structured logging, Linux, containers, CI/CD, observability cost optimization, eBPF-based observability
Senior Data Ops Engineer, Data Activation & Products - Activision
5 days agoKubernetes
5 years of experience, Kubernetes, Databricks, Spark, Airflow, Kafka, Cloud infrastructure, Incident management, Automation frameworks, Observability tools, Event-driven architecture
Site Reliability Engineer (Europe)
6 days agoGo · AWS · GCP
Golang, AWS, Google Cloud, Kubernetes, CI/CD pipelines, Prometheus, Grafana, ELK stack, Advanced Linux internals, Core networking, Security best practices
Site Reliability Engineer (India)
6 days agoGo · AWS · GCP
4 years of experience, AWS, GCP, Kubernetes, Golang, CI/CD, Prometheus, Grafana, ELK stack, Core networking, Security best practices, Systematic troubleshooting
Platform Reliability Engineer
6 days agoPython · Java · Kubernetes
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Observability tooling, CI/CD pipelines, Distributed system design, SLOs and error budgets, Chaos engineering, Cloud platforms
Kubernetes & OpenShift Engineer
6 days agoPython · AWS · Azure
6 years of experience, Kubernetes internals, OpenShift internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Cluster monitoring and logging, Container image security, Red Hat OpenShift experience, Public cloud experience (AWS, Azure, GCP)
Platform Networking Engineer
6 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, Service mesh production experience
Site Observability Engineer
6 days ago6 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, Distributed tracing, High-cardinality metrics, SLOs, CI/CD integration, Linux internals, Container platforms
Azure Infrastructure Engineer
7 days agoPython · Azure · Kubernetes
6 years of experience, Azure core services, Infrastructure-as-code (Terraform, Bicep, ARM), Azure Kubernetes Service (AKS), Azure DevOps or GitHub Actions, PowerShell, Bash, Python, Cloud security principles, Monitoring and observability strategies, Hybrid cloud or multi-cloud experience
Cloud Infrastructure Network Engineer
7 days agoKubernetes
6 years of experience, Cloud networking, Terraform, Hybrid connectivity, Routing and switching, BGP, Cloud security, Kubernetes networking, Multi-cloud networking, SD-WAN familiarity, eBPF-based tools
Monitoring Engineer
7 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, Go, Python, Java, High-cardinality metrics, SLOs, CI/CD integration, Linux internals, eBPF-based tooling, Regulated environments
Observability Engineer
7 days ago6 years of experience, Prometheus, Grafana, Commercial observability platforms, OpenTelemetry, Distributed tracing, Structured logging, High-cardinality metrics, SRE principles, CI/CD integration, Linux internals, Networking, Container platforms
OpenShift Platform Engineer
7 days agoPython · Kubernetes
8 years of experience, OpenShift, Kubernetes internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Container image security, Service mesh (OpenShift Service Mesh, Istio, Linkerd), GitOps workflows (Argo CD, Flux)
Reliability Engineer
7 days agoPython · Java · AWS
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Prometheus, Grafana, CI/CD pipelines, Distributed system design, Incident response, SLOs and error budgets, Chaos engineering, AWS, Azure, GCP
Service Mesh Architect
7 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, Go, Python, Distributed tracing, Cilium, SPIFFE/SPIRE, Zero-trust networking
Site Reliability Engineer (SRE)
7 days agoPython · Kubernetes
10 years of experience, Kubernetes, Python, Go, Prometheus, Grafana, CI/CD pipelines, Chaos engineering, Distributed systems, SLOs and error budgets, Cloud platforms
Senior Site Reliability Engineer, DGX Cloud
8 days agoPython · Kubernetes
8 years of experience, Kubernetes administration, GPU workloads optimization, Infrastructure automation (Terraform, Ansible), High-level programming (Python, Go), Linux operating systems, SRE principles (SLOs, SLIs), Observability stacks (OpenTelemetry, Prometheus), GPU-accelerated clusters (KubeVirt), AI inference workloads (vLLM, PyTorch)
Site Reliability Engineer Technical Lead
8 days ago8 years of experience, Deep SRE and Systems Expertise, Automation and Tooling, Observability and Analysis, Exceptional leadership and communication skills
Systems Observability Specialist
8 days ago6 years of experience, Prometheus, Grafana, Commercial observability platforms, OpenTelemetry, Distributed tracing, Structured logging, High-cardinality metrics, SLOs, SRE principles, CI/CD integration, Linux internals, Networking, Container platforms
Site Reliability Engineer, Tech Lead
9 days agoPython · AWS · Kubernetes
5 years of experience, AWS, Kubernetes, Cloud Computing, SRE/DevOps, Reliability Engineering leadership, CI/CD pipelines, UNIX/Linux, Networking, Automation tools, Python scripting, Monitoring and incident management
Incident Operations Lead (EMEA/AMER)
10 days ago5 years of experience, Incident command experience, FinTech understanding, 24x7 team leadership, Reliability metrics development, AI automation for incident management, Distributed team management, Severity model ownership, Effective communication under pressure
Principal Infrastructure Engineer
10 days agoAWS · Kubernetes
On-call rotation and major-incident recovery, including full outages; related CS/technical bachelor’s degree; 12+ years across infrastructure, platform, SRE or software engineering; AWS, production Kubernetes and Aurora/RDS MySQL or PostgreSQL; production coding and Terraform/IaC; Linux, networking, distributed systems, disaster recovery and observability. Fintech, multi-region and chaos-engineering experience are preferred.