Remote Site Reliability Engineer jobs.
“Remote”? Depends where. $$$? When they say. Stack? Before the essay.
Where can you apply from?
97 openings. None older than 30 days.
Newest first
Site Reliability Engineer
about 8 hours agoKubernetes
8 years of experience, Linux, Kubernetes, CI/CD, Bare metal servers, Data center management, Ansible, Kafka, Observability tools, Networking fundamentals, Automation of provisioning
Senior Staff Site Reliability Engineer
1 day agoPython
15 years of experience, eBPF, XDP, Containerization architectures, Distributed Systems Infrastructure, Terraform, Config Management tools, Go, Python, Linux Kernel Internals, Network Protocols (VLAN/VxLAN/SDN/BGP/Anycast), DNS, LDAP, Microservices architecture, Infrastructure as Code (IaC)
Senior Site Reliability Engineer - Storage
1 day agoPython · Go · AWS
8 years of experience, HPC storage solutions, Enterprise NAS solutions, Distributed filesystems, Python, Bash, Golang, Cloud services (AWS, Azure, GCP), Monitoring stacks (Prometheus, Grafana, etc.), RDMA fabrics, HPC cluster management tools, Containerization technologies (Docker, Kubernetes)
Cloud Infrastructure Network Engineer
2 days agoKubernetes
6 years of experience, Cloud networking, VPC/VNet design, Hybrid connectivity, Infrastructure-as-code (Terraform), Cloud security and network controls, Kubernetes networking, Multi-cloud networking, SD-WAN familiarity, eBPF-based networking tools, Regulated industries exposure
Observability Engineer
2 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, distributed tracing, structured logging, Go, Python, Java, high-cardinality metrics, high-throughput log pipelines, SLOs, error budgets, SRE principles, CI/CD integration, Linux internals, networking, container platforms
OpenShift Administrator
2 days agoPython · Kubernetes
6 years of experience, OpenShift, Kubernetes internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Container image security, Cluster monitoring tools, Service mesh (OpenShift Service Mesh, Istio, Linkerd)
OpenShift Platform Engineer
2 days agoPython · Kubernetes
8 years of experience, OpenShift, Kubernetes internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Container image security, Cluster monitoring tools, Service mesh (OpenShift Service Mesh, Istio, Linkerd), GitOps workflows (Argo CD, Flux)
Service Mesh Architect
2 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, Cilium, SPIFFE/SPIRE, Zero-trust networking
Senior Site Reliability Engineer
5 days agoPython · Kubernetes
6 years of experience, Cloud-native infrastructure, Kubernetes deployment, Terraform, Ansible, CI/CD tooling, Distributed systems troubleshooting, Python, Bash, CUDA-based GPU programs, Security in sensitive environments
Senior Software Engineer, Site Reliability & Security
5 days agoTypeScript · Python · Go
5 years of experience, AWS or GCP, Kubernetes, Docker, Golang, Typescript, Python, Infrastructure as Code, GitOps, Grafana, Prometheus, UNIX shell
Sr. Production Engineer
5 days agoPython · AWS · Azure
4 years of experience, AWS, Azure, GCP, Python, Go, Prometheus, Grafana, OpenTelemetry, AI/ML understanding, Incident management, ITIL frameworks, Infrastructure-as-Code, Chaos engineering, BGP, GRE, IPSec, HAProxy, DNS
Senior Staff Site Reliability Engineer - Compute Core Engineering
5 days ago12 years of experience, eBPF, DPU, Containerization, Distributed Systems, Terraform, Linux Kernel Internals, Network Protocols, Microservices Architecture, Infrastructure as Code
L2 Technician - #35272
5 days agoAzure
2 years of experience, Microsoft technologies, Azure Support, Active Directory, Exchange, Microsoft 365, Windows Servers, Excellent troubleshooting skills, MSP experience
DevOps & SRE Engineer
5 days agoPython · Java · AWS
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Prometheus, Grafana, CI/CD pipelines, SLOs, Chaos engineering, AWS, Azure, GCP, Service mesh technologies
Telemetry Engineer
5 days agoPython · Java
6 years of experience, Prometheus, Grafana, OpenTelemetry, Datadog, New Relic, Splunk, Go, Python, Java, SRE principles, incident management, distributed tracing, structured logging, Linux, containers, CI/CD, observability cost optimization, eBPF-based observability
VMware Infrastructure Engineer
5 days agoAWS · Kubernetes
6 years of experience, vSphere, vSAN, NSX-T, PowerCLI, Tanzu Kubernetes Grid, disaster recovery patterns, VMware Cloud on AWS, VMware Certified Professional (VCP)
Customer Reliability Engineer, Infrastructure
6 days agoPython · Kubernetes
5 years of experience, Kubernetes, Cloud infrastructure management, Production distributed systems, Linux, Customer issue handling, Strong communication, DevOps or CI/CD, Python scripting, Troubleshooting
Staff Site Reliability Engineer
6 days agoAWS
10 years of experience, AWS, Terraform, Puppet, Ansible, Linux administration, Containerization, PCI DSS compliance, Infrastructure as code, Scripting for automation, Technical leadership
Kubernetes Technical Support Engineer (565)
6 days agoAWS · Azure · Kubernetes
2 years of experience, Kubernetes, Kubernetes certification (CKA, CKAD, CKS, AWS, Azure), Technical support (Tier 2 / Tier 3), Troubleshooting (software, infrastructure, networking), Customer-facing technical experience, Independent investigation of technical issues, LATAM-based availability for U.S. East Coast hours
Kubernetes Service Engineer
6 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, mTLS, Kubernetes, Go, Python, Distributed tracing, CNI, Zero-trust networking
Platform Reliability Engineer
6 days agoPython · Java · Kubernetes
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Observability tooling, CI/CD pipelines, Distributed system design, SLOs and error budgets, Chaos engineering, Cloud platforms
Site Observability Engineer
6 days ago6 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, Distributed tracing, High-cardinality metrics, SLOs, CI/CD integration, Linux internals, Container platforms
Systems Observability Specialist
6 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, distributed tracing, structured logging, Go, Python, Java, high-cardinality metrics, SLOs, error budgets, SRE principles, CI/CD integration, Linux internals, networking, container platforms, Thanos, Mimir, Cortex, Loki, Tempo, OpenTelemetry contributions, eBPF, cost optimization, regulated environments
Site Reliability Engineer
7 days agoGCP · Kubernetes
3 years of experience, Kubernetes, GCP, Datadog, Terraform, CI/CD pipelines, SLO/SLI design, Incident response, E-commerce experience, Observability practices
Cloud Systems Engineer
7 days agoPython · AWS · Azure
5 years of experience, AWS, Azure, Microsoft 365, Terraform, PowerShell, Bash, Python, IAM, Security compliance, Linux, Windows Server, Defense sector experience, Automation of workflows, Documentation skills, AI platform administration
Senior Site Reliability Engineer | SRE (Remote, EU/CET)
7 days agoAWS
AWS, Infrastructure as code, Containerized workloads, Observability, Cloud security, Vulnerability management, SOC 2, ISO 27001, Automation, AI-assisted development, Production incident management
Staff Site Reliability Engineer
7 days agoPython · AWS · Kubernetes
7 years of experience, AWS, Kubernetes, Terraform, SLO frameworks, Incident response programs, Observability platforms, Python, Go, AI tools, Healthcare compliance (HIPAA), Mentorship
Senior Database Reliability Engineer
7 days agoGo · SQL · AWS
6 years of experience, Kubernetes, AWS, RDS (MySQL/Postgres), Golang, SQL-based RDBMS, Observability tools, Distributed systems design patterns, AI tooling, Microservices architecture, Engineering best practices
Senior Staff SRE – Compute Platform
7 days agoPython · Kubernetes
10 years of experience, Kubernetes administration, Bare-metal infrastructure, Python or Go, Infrastructure as Code, SRE and observability, HPC or AI infrastructure, VMware vSphere, Generative AI applications, Secure operational platforms, High-impact infrastructure projects
Senior Site Reliability Engineer
7 days agoPython
12 years of experience, eBPF, XDP, Terraform, Go, Python, Linux OS, DNS, LDAP, containerization architecture, distributed-systems infrastructure, network protocols