Remote Site Reliability Engineer jobs.
“Remote”? Depends where. $$$? When they say. Stack? Before the essay.
Where can you apply from?
103 openings. None older than 30 days.
Newest first
Senior Platform Engineer
about 19 hours agoTypeScript · React · AWS
5 years of experience, AWS (IAM, multi-account, EKS), Terraform, Kubernetes, Building internal developer portals, Strong programming ability in modern languages, Experience with CI/CD pipelines, Improving existing workflows, TypeScript, React, AI coding assistants, Platform as a product, Clear writing for documentation
Platform Engineer
about 19 hours agoAWS · Kubernetes
Kubernetes, Bare-metal infrastructure, Infrastructure as code (Terraform), CI/CD (GitLab CI), Linux/Unix administration, AWS (EC2, EKS, IAM), GPU-enabled infrastructure, YAML for Kubernetes manifests, Security clearance eligibility
Software Systems Engineer III,
1 day agoPython · AWS · Azure
8 years of experience, Kubernetes, CI/CD, GitOps, Infrastructure-as-code, Azure, AWS, GCP, Argo CD, Helm, Terraform, Python, MLOps, Kubeflow, AI/ML workloads, Monitoring tools, SRE, Platform engineering
Site Reliability Engineer
1 day agoPython · Go · Ruby
5 years of experience, AWS, Docker, Kubernetes, Infrastructure as Code, CI/CD, Golang, Python, Ruby, Grafana, Prometheus, ELK, PagerDuty, GitLab CI, Snowflake, Redshift, Spark, CDN/edge delivery infrastructure
Site Reliability Engineer
2 days agoKubernetes
8 years of experience, Linux, Kubernetes, CI/CD, Bare metal servers, Data center management, Ansible, Kafka, Observability tools, Networking fundamentals, Automation of provisioning
NCX Senior Engineer
2 days agoKubernetes
8 years of experience, NVIDIA technologies, Kubernetes, GPU infrastructure management, Infrastructure observability tools, Automation for lifecycle management, Linux-based distributed systems, Cloud infrastructure operations, Collaboration with cloud partners
Reliability Engineer
3 days agoPython · Java · AWS
6 years of experience, Python, Go, Java, Kubernetes, Linux, Prometheus, Grafana, CI/CD pipelines, SLOs, Chaos engineering, AWS, Azure, GCP
Senior Staff Site Reliability Engineer
3 days agoPython
15 years of experience, eBPF, XDP, Containerization architectures, Distributed Systems Infrastructure, Terraform, Config Management tools, Go, Python, Linux Kernel Internals, Network Protocols (VLAN/VxLAN/SDN/BGP/Anycast), DNS, LDAP, Microservices architecture, Infrastructure as Code (IaC)
Senior Site Reliability Engineer - Storage
3 days agoPython · Go · AWS
8 years of experience, HPC storage solutions, Enterprise NAS solutions, Distributed filesystems, Python, Bash, Golang, Cloud services (AWS, Azure, GCP), Monitoring stacks (Prometheus, Grafana, etc.), RDMA fabrics, HPC cluster management tools, Containerization technologies (Docker, Kubernetes)
Canada- Jr Software Engineer (SRE)
3 days agoAzure · Kubernetes
2 years of experience, Basic programming, Scripting, Azure, Kubernetes, Observability platforms, AI interest, Automation, Distributed systems
Apache Kafka Developer
3 days agoPython · Kubernetes
6 years of experience, Kafka internals, Kafka security, Kafka Connect, Infrastructure-as-code, HA/DR strategies, Python scripting, Kubernetes Kafka, Managed Kafka services, Stream processing frameworks
Monitoring Engineer
3 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, distributed tracing, structured logging, Go, Python, Java, high-cardinality metrics, SLOs, error budgets, SRE principles, CI/CD integration, Linux internals, networking, container platforms, Thanos, Mimir, Cortex, Loki, Tempo, eBPF-based tooling, observability cost optimization, regulated environments
OpenShift Platform Engineer
3 days agoKubernetes
8 years of experience, OpenShift, Kubernetes internals, Linux administration, GitOps, CI/CD pipelines, Infrastructure-as-code, Container security, Monitoring tools, Disaster recovery strategies, Public cloud experience
Engineering / Platform Professionals
3 days ago2 years of experience, Technical document analysis, Cloud infrastructure experience, Attention to detail, Ability to interpret technical concepts, Experience reviewing engineering reports, Familiarity with technical architecture documentation
Datadog Administration and Operations (Servicenow)
3 days agoPython · SQL
5 years of experience, Datadog expertise, ServiceNow integration, Observability strategy design, Scripting (Python/PowerShell/Bash), Cloud cost optimization, CI/CD integration, Network performance monitoring, Database monitoring (Postgres/SQL Server/Oracle/MySQL), Automation of monitoring processes, ITIL v4 Foundation certification
Cloud Infrastructure Network Engineer
4 days agoKubernetes
6 years of experience, Cloud networking, VPC/VNet design, Hybrid connectivity, Infrastructure-as-code (Terraform), Cloud security and network controls, Kubernetes networking, Multi-cloud networking, SD-WAN familiarity, eBPF-based networking tools, Regulated industries exposure
Observability Engineer
4 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, distributed tracing, structured logging, Go, Python, Java, high-cardinality metrics, high-throughput log pipelines, SLOs, error budgets, SRE principles, CI/CD integration, Linux internals, networking, container platforms
OpenShift Administrator
4 days agoPython · Kubernetes
6 years of experience, OpenShift, Kubernetes internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Container image security, Cluster monitoring tools, Service mesh (OpenShift Service Mesh, Istio, Linkerd)
OpenShift Platform Engineer
4 days agoPython · Kubernetes
8 years of experience, OpenShift, Kubernetes internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Container image security, Cluster monitoring tools, Service mesh (OpenShift Service Mesh, Istio, Linkerd), GitOps workflows (Argo CD, Flux)
Service Mesh Architect
4 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, Cilium, SPIFFE/SPIRE, Zero-trust networking
Senior Site Reliability Engineer
7 days agoPython · Kubernetes
6 years of experience, Cloud-native infrastructure, Kubernetes deployment, Terraform, Ansible, CI/CD tooling, Distributed systems troubleshooting, Python, Bash, CUDA-based GPU programs, Security in sensitive environments
Senior Software Engineer, Site Reliability & Security
7 days agoTypeScript · Python · Go
5 years of experience, AWS or GCP, Kubernetes, Docker, Golang, Typescript, Python, Infrastructure as Code, GitOps, Grafana, Prometheus, UNIX shell
Sr. Production Engineer
7 days agoPython · AWS · Azure
4 years of experience, AWS, Azure, GCP, Python, Go, Prometheus, Grafana, OpenTelemetry, AI/ML understanding, Incident management, ITIL frameworks, Infrastructure-as-Code, Chaos engineering, BGP, GRE, IPSec, HAProxy, DNS
Senior Staff Site Reliability Engineer - Compute Core Engineering
7 days ago12 years of experience, eBPF, DPU, Containerization, Distributed Systems, Terraform, Linux Kernel Internals, Network Protocols, Microservices Architecture, Infrastructure as Code
L2 Technician - #35272
7 days agoAzure
2 years of experience, Microsoft technologies, Azure Support, Active Directory, Exchange, Microsoft 365, Windows Servers, Excellent troubleshooting skills, MSP experience
DevOps & SRE Engineer
7 days agoPython · Java · AWS
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Prometheus, Grafana, CI/CD pipelines, SLOs, Chaos engineering, AWS, Azure, GCP, Service mesh technologies
Telemetry Engineer
7 days agoPython · Java
6 years of experience, Prometheus, Grafana, OpenTelemetry, Datadog, New Relic, Splunk, Go, Python, Java, SRE principles, incident management, distributed tracing, structured logging, Linux, containers, CI/CD, observability cost optimization, eBPF-based observability
VMware Infrastructure Engineer
7 days agoAWS · Kubernetes
6 years of experience, vSphere, vSAN, NSX-T, PowerCLI, Tanzu Kubernetes Grid, disaster recovery patterns, VMware Cloud on AWS, VMware Certified Professional (VCP)
Customer Reliability Engineer, Infrastructure
8 days agoPython · Kubernetes
5 years of experience, Kubernetes, Cloud infrastructure management, Production distributed systems, Linux, Customer issue handling, Strong communication, DevOps or CI/CD, Python scripting, Troubleshooting
Staff Site Reliability Engineer
8 days agoAWS
10 years of experience, AWS, Terraform, Puppet, Ansible, Linux administration, Containerization, PCI DSS compliance, Infrastructure as code, Scripting for automation, Technical leadership