Remote Site Reliability Engineer jobs.
“Remote”? Depends where. $$$? When they say. Stack? Before the essay.
Where can you apply from?
86 openings. None older than 30 days.
Newest first
Site Reliability Engineer, Infrastructure Platforms — UK (Intermediate to Senior Staff)
1 day agoAWS · GCP · Kubernetes
Kubernetes, Terraform, Go, AWS, GCP, Infrastructure as Code, Observability practices, Automation, Incident response, Strong written communication
Senior Software Engineer - SRE
1 day agoSite Reliability Engineering experience, PostgreSQL, Temporal workflows, Grafana or Honeycomb, OpenTelemetry
Senior Storage Production Engineer - DGX Cloud
1 day agoKubernetes
8 years of experience, Distributed storage solutions, High-performance storage systems, Storage networking protocols, Linux-based storage automation, Infrastructure configuration management, Observability tools, Capacity planning, Disaster recovery strategies, Kubernetes storage solutions, Replication strategies
Senior Site Reliability Engineer, Compute
1 day agoKubernetes
6 years of experience, Go, Kubernetes, Nomad, Vault, Consul, Fault-tolerance design, Performance Monitoring Services, Production readiness analysis
Senior Site Reliability Engineer
3 days agoPython · AWS · Kubernetes
4 years of experience, Kubernetes, Terraform, AWS, Production observability, Python, Go, AI tools, HIPAA compliance, Cloud cost-efficiency
Cloud Infrastructure Engineer – AWS
3 days agoPython · AWS · Kubernetes
10 years of experience, AWS, Terraform, AWS CloudFormation, Kubernetes, DevOps, Cloud Security, Site Reliability Engineering, CI/CD, Python, Bash, AWS Organizations, FinOps, Disaster Recovery, Cloud Governance
Cloud Infrastructure Network Engineer
3 days agoKubernetes
6 years of experience, Cloud networking, VPC/VNet design, Hybrid connectivity, Infrastructure-as-code (Terraform), Cloud security and network controls, Kubernetes networking, Multi-cloud networking, SD-WAN familiarity, eBPF-based networking tools, Regulated industries exposure
Kafka Engineer
3 days agoKubernetes
7 years of experience, Kafka internals, Kafka security, Kafka Connect, Infrastructure-as-code, Observability tooling, DevOps practices, Kubernetes (Strimzi, Confluent Operator)
Observability Engineer
3 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, distributed tracing, structured logging, Go, Python, Java, high-cardinality metrics, SLOs, error budgets, SRE principles, CI/CD integration, Linux internals, networking, container platforms, Thanos, Mimir, Cortex, Loki, Tempo, eBPF, cost optimization, regulated environments
OpenShift Platform Engineer
3 days agoPython · Kubernetes
8 years of experience, OpenShift, Kubernetes internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Container image security, Cluster monitoring and logging, Service mesh (OpenShift Service Mesh, Istio, Linkerd), GitOps workflows (Argo CD, Flux)
SAP Basis Administrator
3 days agoPython · AWS · Azure
10 years of experience, SAP HANA, S/4HANA, ECC, BW/4HANA, Cloud migrations (AWS, Azure, GCP), SAP upgrades, HANA System Replication, SAP transport management, Ansible, PowerShell, Python, Strong troubleshooting skills
Service Mesh Architect
3 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, Cilium, SPIFFE/SPIRE, Zero-trust networking
Staff Site Reliability Engineer - AI Platform Runtime
3 days agoTypeScript · JavaScript · Python
10 years of experience, Python, Typescript, JavaScript, Go, AWS CDK, AWS CloudFormation, Terraform, CrossPlane, OpenTelemetry, Kubernetes, AWS, Azure, GCP, Distributed systems design, Automation, Observability, Incident management, Postmortem analysis, Technical strategy, Mentorship
Senior Infrastructure Engineer, SRE
5 days agoPython · AWS
5 years of experience, Cloud infrastructure, Reliability engineering, SLIs and SLOs, Datadog, Python, Terraform, AWS, Disaster recovery planning, On-call experience
Senior Staff Site Reliability Engineer
7 days agoPython · Java · AWS
10 years of experience, Incident Commander experience, Distributed systems expertise, Kubernetes and cloud-native infrastructure, Automation for incident management, AI/ML applied to operations, Infrastructure-as-code tooling, Observability tooling expertise, Strong programming skills (Python, Go, Java), Experience with public cloud platforms (AWS, Azure, GCP), Scaling reliability across distributed teams
Senior Site Reliability Engineer - HPC
7 days agoPython · Ruby · Kubernetes
5 years of experience, HPC cluster support, Slurm or LSF or Kubernetes, Infrastructure as Code (IaC), CI/CD techniques, Automated host lifecycle management, E2E observability, Python or Go or Perl or Ruby, Technical mentoring, Published technical write-ups, Open source component maintenance
Infiniband Network Engineer
7 days ago5 years of experience, Infiniband Network troubleshooting, Infiniband Network certifications, Linux background, Ethernet networking
Cloud Networking Engineer
7 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Distributed tracing, Go, Python, Multi-cluster deployments, Cilium, SPIFFE/SPIRE, Zero-trust networking
DevOps & SRE Engineer
7 days agoPython · Java · AWS
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Prometheus, Grafana, CI/CD pipelines, SLOs, Chaos engineering, AWS, Azure, GCP
Telemetry Engineer
7 days agoPython · Java
6 years of experience, Prometheus, Grafana, OpenTelemetry, Datadog, New Relic, Splunk, Go, Python, Java, SRE principles, distributed tracing, structured logging, Linux, containers, CI/CD
VMware Infrastructure Engineer
7 days agoAWS · Kubernetes
6 years of experience, vSphere, vSAN, NSX-T, PowerCLI, Tanzu Kubernetes Grid, disaster recovery patterns, VMware Cloud on AWS, VMware Certified Professional (VCP), Aria Operations
Platform Engineer
8 days agoAWS · Kubernetes
Kubernetes, Bare-metal infrastructure, Infrastructure as code (Terraform), CI/CD (GitLab CI), Linux/Unix administration, AWS (EC2, EKS, IAM), GPU-enabled infrastructure, YAML for Kubernetes manifests, Security clearance eligibility
Kubernetes Service Engineer
8 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, mTLS, Kubernetes, Go, Python, Distributed tracing, Traffic management policies, Service mesh architecture
Platform Reliability Engineer
8 days agoPython · Java · AWS
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Prometheus, Grafana, CI/CD pipelines, Distributed system design, SLOs and error budgets, Chaos engineering, AWS, Azure, GCP
Systems Observability Specialist
8 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, distributed tracing, structured logging, Go, Python, Java, high-cardinality metrics, high-throughput log pipelines, SLOs, error budgets, SRE principles, CI/CD integration, Linux internals, networking, container platforms
AWS Senior SRE Consultant
8 days agoPython · AWS · Kubernetes
AWS observability solutions, Automated disaster recovery, Infrastructure-as-Code, Python, Terraform, Docker, Kubernetes, DataDog, Splunk, Grafana, AWS CloudFormation, AWS CLI, AWS CDK, Customer-facing experience
Software Systems Engineer III,
9 days agoPython · AWS · Azure
8 years of experience, Kubernetes, CI/CD, GitOps, Infrastructure-as-code, Azure, AWS, GCP, Argo CD, Helm, Terraform, Python, MLOps, Kubeflow, AI/ML workloads, Monitoring tools, SRE, Platform engineering
Site Reliability Engineer
9 days agoPython · Go · Ruby
5 years of experience, AWS, Docker, Kubernetes, Infrastructure as Code, CI/CD, Golang, Python, Ruby, Grafana, Prometheus, ELK, PagerDuty, GitLab CI, Snowflake, Redshift, Spark, CDN/edge delivery infrastructure
NCX Senior Engineer
9 days agoKubernetes
8 years of experience, NVIDIA technologies, Kubernetes, GPU infrastructure management, Infrastructure observability tools, Automation for lifecycle management, Linux-based distributed systems, Cloud infrastructure operations, Collaboration with cloud partners
Senior IT Infrastructure Engineer
9 days agoAWS
5 years of experience, AWS administration, Linux server administration, Infrastructure as code, macOS fleet management, Identity provider administration, Terraform, Incident response, Data platform infrastructure