Remote Site Reliability Engineer jobs.
“Remote”? Depends where. $$$? When they say. Stack? Before the essay.
Where can you apply from?
78 openings. None older than 30 days.
Newest first
Senior Platform Engineer - Platform Metal | Ireland | Remote
1 day agoPython · Kubernetes
Kubernetes, Terraform, Datacenter experience, Distributed systems, Go, Python, Shell, Crossplane, Cluster-API, Tinkerbell, Talos, Ceph, CSP experience
Senior HPC Cluster Engineer - AI, ML
2 days agoPython · Docker
5 years of experience, AI/HPC advanced job schedulers, Slurm, Centos/RHEL, Ubuntu Linux, Cluster configuration management, Docker, Python, MPI, NVIDIA GPUs, CUDA Programming, InfiniBand, Lustre
Site Reliability Engineer – Localization
3 days agoSQL · AWS · Azure
3 years of experience, AWS, Azure, Infrastructure as Code, Linux administration, Networking fundamentals, SQL, NoSQL, Containerized architectures, Serverless architectures
Site Reliability Engineer / DevOps – Localization
3 days agoSQL · AWS · Azure
3 years of experience, AWS, Azure, Infrastructure as Code, Linux administration, Core networking fundamentals, SQL, NoSQL, Containerized architectures, Serverless architectures
Site Reliability Engineer
3 days agoGo · AWS · GCP
Kubernetes, Golang, AWS, Google Cloud, Cloud-native environment, Distributed database systems, CI/CD optimization, Monitoring and observability tools, Disaster recovery strategies, Troubleshooting complex systems
Operations Support Engineer
3 days ago5 years of experience, ITIL-aligned frameworks, Second-level support for mission-critical systems, Observability tools, Linux/Unix environments, Container orchestration, CI/CD pipelines, Middleware and integration technologies, Database performance monitoring, System hardening and patching, English fluency
Azure Infrastructure Engineer
3 days agoPython · Azure · Kubernetes
6 years of experience, Azure core services, Infrastructure-as-code (Terraform, Bicep, ARM), Azure Kubernetes Service (AKS), Azure DevOps or GitHub Actions, PowerShell, Bash, Python, Cloud security principles, Monitoring and observability strategies, Hybrid cloud or multi-cloud experience, FinOps practices, Regulated environments (HIPAA, PCI-DSS, SOC 2, FedRAMP)
Observability Engineer
3 days ago12 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, Distributed tracing, High-cardinality metrics, SLOs, CI/CD integration, Linux internals, eBPF-based observability
Site Reliability Engineer (SRE)
3 days agoPython · Kubernetes
10 years of experience, Kubernetes, Python, Go, Chaos engineering, SLOs and error budgets, Prometheus, Grafana, CI/CD pipelines, Distributed systems design, Incident response leadership
Messaging Systems Engineer
3 days agoPython
6 years of experience, Apache Kafka, Confluent Platform, Kafka internals, Kafka security, Kafka Connect, Schema Registry, Kafka Streams, Infrastructure-as-code, Observability tooling, Scripting in Python, Bash, or Go
Principal Engineer, Cloud Site Reliability Engineering
3 days agoPython · Java · SQL
15 years of experience, AI development, Cloud infrastructure maintenance, JAVA, Python, Shell scripting, Distributed systems, SQL/NoSQL databases, Docker, Kubernetes, OpenStack, Machine Learning, High-performance software design, Scalable software systems
Senior Systems Software Engineer - Infrastructure
4 days agoPython · Kubernetes
12 years of experience, CI/CD architecture, Test automation architecture, Python fluency, Infrastructure-as-code, Distributed systems, Cloud computing, Kubernetes, Database schema design, Technical proposal writing, Device BIOS development
Monitoring Engineer
4 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, distributed tracing, structured logging, Go, Python, Java, high-cardinality metrics, SLOs, error budgets, SRE principles, CI/CD integration, Linux internals, networking, container platforms, Thanos, Mimir, Cortex, Loki, Tempo, eBPF, observability cost optimization, regulated environments
OpenShift Platform Engineer
4 days agoPython · Kubernetes
8 years of experience, OpenShift, Kubernetes internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Monitoring and logging (Prometheus, Grafana, EFK, Tempo), Container image security, Multi-tenant platform design, Disaster recovery strategies
Reliability Engineer
4 days agoPython · Java · AWS
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Prometheus, Grafana, CI/CD pipelines, Distributed system design, Incident response, SLOs and error budgets, Chaos engineering, AWS, Azure, GCP
Senior AI Infrastructure & Platform Operations Engineer (remote in the EU)
5 days agoKubernetes
7 years of experience, NVIDIA GPU infrastructure, Kubernetes in production environments, Technical leadership, Root cause analysis, Infrastructure automation technologies, Observability platforms, High-performance networking, AI infrastructure environments, Performance analysis of distributed platforms
Senior Site Reliability Engineer - Cloud
5 days agoPython · AWS · Kubernetes
8 years of experience, Akamai CDN, AWS Cloud Platform, Kubernetes, Python scripting, Generative AI/LLM applications, Edge computing/edge AI, Incident management process
Cloud Infrastructure Engineer – AWS
5 days agoPython · AWS · Kubernetes
10 years of experience, AWS, Terraform, AWS CloudFormation, Kubernetes, DevOps, Cloud Security, Site Reliability Engineering, CI/CD, Python, Bash, AWS Organizations, FinOps, Disaster Recovery, Cloud Governance
Cloud Infrastructure Network Engineer
5 days agoKubernetes
6 years of experience, Cloud networking, VPC/VNet design, Hybrid connectivity, Infrastructure-as-code (Terraform), Cloud security, Kubernetes networking, Multi-cloud networking, SD-WAN familiarity, eBPF-based networking tools, Regulated industries exposure
Observability Engineer
5 days ago6 years of experience, Prometheus, Grafana, Commercial observability platforms, OpenTelemetry, Distributed tracing, Structured logging, High-cardinality metrics, SLOs, CI/CD integration, Linux internals, Networking, Container platforms, Thanos, Mimir, Cortex, Loki, Tempo, eBPF-based tooling, Cost optimization initiatives, Regulated environments
OpenShift Platform Engineer
5 days agoPython · Kubernetes
6 years of experience, OpenShift, Kubernetes internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Container image security, Cluster monitoring tools, Service mesh (OpenShift Service Mesh, Istio, Linkerd), GitOps workflows (Argo CD, Flux)
Service Mesh Architect
5 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, Cilium, SPIFFE/SPIRE, Zero-trust networking
Senior Site Reliability Engineer (SRE)
6 days agoJava · GCP · Kubernetes
3 years of experience, Java / Spring Boot, Terraform, Go, Kubernetes, Docker, OpenTelemetry, Gradle, JVM performance tuning, GCP experience
Reliability Monitoring Engineer
6 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, Go, Python, Java, SLOs, CI/CD integration, Linux internals, High-cardinality metrics, Observability cost optimization
Systems Reliability Engineer
6 days agoPython · Java · AWS
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Prometheus, Grafana, CI/CD pipelines, Distributed system design, Incident response, SLOs and error budgets, Chaos engineering, AWS, Azure, GCP
Site Reliability Engineer
7 days agoPython · AWS · Kubernetes
5 years of experience, AWS, Terraform, Kubernetes, Helm, Distributed systems, Observability practices, CI/CD pipelines, Python, Healthcare technology experience, Incident response
Senior Engineer
8 days agoAWS · Kubernetes
5 years of experience, Kubernetes, AWS, Infrastructure as Code, CI/CD automation, GitOps, Container networking technologies, Prometheus, Grafana, DevSecOps, Active Secret clearance
Senior Solutions Engineer
9 days agoPython · Kubernetes
7 years of experience, Kubernetes, AI/GPU Infrastructure, RDMA/RoCEv2, Linux Networking, Python, Ansible, Executive Communication
Infrastructure Engineer - Virtualization
9 days ago4 years of experience, KVM/QEMU, Proxmox, Linux-based virtualization, Infrastructure automation (Ansible), GPU-based systems, High-throughput networking, VM lifecycle management, Incident response
Infrastructure Engineer – Storage Platform
9 days agoKubernetes
4 years of experience, Ceph, Kubernetes, High-performance NAS, Weka, VAST Data, Storage performance analysis, Ansible, Terraform, RDMA, AI/ML workload support, Multi-data center storage operations