Remote Site Reliability Engineer jobs.
“Remote”? Depends where. $$$? When they say. Stack? Before the essay.
Where can you apply from?
90 openings. None older than 30 days.
Newest first
Cloud Operations Engineer
2 days agoAWS · Azure
7 years of experience, Microsoft Azure, AWS, Linux administration, Office 365 migrations, Intune configuration, AWS GovCloud, Cloud security practices, System troubleshooting, Automation tools, Security tools knowledge, Analytical skills, Interpersonal communication, Team collaboration
Senior Staff Site Reliability Engineer
2 days agoPython · AWS · Kubernetes
10 years of experience, Kubernetes, AI/ML infrastructure, OpenTelemetry, AWS CDK, Terraform, Python, Go, Distributed systems debugging, Agentic AI platforms, Automation systems, Technical strategy steering
Systems Operations and Administrator
2 days ago5 years of experience, Linux administration, Enterprise networking standards, Data-center infrastructure, Networking knowledge (TCP/IP, DNS, DHCP), Scripting and automation, Infrastructure monitoring tools, Capacity planning, Operational process establishment
Sr. Staff Platform/Data Reliability Engineer, Databricks (R5537)
2 days ago12 years of experience, Databricks, CI/CD for data platforms, Observability, Incident management, Compute policy design, Regulated environment experience, Delta Lake, Infrastructure-as-code, Defense or aerospace industry experience
Infrastructure Security Engineer
3 days agoAWS · Kubernetes
5 years of experience, AWS, Kubernetes, CI/CD security integrations, Terraform, Helm, Flux, ArgoCD, Okta, Web Application Firewalls, monitoring and logging platforms, security policies as code, incident response, active Secret clearance
Apache Kafka Developer
3 days agoPython · Kubernetes
6 years of experience, Kafka internals, Kafka security, Kafka Connect, Schema Registry, Kafka Streams, HA/DR strategies, Python scripting, Terraform, Observability tooling, Confluent Certified Administrator, Kafka on Kubernetes
Monitoring Engineer
3 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, distributed tracing, structured logging, Go, Python, Java, high-cardinality metrics, SLOs, error budgets, SRE principles, CI/CD integration, Linux internals, networking, container platforms, Thanos, Mimir, Cortex, Loki, Tempo, eBPF-based tooling, observability cost optimization, regulated environments
Observability Engineer
3 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, Distributed tracing, Structured logging, Go, Python, Java, High-cardinality metrics, SLOs, Error budgets, SRE principles, CI/CD integration, Linux internals, Networking, Container platforms, Thanos, Mimir, Cortex, Loki, Tempo, eBPF, Cost optimization, Regulated environments
Reliability Engineer
3 days agoPython · Java · AWS
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Prometheus, Grafana, CI/CD pipelines, Distributed system design, Incident response, SLOs and error budgets, Chaos engineering, AWS, Azure, GCP
IT Systems Engineer
4 days agoPython · AWS · GCP
3 years of experience, Python, PowerShell, Infrastructure as Code (IaC), CI/CD tools, AWS, GCP, Ansible, Terraform, Zabbix, Prometheus, Grafana, Analytical thinking, Bilingual (Polish and English)
Service Reliability Engineer
4 days agoPython · AWS · Azure
8 years of experience, Kubernetes, SLURM, large-scale cluster management, GPU hardware, high-performance computing, observability tools, incident management tools, AWS, Azure, GCP, OCI, Linux system administration, Ansible, Python, shell scripting, DNS, DHCP, storage systems, core networking, problem-solving, bare-metal infrastructure
Solutions Architect, Ethernet Networking - NVIS
4 days agoPython · Kubernetes · Docker
5 years of experience, Ethernet networking expertise, BGP, VxLAN, EVPN, Data center architecture, Network automation (Ansible, Salt, Python), Advanced network troubleshooting, Linux administration, Customer-facing experience, AI tools usage, Kubernetes, Docker, Networking simulation tools (NVIDIA Air, GNS3, EVE-NG), Network management tools (Grafana, Prometheus, Datadog)
Senior Site Reliability Engineer - Storage
4 days agoPython · Go · AWS
8 years of experience, HPC storage solutions, Enterprise NAS solutions, Distributed filesystems (Lustre, GPFS), Python/Bash/Golang, Cloud services (AWS, Azure, GCP), Monitoring stacks (Prometheus, Grafana, etc.), RDMA fabrics (InfiniBand, RoCE), HPC cluster management tools (Slurm, PBS, LSF), Containerization (Docker, Kubernetes)
Site Reliability Engineer
4 days agoAWS · Kubernetes
7 years of experience, Event-driven architecture, Deep AWS, Infrastructure as Code, Kubernetes, Observability and SLOs, Chaos engineering, Distributed systems debugging, Proven technical leadership, AI / MLOps infrastructure, Experience in payments industry
Site Reliability Engineer
4 days agoNode.js · Python · Java
Ansible, Terraform, Kubernetes, Linux, Windows, Python, Java, Golang, Node.js, Nginx, HAProxy, Docker, cloud-first mindset, security-first mindset
Cloud Engineer – Windows & Linux Platform Automation
4 days agoPython · SQL · AWS
4 years of experience, Microsoft platform management, Linux platform management, Public cloud services (Azure, AWS, GCP), Production networking, Production automation, SQL Server, PostgreSQL, MongoDB, Elasticsearch, Python, Terraform, PowerShell, Container scheduling engines (Mesos, Docker, Kubernetes), CI/CD solutions (GitHub Actions, Jenkins, CircleCI, ArgoCD), Level-3 support in ticketing systems
Cloud Infrastructure Engineer – AWS
4 days agoPython · AWS · Kubernetes
10 years of experience, AWS, Terraform, AWS CloudFormation, Kubernetes, DevOps, Site Reliability Engineering, Cloud security, CI/CD pipelines, Python, Bash, FinOps, AWS Organizations, Zero Trust architecture, Observability solutions
Cloud Infrastructure Network Engineer
4 days agoKubernetes
6 years of experience, Cloud networking, VPC/VNet design, Hybrid connectivity, Infrastructure-as-code (Terraform), Cloud security, Kubernetes networking, Multi-cloud networking, SD-WAN familiarity, eBPF-based networking tools, Regulated industries exposure
DevOps & SRE Engineer
4 days agoPython · Java · Kubernetes
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Observability tooling, CI/CD pipelines, SLOs, Chaos engineering, Cloud platforms
OpenShift Administrator
4 days agoPython · Kubernetes
6 years of experience, OpenShift, Kubernetes internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Cluster monitoring and logging, Container image security, Service mesh (OpenShift Service Mesh, Istio, Linkerd), GitOps workflows (Argo CD, Flux)
Reliability Monitoring Engineer
4 days ago6 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, SLOs, High-cardinality metrics, CI/CD integration, Linux internals, Container platforms, Observability cost optimization
Service Mesh Architect
4 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, Cilium, SPIFFE/SPIRE, Zero-trust networking
Systems Reliability Engineer
4 days agoPython · Java · AWS
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Prometheus, Grafana, CI/CD pipelines, Distributed system design, Incident response, SLOs and error budgets, Chaos engineering, AWS, Azure, GCP
Telemetry Engineer
4 days agoPython · Java
6 years of experience, Prometheus, Grafana, OpenTelemetry, Datadog, New Relic, Splunk, Go, Python, Java, SRE principles, incident management, distributed tracing, structured logging, Linux, containers, CI/CD, observability cost optimization, eBPF-based observability
Staff Platform Engineer
5 days agoKubernetes
10 years of experience, Kubernetes, Infrastructure-as-code (Terraform), Database proficiency (Postgres), GitOps model (Argo), Security mindset, AI in engineering, Architectural judgment, Multi-tenant isolation design, Compliance-heavy deployments (FedRAMP), Excellent communication and collaboration
Senior Site Reliability Engineer (SRE & Platform Reliability)
5 days agoPython · Kotlin · AWS
4 years of experience, Bash, Python, Kotlin, AWS, MySQL, Kubernetes, Incident Lifecycle, Distributed systems, Capacity management, Automation, Observability, Configuration management
Senior Site Reliability Engineer (SRE & Platform Reliability)
5 days agoPython · Kotlin · AWS
4 years of experience, Bash, Python, Kotlin, AWS, MySQL, Kubernetes, Incident Lifecycle, Distributed systems, Capacity management, Automation, Observability, Configuration management
Cloud Engineer – Windows & Linux Platform Automation
5 days agoPython · SQL · AWS
4 years of experience, Microsoft platform management, Linux platform management, Public cloud services (Azure, AWS, GCP), Production networking, Configuration-management tools, SQL Server, PostgreSQL, MongoDB, Elasticsearch, Python, Terraform, PowerShell, Container scheduling engines (Mesos, Docker, Kubernetes), CI/CD solutions (GitHub Actions, Jenkins, CircleCI, ArgoCD), Level-3 support in ticketing systems
Principal Site Reliability Engineer
7 days agoPython
10 years of experience, Openshift, Nutanix AHV, VMware vSphere, RedHat OpenShift, Python, Go, Ansible, GitHub, 24x7 operations
Senior Platform Engineer - Platform Metal | Ireland | Remote
7 days agoPython · Kubernetes
Kubernetes, Terraform, Datacenter experience, Distributed systems, Go, Python, Shell, Crossplane, Cluster-API, Tinkerbell, Talos, Ceph, CSP experience