Remote Site Reliability Engineer jobs.
“Remote”? Depends where. $$$? When they say. Stack? Before the essay.
Where can you apply from?
84 openings. None older than 30 days.
Newest first
Infrastructure Security Engineer
1 day agoAWS · Kubernetes
5 years of experience, AWS, Kubernetes, CI/CD security integrations, Terraform, Helm, Flux, ArgoCD, Okta, Web Application Firewalls, monitoring and logging platforms, security policies as code, incident response, active Secret clearance
IT Systems Engineer
1 day agoPython · AWS · GCP
3 years of experience, Python, PowerShell, Infrastructure as Code (IaC), CI/CD tools, AWS, GCP, Ansible, Terraform, Zabbix, Prometheus, Grafana, Analytical thinking, Bilingual (Polish and English)
Service Reliability Engineer
2 days agoPython · AWS · Azure
8 years of experience, Kubernetes, SLURM, large-scale cluster management, GPU hardware, high-performance computing, observability tools, incident management tools, AWS, Azure, GCP, OCI, Linux system administration, Ansible, Python, shell scripting, DNS, DHCP, storage systems, core networking, problem-solving, bare-metal infrastructure
Solutions Architect, Ethernet Networking - NVIS
2 days agoPython · Kubernetes · Docker
5 years of experience, Ethernet networking expertise, BGP, VxLAN, EVPN, Data center architecture, Network automation (Ansible, Salt, Python), Advanced network troubleshooting, Linux administration, Customer-facing experience, AI tools usage, Kubernetes, Docker, Networking simulation tools (NVIDIA Air, GNS3, EVE-NG), Network management tools (Grafana, Prometheus, Datadog)
Senior Site Reliability Engineer - Storage
2 days agoPython · Go · AWS
8 years of experience, HPC storage solutions, Enterprise NAS solutions, Distributed filesystems (Lustre, GPFS), Python/Bash/Golang, Cloud services (AWS, Azure, GCP), Monitoring stacks (Prometheus, Grafana, etc.), RDMA fabrics (InfiniBand, RoCE), HPC cluster management tools (Slurm, PBS, LSF), Containerization (Docker, Kubernetes)
Site Reliability Engineer
2 days agoAWS · Kubernetes
7 years of experience, Event-driven architecture, Deep AWS, Infrastructure as Code, Kubernetes, Observability and SLOs, Chaos engineering, Distributed systems debugging, Proven technical leadership, AI / MLOps infrastructure, Experience in payments industry
Staff Platform Engineer
3 days agoKubernetes
10 years of experience, Kubernetes, Infrastructure-as-code (Terraform), Database proficiency (Postgres), GitOps model (Argo), Security mindset, AI in engineering, Architectural judgment, Multi-tenant isolation design, Compliance-heavy deployments (FedRAMP), Excellent communication and collaboration
Senior Site Reliability Engineer (SRE & Platform Reliability)
3 days agoPython · Kotlin · AWS
4 years of experience, Bash, Python, Kotlin, AWS, MySQL, Kubernetes, Incident Lifecycle, Distributed systems, Capacity management, Automation, Observability, Configuration management
Senior Site Reliability Engineer (SRE & Platform Reliability)
3 days agoPython · Kotlin · AWS
4 years of experience, Bash, Python, Kotlin, AWS, MySQL, Kubernetes, Incident Lifecycle, Distributed systems, Capacity management, Automation, Observability, Configuration management
Cloud Engineer – Windows & Linux Platform Automation
3 days agoPython · SQL · AWS
4 years of experience, Microsoft platform management, Linux platform management, Public cloud services (Azure, AWS, GCP), Production networking, Configuration-management tools, SQL Server, PostgreSQL, MongoDB, Elasticsearch, Python, Terraform, PowerShell, Container scheduling engines (Mesos, Docker, Kubernetes), CI/CD solutions (GitHub Actions, Jenkins, CircleCI, ArgoCD), Level-3 support in ticketing systems
Principal Site Reliability Engineer
5 days agoPython
10 years of experience, Openshift, Nutanix AHV, VMware vSphere, RedHat OpenShift, Python, Go, Ansible, GitHub, 24x7 operations
Senior Platform Engineer - Platform Metal | Ireland | Remote
5 days agoPython · Kubernetes
Kubernetes, Terraform, Datacenter experience, Distributed systems, Go, Python, Shell, Crossplane, Cluster-API, Tinkerbell, Talos, Ceph, CSP experience
Production Engineer (IC4)
5 days agoPython
3 years of experience, Python, RPM packaging, Enterprise OS modernization, Configuration management, CI/CD pipelines, Tier-2 operational support, Service onboarding, Monitoring and logging, Independent bug ownership
AWS Platform Specialist
5 days agoAWS
6 years of experience, AWS core services, Infrastructure-as-code (Terraform, AWS CDK), Amazon EKS or ECS clusters, CI/CD pipelines, Cloud security and compliance, Observability and monitoring, AWS Certified Solutions Architect – Professional, Multi-account AWS Organizations, FinOps practices, Regulated workloads (HIPAA, PCI-DSS)
Platform Infrastructure Engineer
5 days agoPython · AWS · Azure
6 years of experience, OpenShift, Kubernetes internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Cluster monitoring and logging, Container image security, Red Hat Certified Specialist in OpenShift Administration, Public cloud experience (AWS, Azure, GCP)
Service Infrastructure Engineer
5 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, CNI, Ingress
Sr Cloud Infrastructure & Platform Engineer
5 days agoPython · AWS · Kubernetes
AWS, Terraform, Kubernetes/EKS, Python, CI/CD, Cloud security best practices, Observability, Automation, Networking
Infrastructure Engineer, APAC
6 days agoPython · AWS · GCP
3 years of experience, Python, Terraform, AWS, GCP, Linux internals, Bash scripting, Ansible, Saltstack, Git, CI/CD pipelines, Containerization, Zabbix, Prometheus, Grafana, Opensearch, Elasticsearch
Network Automation & Reliability Engineer
6 days agoPython
5 years of experience, Python development, BGP policy management, Arista EOS, Juniper Junos, Cisco IOS-XR, Ansible, gNMI/gRPC telemetry, NETCONF/YANG, Linux operations, Git-based workflows, Production on-call experience
Cloud Networking Engineer
6 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, Go, Python, Distributed tracing, CNI, Ingress
Kubernetes Service Engineer
6 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, mTLS, Kubernetes, Go, Python, Distributed tracing, CNI, API gateway integration
Platform Reliability Engineer
6 days agoPython · Java · Kubernetes
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Observability tooling, CI/CD pipelines, Distributed system design, SLOs and error budgets, Chaos engineering, Cloud platforms
Site Observability Engineer
6 days ago6 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, Distributed tracing, High-cardinality metrics, SLOs, CI/CD integration, Linux internals, Container platforms
Systems Observability Specialist
6 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, distributed tracing, structured logging, Go, Python, Java, high-cardinality metrics, SLOs, error budgets, CI/CD integration, Linux internals, networking, container platforms
VMware Infrastructure Engineer
6 days agoAWS · Kubernetes
6 years of experience, vSphere, vSAN, NSX-T, PowerCLI, Tanzu Kubernetes Grid, disaster recovery, VMware Cloud on AWS, VMware Certified Professional (VCP)
Senior HPC Cluster Engineer - AI, ML
6 days agoPython · Docker
5 years of experience, AI/HPC advanced job schedulers, Slurm, Centos/RHEL, Ubuntu Linux, Cluster configuration management, Docker, Python, MPI, NVIDIA GPUs, CUDA Programming, InfiniBand, Lustre
Site Reliability Engineer – Localization
7 days agoSQL · AWS · Azure
3 years of experience, AWS, Azure, Infrastructure as Code, Linux administration, Networking fundamentals, SQL, NoSQL, Containerized architectures, Serverless architectures
Site Reliability Engineer / DevOps – Localization
7 days agoSQL · AWS · Azure
3 years of experience, AWS, Azure, Infrastructure as Code, Linux administration, Core networking fundamentals, SQL, NoSQL, Containerized architectures, Serverless architectures
Site Reliability Engineer
7 days agoGo · AWS · GCP
Kubernetes, Golang, AWS, Google Cloud, Cloud-native environment, Distributed database systems, CI/CD optimization, Monitoring and observability tools, Disaster recovery strategies, Troubleshooting complex systems
Operations Support Engineer
7 days ago5 years of experience, ITIL-aligned frameworks, Second-level support for mission-critical systems, Observability tools, Linux/Unix environments, Container orchestration, CI/CD pipelines, Middleware and integration technologies, Database performance monitoring, System hardening and patching, English fluency