Remote Site Reliability Engineer jobs.
“Remote”? Depends where. $$$? When they say. Stack? Before the essay.
Where can you apply from?
98 openings. None older than 30 days.
Newest first · page 3
Senior Site Reliability Engineer
21 days agoTypeScript · Python · Kubernetes
5 years of experience, Kubernetes (EKS, GKE), Terraform, Pulumi, GitOps (ArgoCD, Flux), Datadog, Prometheus, Grafana, SLI/SLO/error budget fluency, Go, Python, TypeScript, Linux internals, TCP/IP networking, Incident response leadership, Live-service game experience, Service mesh (Istio, Cilium), FinOps, Cloud certifications
Site Reliability Engineer
21 days agoPython · Azure · Kubernetes
3 years of experience, SLIs/SLOs definition, Multi-tenant SaaS platforms, Datadog, Grafana, Elastic Stack, High-availability architectures, Kubernetes, Python, Bash, Incident response, Cloud experience (Azure)
Site Reliability Engineer
21 days agoGo · AWS · GCP
Kubernetes, AWS, Google Cloud, Golang, CI/CD optimization, Distributed systems management, Monitoring tools (Prometheus, Grafana), Troubleshooting complex infrastructure issues, Disaster recovery strategies, Cloud security practices
Senior Site Reliability Engineer, BCM - DGX Cloud
21 days agoPython · Kubernetes
8 years of experience, Fluency in Python, In-depth knowledge of Linux, Cluster networking proficiency, Experience with Kubernetes, High-performance computing experience, System administration experience with BCM
Senior Site Reliability Engineer in Test, SDET
21 days agoPython · Kubernetes
8 years of experience, GitLab CI, ArgoCD, Kubernetes, Linux, Python, SRE principles, SBOM tooling, AI/ML techniques, test environment management, chaos engineering
Site Reliability Engineer
21 days agoAWS · Azure · Kubernetes
7 years of experience, AWS, Azure, Kubernetes, Cloud-native transformation, Cloud networking, Terraform, Security best practices
Site Reliability Engineer (Europe)
21 days agoGo · AWS · GCP
Golang, AWS, Google Cloud, Kubernetes, CI/CD pipelines, Prometheus, Grafana, ELK stack, Core networking, Security best practices, Automation tools
Cloud Infrastructure Engineer
22 days agoAzure
4 years of experience, Azure administration, PowerShell scripting, Microsoft Entra ID, DevSecOps collaboration, Cloud Operations experience, Troubleshooting complex infrastructure issues, Automation solutions development, Networking fundamentals
Senior Network Engineer - Operations
22 days agoPython
5 years of experience, Large Ethernet fabrics, Multi-Vendor & NOS (Juniper, Cisco, Arista), Scale-up vs scale-out architectures, Production networks at scale, Automation (Python/Go, Ansible), 100G+ environments, AI, GPU, or HPC exposure
Senior Support Engineer - Toronto
22 days agoPython
8 years of experience, API platform expertise, Automation in support operations, Advanced monitoring and alerting, Incident response leadership, Scripting (Python), Cloud infrastructure knowledge, Cross-functional communication
Senior Site Reliability Engineer
23 days agoTypeScript · Python · Kubernetes
5 years of experience, Kubernetes (EKS, GKE), Terraform, Pulumi, GitOps (ArgoCD, Flux), Observability stack (Prometheus, Grafana, Datadog), SLI/SLO/error budget fluency, Production-quality code in Go, Python, or TypeScript, Incident response leadership, Service mesh (Istio, Cilium), Live-service game experience
Lab Support Engineer 2, Hopkinton, MA
23 days agoPython
Python, Bash, PowerShell, Dell PowerEdge, PowerStore, Linux administration, VMware, vCenter, datacenter operations, networking fundamentals, problem-solving skills
Staff Platform Engineer, Core Cloud Platform
24 days agoPython · Go · AWS
Kubernetes expertise, AWS cloud proficiency, Production SaaS systems experience, Networking and service mesh knowledge, Operational troubleshooting skills, Python and Golang proficiency, CI/CD practices advocacy, Technical roadmap design and leadership
Cloud 2nd line ops Engineer
24 days agoAzure
Terraform, Azure DevOps, Azure cloud services, Azure resources, Azure RBAC, Azure Networking, Problem-Solving, Effective Communication, Documentation, Customer centric approach, Adaptability
Senior Cloud / DevSecOps Engineer (R-00200)
24 days agoPython · AWS · Kubernetes
AWS GovCloud, AWS CloudFormation, Terraform, Python, Bash, PowerShell, Ansible, CI/CD, Kubernetes, Amazon EKS, Amazon ECS, Cloud Security, DoD RMF, DISA STIGs, Infrastructure as Code, Hybrid Cloud Engineering, Windows Server, RHEL, Technical Leadership
Platform Infrastructure Engineer
24 days agoPython · Kubernetes
6 years of experience, OpenShift, Kubernetes, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Cluster monitoring and logging, Container image security, Service mesh (Istio, Linkerd), GitOps workflows (Argo CD, Flux)
Service Infrastructure Engineer
24 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, Service mesh production experience
WFH Sr. Network and Systems Administrator (Firewall) - #35234
25 days agoAzure
5 years of experience, Firewall management, Active Directory, Network security, VPN configuration, Enterprise-scale software deployments, Advanced PowerShell scripting, Hospitality industry experience, Microsoft Azure, Intrusion detection/prevention, Automation frameworks
Cloud Operations Engineer
26 days agoAWS · Azure
7 years of experience, Microsoft Azure, AWS, Linux administration, Office 365 migrations, Intune configuration, AWS GovCloud, Cloud security practices, System troubleshooting, Automation tools, Security tools knowledge, Analytical skills, Interpersonal communication, Team collaboration
Senior Staff Site Reliability Engineer
26 days agoPython · AWS · Kubernetes
10 years of experience, Kubernetes, AI/ML infrastructure, OpenTelemetry, AWS CDK, Terraform, Python, Go, Distributed systems debugging, Agentic AI platforms, Automation systems, Technical strategy steering
Systems Operations and Administrator
26 days ago5 years of experience, Linux administration, Enterprise networking standards, Data-center infrastructure, Networking knowledge (TCP/IP, DNS, DHCP), Scripting and automation, Infrastructure monitoring tools, Capacity planning, Operational process establishment
Sr. Staff Platform/Data Reliability Engineer, Databricks (R5537)
26 days ago12 years of experience, Databricks, CI/CD for data platforms, Observability, Incident management, Compute policy design, Regulated environment experience, Delta Lake, Infrastructure-as-code, Defense or aerospace industry experience
SRE L1 Support/Cloud Platform Ops Engineer
26 days ago2 years of experience, Basic Linux, Monitoring tools (Prometheus, Grafana, Nagios), ServiceNow/Jira, Physical data center tasks, Curiosity about automation, Structured data handling, Strong communication skills
Infrastructure Security Engineer
27 days agoAWS · Kubernetes
5 years of experience, AWS, Kubernetes, CI/CD security integrations, Terraform, Helm, Flux, ArgoCD, Okta, Web Application Firewalls, monitoring and logging platforms, security policies as code, incident response, active Secret clearance
IT Systems Engineer
28 days agoPython · AWS · GCP
3 years of experience, Python, PowerShell, Infrastructure as Code (IaC), CI/CD tools, AWS, GCP, Ansible, Terraform, Zabbix, Prometheus, Grafana, Analytical thinking, Bilingual (Polish and English)
Service Reliability Engineer
28 days agoPython · AWS · Azure
8 years of experience, Kubernetes, SLURM, large-scale cluster management, GPU hardware, high-performance computing, observability tools, incident management tools, AWS, Azure, GCP, OCI, Linux system administration, Ansible, Python, shell scripting, DNS, DHCP, storage systems, core networking, problem-solving, bare-metal infrastructure
Solutions Architect, Ethernet Networking - NVIS
28 days agoPython · Kubernetes · Docker
5 years of experience, Ethernet networking expertise, BGP, VxLAN, EVPN, Data center architecture, Network automation (Ansible, Salt, Python), Advanced network troubleshooting, Linux administration, Customer-facing experience, AI tools usage, Kubernetes, Docker, Networking simulation tools (NVIDIA Air, GNS3, EVE-NG), Network management tools (Grafana, Prometheus, Datadog)
Senior Site Reliability Engineer - Storage
28 days agoPython · Go · AWS
8 years of experience, HPC storage solutions, Enterprise NAS solutions, Distributed filesystems (Lustre, GPFS), Python/Bash/Golang, Cloud services (AWS, Azure, GCP), Monitoring stacks (Prometheus, Grafana, etc.), RDMA fabrics (InfiniBand, RoCE), HPC cluster management tools (Slurm, PBS, LSF), Containerization (Docker, Kubernetes)
Site Reliability Engineer
28 days agoAWS · Kubernetes
7 years of experience, Event-driven architecture, Deep AWS, Infrastructure as Code, Kubernetes, Observability and SLOs, Chaos engineering, Distributed systems debugging, Proven technical leadership, AI / MLOps infrastructure, Experience in payments industry
Site Reliability Engineer
28 days agoNode.js · Python · Java
Ansible, Terraform, Kubernetes, Linux, Windows, Python, Java, Golang, Node.js, Nginx, HAProxy, Docker, cloud-first mindset, security-first mindset