Remote Site Reliability Engineer jobs.
“Remote”? Depends where. $$$? When they say. Stack? Before the essay.
Where can you apply from?
106 openings. None older than 30 days.
Newest first
Senior Site Reliability Engineer, DGX Cloud
1 day agoPython · Kubernetes
10 years of experience, Kubernetes administration, GPU workloads optimization, Infrastructure automation (Terraform, Ansible), High-level programming (Python, Go), Linux operating systems, SRE principles (SLOs, SLIs), Observability stacks (OpenTelemetry, Prometheus), GPU-accelerated clusters with KubeVirt, Generative-AI techniques, Workflow orchestration (Temporal, Airflow)
Senior Site Reliability Engineer
2 days agoTypeScript · Python · Kubernetes
5 years of experience, Kubernetes (EKS, GKE), Terraform, Pulumi, GitOps (ArgoCD, Flux), Datadog, Prometheus, Grafana, SLI/SLO/error budget fluency, Go, Python, TypeScript, Linux internals, TCP/IP networking, Incident response leadership, Live-service game experience, Service mesh (Istio, Cilium), FinOps, Cloud certifications
Observability Engineer
2 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, distributed tracing, structured logging, Go, Python, Java, high-cardinality metrics, SLOs, error budgets, SRE principles, CI/CD integration, Linux internals, networking, container platforms, Thanos, Mimir, Cortex, Loki, Tempo, eBPF, cost optimization, regulated environments
OpenShift Platform Engineer
2 days agoPython · Kubernetes
6 years of experience, OpenShift, Kubernetes internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Container image security, Cluster monitoring tools, Service mesh (OpenShift Service Mesh, Istio, Linkerd), GitOps workflows (Argo CD, Flux)
Service Mesh Engineer
2 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, Observability
Site Reliability Engineer
2 days agoPython · Azure · Kubernetes
3 years of experience, SLIs/SLOs definition, Multi-tenant SaaS platforms, Datadog, Grafana, Elastic Stack, High-availability architectures, Kubernetes, Python, Bash, Incident response, Cloud experience (Azure)
Site Reliability Engineer
2 days agoGo · AWS · GCP
Kubernetes, AWS, Google Cloud, Golang, CI/CD optimization, Distributed systems management, Monitoring tools (Prometheus, Grafana), Troubleshooting complex infrastructure issues, Disaster recovery strategies, Cloud security practices
Senior Site Reliability Engineer, BCM - DGX Cloud
2 days agoPython · Kubernetes
8 years of experience, Fluency in Python, In-depth knowledge of Linux, Cluster networking proficiency, Experience with Kubernetes, High-performance computing experience, System administration experience with BCM
Senior Site Reliability Engineer in Test, SDET
2 days agoPython · Kubernetes
8 years of experience, GitLab CI, ArgoCD, Kubernetes, Linux, Python, SRE principles, SBOM tooling, AI/ML techniques, test environment management, chaos engineering
Site Reliability Engineer
2 days agoAWS · Azure · Kubernetes
7 years of experience, AWS, Azure, Kubernetes, Cloud-native transformation, Cloud networking, Terraform, Security best practices
Cloud Infrastructure Engineer
3 days agoAzure
4 years of experience, Azure administration, PowerShell scripting, Microsoft Entra ID, DevSecOps collaboration, Cloud Operations experience, Troubleshooting complex infrastructure issues, Automation solutions development, Networking fundamentals
Senior Network Engineer - Operations
3 days agoPython
5 years of experience, Large Ethernet fabrics, Multi-Vendor & NOS (Juniper, Cisco, Arista), Scale-up vs scale-out architectures, Production networks at scale, Automation (Python/Go, Ansible), 100G+ environments, AI, GPU, or HPC exposure
Lead Site Reliability Engineer (Performance & Scalability) | Contract | Remote US
3 days agoSite Reliability Engineering, Performance engineering, Capacity planning, Observability, Production distributed systems, SLOs and SLIs, Load and stress testing, Incident management, Graceful degradation design, Technical performance translation
Senior Support Engineer - Toronto
3 days agoPython
8 years of experience, API platform expertise, Automation in support operations, Advanced monitoring and alerting, Incident response leadership, Scripting (Python), Cloud infrastructure knowledge, Cross-functional communication
Senior Site Reliability Engineer
4 days agoTypeScript · Python · Kubernetes
5 years of experience, Kubernetes (EKS, GKE), Terraform, Pulumi, GitOps (ArgoCD, Flux), Observability stack (Prometheus, Grafana, Datadog), SLI/SLO/error budget fluency, Production-quality code in Go, Python, or TypeScript, Incident response leadership, Service mesh (Istio, Cilium), Live-service game experience
Lab Support Engineer 2, Hopkinton, MA
5 days agoPython
Python, Bash, PowerShell, Dell PowerEdge, PowerStore, Linux administration, VMware, vCenter, datacenter operations, networking fundamentals, problem-solving skills
Associate Infrastructure Engineer
5 days agoPython · AWS · GCP
2 years of experience, Kubernetes, AWS or GCP, GitOps tools, Observability stack, Infrastructure-as-code tools, Node, Python or Go, Debugging distributed systems, On-call rotation participation, Proactive embrace of AI
Staff Platform Engineer, Core Cloud Platform
5 days agoPython · Go · AWS
Kubernetes expertise, AWS cloud proficiency, Production SaaS systems experience, Networking and service mesh knowledge, Operational troubleshooting skills, Python and Golang proficiency, CI/CD practices advocacy, Technical roadmap design and leadership
Cloud 2nd line ops Engineer
5 days agoAzure
Terraform, Azure DevOps, Azure cloud services, Azure resources, Azure RBAC, Azure Networking, Problem-Solving, Effective Communication, Documentation, Customer centric approach, Adaptability
Senior Cloud / DevSecOps Engineer (R-00200)
5 days agoPython · AWS · Kubernetes
AWS GovCloud, AWS CloudFormation, Terraform, Python, Bash, PowerShell, Ansible, CI/CD, Kubernetes, Amazon EKS, Amazon ECS, Cloud Security, DoD RMF, DISA STIGs, Infrastructure as Code, Hybrid Cloud Engineering, Windows Server, RHEL, Technical Leadership
Cloud Networking Engineer
5 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, CNI, Ingress
DevOps & SRE Engineer
5 days agoPython · Java · Kubernetes
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Observability tooling, CI/CD pipelines, SLOs and error budgets, Chaos engineering, Cloud platforms
Platform Infrastructure Engineer
5 days agoPython · Kubernetes
6 years of experience, OpenShift, Kubernetes, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Cluster monitoring and logging, Container image security, Service mesh (Istio, Linkerd), GitOps workflows (Argo CD, Flux)
Service Infrastructure Engineer
5 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, Service mesh production experience
Telemetry Engineer
5 days agoPython · Java
6 years of experience, Prometheus, Grafana, OpenTelemetry, Datadog, New Relic, Splunk, Go, Python, Java, SRE principles, distributed tracing, structured logging, Linux, containers, CI/CD, observability cost optimization, eBPF-based observability
VMware Infrastructure Engineer
5 days agoAWS · Kubernetes
6 years of experience, vSphere, vSAN, NSX-T, PowerCLI, Tanzu Kubernetes Grid, disaster recovery, VMware Cloud on AWS, VMware Certified Professional (VCP)
Kubernetes & OpenShift Engineer
6 days agoPython · Kubernetes
6 years of experience, Kubernetes internals, OpenShift internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Cluster monitoring and logging, Container image security
Kubernetes Service Engineer
6 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, Observability
Platform Reliability Engineer
6 days agoPython · Java · AWS
6 years of experience, Python, Go, Java, Kubernetes, Linux, Prometheus, Grafana, CI/CD pipelines, SLOs, Chaos engineering, AWS, Azure, GCP
Systems Observability Specialist
6 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, distributed tracing, structured logging, Go, Python, Java, high-cardinality metrics, SLOs, error budgets, CI/CD integration, Linux internals, networking, container platforms