Remote Site Reliability Engineer jobs.
“Remote”? Depends where. $$$? When they say. Stack? Before the essay.
Where can you apply from?
102 openings. None older than 30 days.
Newest first
Cloud Infrastructure Engineer
about 8 hours agoAzure
4 years of experience, Azure administration, PowerShell scripting, Microsoft Entra ID, DevSecOps collaboration, Cloud Operations experience, Troubleshooting complex infrastructure issues, Automation solutions development, Networking fundamentals
Senior Network Engineer - Operations
about 8 hours agoPython
5 years of experience, Large Ethernet fabrics, Multi-Vendor & NOS (Juniper, Cisco, Arista), Scale-up vs scale-out architectures, Production networks at scale, Automation (Python/Go, Ansible), 100G+ environments, AI, GPU, or HPC exposure
Lead Site Reliability Engineer (Performance & Scalability) | Contract | Remote US
about 8 hours agoSite Reliability Engineering, Performance engineering, Capacity planning, Observability, Production distributed systems, SLOs and SLIs, Load and stress testing, Incident management, Graceful degradation design, Technical performance translation
Senior Support Engineer - Toronto
about 8 hours agoPython
8 years of experience, API platform expertise, Automation in support operations, Advanced monitoring and alerting, Incident response leadership, Scripting (Python), Cloud infrastructure knowledge, Cross-functional communication
Senior Site Reliability Engineer
1 day agoTypeScript · Python · Kubernetes
5 years of experience, Kubernetes (EKS, GKE), Terraform, Pulumi, GitOps (ArgoCD, Flux), Observability stack (Prometheus, Grafana, Datadog), SLI/SLO/error budget fluency, Production-quality code in Go, Python, or TypeScript, Incident response leadership, Service mesh (Istio, Cilium), Live-service game experience
Associate Infrastructure Engineer
2 days agoPython · AWS · GCP
2 years of experience, Kubernetes, AWS or GCP, GitOps tools, Observability stack, Infrastructure-as-code tools, Node, Python or Go, Debugging distributed systems, On-call rotation participation, Proactive embrace of AI
Staff Platform Engineer, Core Cloud Platform
2 days agoPython · Go · AWS
Kubernetes expertise, AWS cloud proficiency, Production SaaS systems experience, Networking and service mesh knowledge, Operational troubleshooting skills, Python and Golang proficiency, CI/CD practices advocacy, Technical roadmap design and leadership
Cloud 2nd line ops Engineer
2 days agoAzure
Terraform, Azure DevOps, Azure cloud services, Azure resources, Azure RBAC, Azure Networking, Problem-Solving, Effective Communication, Documentation, Customer centric approach, Adaptability
Senior Cloud / DevSecOps Engineer (R-00200)
2 days agoPython · AWS · Kubernetes
AWS GovCloud, AWS CloudFormation, Terraform, Python, Bash, PowerShell, Ansible, CI/CD, Kubernetes, Amazon EKS, Amazon ECS, Cloud Security, DoD RMF, DISA STIGs, Infrastructure as Code, Hybrid Cloud Engineering, Windows Server, RHEL, Technical Leadership
Cloud Networking Engineer
2 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, CNI, Ingress
DevOps & SRE Engineer
2 days agoPython · Java · Kubernetes
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Observability tooling, CI/CD pipelines, SLOs and error budgets, Chaos engineering, Cloud platforms
Platform Infrastructure Engineer
2 days agoPython · Kubernetes
6 years of experience, OpenShift, Kubernetes, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Cluster monitoring and logging, Container image security, Service mesh (Istio, Linkerd), GitOps workflows (Argo CD, Flux)
Service Infrastructure Engineer
2 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, Service mesh production experience
Telemetry Engineer
2 days agoPython · Java
6 years of experience, Prometheus, Grafana, OpenTelemetry, Datadog, New Relic, Splunk, Go, Python, Java, SRE principles, distributed tracing, structured logging, Linux, containers, CI/CD, observability cost optimization, eBPF-based observability
VMware Infrastructure Engineer
2 days agoAWS · Kubernetes
6 years of experience, vSphere, vSAN, NSX-T, PowerCLI, Tanzu Kubernetes Grid, disaster recovery, VMware Cloud on AWS, VMware Certified Professional (VCP)
Kubernetes & OpenShift Engineer
3 days agoPython · Kubernetes
6 years of experience, Kubernetes internals, OpenShift internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Cluster monitoring and logging, Container image security
Kubernetes Service Engineer
3 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, Observability
Platform Reliability Engineer
3 days agoPython · Java · AWS
6 years of experience, Python, Go, Java, Kubernetes, Linux, Prometheus, Grafana, CI/CD pipelines, SLOs, Chaos engineering, AWS, Azure, GCP
Systems Observability Specialist
3 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, distributed tracing, structured logging, Go, Python, Java, high-cardinality metrics, SLOs, error budgets, CI/CD integration, Linux internals, networking, container platforms
Cloud Operations Engineer
4 days agoAWS · Azure
7 years of experience, Microsoft Azure, AWS, Linux administration, Office 365 migrations, Intune configuration, AWS GovCloud, Cloud security practices, System troubleshooting, Automation tools, Security tools knowledge, Analytical skills, Interpersonal communication, Team collaboration
Senior Staff Site Reliability Engineer
4 days agoPython · AWS · Kubernetes
10 years of experience, Kubernetes, AI/ML infrastructure, OpenTelemetry, AWS CDK, Terraform, Python, Go, Distributed systems debugging, Agentic AI platforms, Automation systems, Technical strategy steering
Systems Operations and Administrator
4 days ago5 years of experience, Linux administration, Enterprise networking standards, Data-center infrastructure, Networking knowledge (TCP/IP, DNS, DHCP), Scripting and automation, Infrastructure monitoring tools, Capacity planning, Operational process establishment
Sr. Staff Platform/Data Reliability Engineer, Databricks (R5537)
4 days ago12 years of experience, Databricks, CI/CD for data platforms, Observability, Incident management, Compute policy design, Regulated environment experience, Delta Lake, Infrastructure-as-code, Defense or aerospace industry experience
Site Reliability Engineer (SRE)
4 days agoPython · AWS · Azure
10 years of experience, Kubernetes, Python, Go, Prometheus, Grafana, CI/CD pipelines, Chaos engineering, Distributed systems, SLOs and error budgets, Cloud platforms (AWS, Azure, GCP)
Azure Infrastructure Engineer
4 days agoPython · Azure · Kubernetes
6 years of experience, Azure core services, Infrastructure-as-code (Terraform, Bicep, ARM), Azure Kubernetes Service (AKS), Azure DevOps or GitHub Actions, Scripting (PowerShell, Bash, Python), Cloud security principles, Monitoring and observability strategies, Hybrid cloud or multi-cloud experience, FinOps practices, Regulated environments (HIPAA, PCI-DSS, SOC 2, FedRAMP)
Observability Engineer
4 days ago12 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, Distributed tracing, High-cardinality metrics, SLOs, CI/CD integration, Linux internals, Container platforms
SRE L1 Support/Cloud Platform Ops Engineer
4 days ago2 years of experience, Basic Linux, Monitoring tools (Prometheus, Grafana, Nagios), ServiceNow/Jira, Physical data center tasks, Curiosity about automation, Structured data handling, Strong communication skills
Infrastructure Security Engineer
5 days agoAWS · Kubernetes
5 years of experience, AWS, Kubernetes, CI/CD security integrations, Terraform, Helm, Flux, ArgoCD, Okta, Web Application Firewalls, monitoring and logging platforms, security policies as code, incident response, active Secret clearance
Apache Kafka Developer
5 days agoPython · Kubernetes
6 years of experience, Kafka internals, Kafka security, Kafka Connect, Schema Registry, Kafka Streams, HA/DR strategies, Python scripting, Terraform, Observability tooling, Confluent Certified Administrator, Kafka on Kubernetes
Monitoring Engineer
5 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, distributed tracing, structured logging, Go, Python, Java, high-cardinality metrics, SLOs, error budgets, SRE principles, CI/CD integration, Linux internals, networking, container platforms, Thanos, Mimir, Cortex, Loki, Tempo, eBPF-based tooling, observability cost optimization, regulated environments