Remote Site Reliability Engineer jobs open to North America.
North America? Yes. $$$? When they say. Stack? Before the essay.
Can you be more specific?
66 openings. None older than 30 days.
Newest first
Senior Site Reliability Engineer
4 days agoPython · Kubernetes
6 years of experience, Cloud-native infrastructure, Kubernetes deployment, Terraform, Ansible, CI/CD tooling, Distributed systems troubleshooting, Python, Bash, CUDA-based GPU programs, Security in sensitive environments
Senior Software Engineer, Site Reliability & Security
4 days agoTypeScript · Python · Go
5 years of experience, AWS or GCP, Kubernetes, Docker, Golang, Typescript, Python, Infrastructure as Code, GitOps, Grafana, Prometheus, UNIX shell
Sr. Production Engineer
4 days agoPython · AWS · Azure
4 years of experience, AWS, Azure, GCP, Python, Go, Prometheus, Grafana, OpenTelemetry, AI/ML understanding, Incident management, ITIL frameworks, Infrastructure-as-Code, Chaos engineering, BGP, GRE, IPSec, HAProxy, DNS
Senior Staff Site Reliability Engineer - Compute Core Engineering
4 days ago12 years of experience, eBPF, DPU, Containerization, Distributed Systems, Terraform, Linux Kernel Internals, Network Protocols, Microservices Architecture, Infrastructure as Code
DevOps & SRE Engineer
4 days agoPython · Java · AWS
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Prometheus, Grafana, CI/CD pipelines, SLOs, Chaos engineering, AWS, Azure, GCP, Service mesh technologies
Telemetry Engineer
4 days agoPython · Java
6 years of experience, Prometheus, Grafana, OpenTelemetry, Datadog, New Relic, Splunk, Go, Python, Java, SRE principles, incident management, distributed tracing, structured logging, Linux, containers, CI/CD, observability cost optimization, eBPF-based observability
VMware Infrastructure Engineer
4 days agoAWS · Kubernetes
6 years of experience, vSphere, vSAN, NSX-T, PowerCLI, Tanzu Kubernetes Grid, disaster recovery patterns, VMware Cloud on AWS, VMware Certified Professional (VCP)
Customer Reliability Engineer, Infrastructure
5 days agoPython · Kubernetes
5 years of experience, Kubernetes, Cloud infrastructure management, Production distributed systems, Linux, Customer issue handling, Strong communication, DevOps or CI/CD, Python scripting, Troubleshooting
Staff Site Reliability Engineer
5 days agoAWS
10 years of experience, AWS, Terraform, Puppet, Ansible, Linux administration, Containerization, PCI DSS compliance, Infrastructure as code, Scripting for automation, Technical leadership
Kubernetes Service Engineer
5 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, mTLS, Kubernetes, Go, Python, Distributed tracing, CNI, Zero-trust networking
Platform Reliability Engineer
5 days agoPython · Java · Kubernetes
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Observability tooling, CI/CD pipelines, Distributed system design, SLOs and error budgets, Chaos engineering, Cloud platforms
Site Observability Engineer
5 days ago6 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, Distributed tracing, High-cardinality metrics, SLOs, CI/CD integration, Linux internals, Container platforms
Systems Observability Specialist
5 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, distributed tracing, structured logging, Go, Python, Java, high-cardinality metrics, SLOs, error budgets, SRE principles, CI/CD integration, Linux internals, networking, container platforms, Thanos, Mimir, Cortex, Loki, Tempo, OpenTelemetry contributions, eBPF, cost optimization, regulated environments
Cloud Systems Engineer
6 days agoPython · AWS · Azure
5 years of experience, AWS, Azure, Microsoft 365, Terraform, PowerShell, Bash, Python, IAM, Security compliance, Linux, Windows Server, Defense sector experience, Automation of workflows, Documentation skills, AI platform administration
Staff Site Reliability Engineer
6 days agoPython · AWS · Kubernetes
7 years of experience, AWS, Kubernetes, Terraform, SLO frameworks, Incident response programs, Observability platforms, Python, Go, AI tools, Healthcare compliance (HIPAA), Mentorship
Kubernetes & OpenShift Engineer
6 days agoPython · AWS · Azure
6 years of experience, Kubernetes internals, OpenShift internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Cluster monitoring and logging, Container image security, Red Hat OpenShift experience, Public cloud experience (AWS, Azure, GCP)
Observability Engineer
6 days ago12 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, Distributed tracing, High-cardinality metrics, SLOs, CI/CD integration, Linux internals, Container platforms
Site Reliability Engineer (SRE)
6 days agoPython · AWS · Azure
10 years of experience, Kubernetes, Python, Go, Prometheus, Grafana, CI/CD pipelines, Chaos engineering, Distributed systems, SLOs and error budgets, Cloud platforms (AWS, Azure, GCP)
Azure Infrastructure Engineer
6 days agoPython · Azure · Kubernetes
6 years of experience, Azure core services, Infrastructure-as-code (Terraform, Bicep, ARM), Azure Kubernetes Service (AKS), Azure DevOps or GitHub Actions, Scripting (PowerShell, Bash, Python), Cloud security principles, Monitoring and observability strategies, Hybrid cloud or multi-cloud experience, FinOps practices, Regulated environments (HIPAA, PCI-DSS, SOC 2, FedRAMP)
Monitoring Engineer
7 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, distributed tracing, structured logging, Go, Python, Java, high-cardinality metrics, SLOs, error budgets, SRE principles, CI/CD integration, Linux internals, networking, container platforms
OpenShift Platform Engineer
7 days agoKubernetes
8 years of experience, OpenShift, Kubernetes internals, GitOps, CI/CD pipelines, Linux administration, Infrastructure-as-code, Container security, Disaster recovery strategies, Multi-tenant platform design, Monitoring and observability tools
Reliability Engineer
7 days agoPython · Java · Kubernetes
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Observability tooling, CI/CD pipelines, Distributed system design, SLOs and error budgets, Chaos engineering, Cloud platforms
Apache Kafka Developer
7 days agoKubernetes
6 years of experience, Kafka internals, Kafka security, Kafka Connect, Infrastructure-as-code, Observability tooling, Confluent Certified Administrator, Kafka on Kubernetes, Managed Kafka services, Stream processing frameworks, Data governance
Site Reliability Engineer Technical Lead
8 days agoPython · AWS · Azure
8 years of experience, Deep SRE and Systems Expertise, AWS, GCP, Azure, Kubernetes, Python, Automation and Tooling, Dynatrace, Splunk, ELK Stack, Observability and Analysis, Exceptional leadership, Communication skills
NCX Senior Engineer
9 days agoKubernetes
8 years of experience, GPU infrastructure management, NVIDIA technologies, Kubernetes expertise, Infrastructure observability tools, Automation for lifecycle management, Linux-based distributed systems, Collaboration with cloud partners, Failure modes in distributed AI workloads
Staff Production Operations Engineer
12 days agoAWS · Kubernetes
10 years of experience, AWS, Kubernetes, AI-assisted tools, Distributed service architecture, Incident management, Automation frameworks, Observability, CI/CD, Full-stack engineering, Production operations
Senior Software Engineer, Platform & Infrastructure - Riot Technology
13 days agoPython · AWS · GCP
3 years of experience, Kubernetes, AWS or GCP, Infrastructure-as-code, CI/CD, GPU compute infrastructure, Python, Distributed systems, MLOps workflows, Multi-node orchestration
Infrastructure Solutions Architect - OEM Deployment
13 days agoPython
2 years of experience, Large-scale datacenter rollouts, NVIDIA Cloud Partner integration, TCP/IP networking expertise, Bash scripting, Ansible, Python programming, GPU and DPU technologies, Signal integrity principles, Thermal management, Power distribution, Cabling and rack layout
Monitoring Platform Engineer (Mid-Level)
13 days ago3 years of experience, Splunk, DataDog, AppDynamics, Monitoring solutions, Attention to detail, Written communication
Monitoring Platform Engineer (Senior)
13 days ago6 years of experience, Splunk SOAR, Ansible, Enterprise monitoring tools, Secured environments, Strong troubleshooting skills, ITSM tools