Remote Site Reliability Engineer jobs open to North America.
North America? Yes. $$$? When they say. Stack? Before the essay.
Can you be more specific?
66 openings. None older than 30 days.
Newest first · page 2
Site Reliability Engineer
15 days agoPython · Azure · Kubernetes
3 years of experience, SLIs/SLOs definition, Multi-tenant SaaS platforms, Datadog, Grafana, Elastic Stack, High-availability architectures, Kubernetes, Python, Bash, Incident response, Cloud experience (Azure)
Senior Site Reliability Engineer, BCM - DGX Cloud
15 days agoPython · Kubernetes
8 years of experience, Fluency in Python, In-depth knowledge of Linux, Cluster networking proficiency, Experience with Kubernetes, High-performance computing experience, System administration experience with BCM
Site Reliability Engineer
15 days agoAWS · Azure · Kubernetes
7 years of experience, AWS, Azure, Kubernetes, Cloud-native transformation, Cloud networking, Terraform, Security best practices
Cloud Infrastructure Engineer
16 days agoAzure
4 years of experience, Azure administration, PowerShell scripting, Microsoft Entra ID, DevSecOps collaboration, Cloud Operations experience, Troubleshooting complex infrastructure issues, Automation solutions development, Networking fundamentals
Senior Network Engineer - Operations
16 days agoPython
5 years of experience, Large Ethernet fabrics, Multi-Vendor & NOS (Juniper, Cisco, Arista), Scale-up vs scale-out architectures, Production networks at scale, Automation (Python/Go, Ansible), 100G+ environments, AI, GPU, or HPC exposure
Senior Support Engineer - Toronto
16 days agoPython
8 years of experience, API platform expertise, Automation in support operations, Advanced monitoring and alerting, Incident response leadership, Scripting (Python), Cloud infrastructure knowledge, Cross-functional communication
Senior Site Reliability Engineer
17 days agoTypeScript · Python · Kubernetes
5 years of experience, Kubernetes (EKS, GKE), Terraform, Pulumi, GitOps (ArgoCD, Flux), Observability stack (Prometheus, Grafana, Datadog), SLI/SLO/error budget fluency, Production-quality code in Go, Python, or TypeScript, Incident response leadership, Service mesh (Istio, Cilium), Live-service game experience
Lab Support Engineer 2, Hopkinton, MA
18 days agoPython
Python, Bash, PowerShell, Dell PowerEdge, PowerStore, Linux administration, VMware, vCenter, datacenter operations, networking fundamentals, problem-solving skills
Associate Infrastructure Engineer
18 days agoPython · AWS · GCP
2 years of experience, Kubernetes, AWS or GCP, GitOps tools, Observability stack, Infrastructure-as-code tools, Node, Python or Go, Debugging distributed systems, On-call rotation participation, Proactive embrace of AI
Staff Platform Engineer, Core Cloud Platform
18 days agoPython · Go · AWS
Kubernetes expertise, AWS cloud proficiency, Production SaaS systems experience, Networking and service mesh knowledge, Operational troubleshooting skills, Python and Golang proficiency, CI/CD practices advocacy, Technical roadmap design and leadership
Senior Cloud / DevSecOps Engineer (R-00200)
18 days agoPython · AWS · Kubernetes
AWS GovCloud, AWS CloudFormation, Terraform, Python, Bash, PowerShell, Ansible, CI/CD, Kubernetes, Amazon EKS, Amazon ECS, Cloud Security, DoD RMF, DISA STIGs, Infrastructure as Code, Hybrid Cloud Engineering, Windows Server, RHEL, Technical Leadership
Platform Infrastructure Engineer
18 days agoPython · Kubernetes
6 years of experience, OpenShift, Kubernetes, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Cluster monitoring and logging, Container image security, Service mesh (Istio, Linkerd), GitOps workflows (Argo CD, Flux)
Service Infrastructure Engineer
18 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, Service mesh production experience
Cloud Operations Engineer
20 days agoAWS · Azure
7 years of experience, Microsoft Azure, AWS, Linux administration, Office 365 migrations, Intune configuration, AWS GovCloud, Cloud security practices, System troubleshooting, Automation tools, Security tools knowledge, Analytical skills, Interpersonal communication, Team collaboration
Systems Operations and Administrator
20 days ago5 years of experience, Linux administration, Enterprise networking standards, Data-center infrastructure, Networking knowledge (TCP/IP, DNS, DHCP), Scripting and automation, Infrastructure monitoring tools, Capacity planning, Operational process establishment
Sr. Staff Platform/Data Reliability Engineer, Databricks (R5537)
20 days ago12 years of experience, Databricks, CI/CD for data platforms, Observability, Incident management, Compute policy design, Regulated environment experience, Delta Lake, Infrastructure-as-code, Defense or aerospace industry experience
SRE L1 Support/Cloud Platform Ops Engineer
20 days ago2 years of experience, Basic Linux, Monitoring tools (Prometheus, Grafana, Nagios), ServiceNow/Jira, Physical data center tasks, Curiosity about automation, Structured data handling, Strong communication skills
Infrastructure Security Engineer
21 days agoAWS · Kubernetes
5 years of experience, AWS, Kubernetes, CI/CD security integrations, Terraform, Helm, Flux, ArgoCD, Okta, Web Application Firewalls, monitoring and logging platforms, security policies as code, incident response, active Secret clearance
Service Reliability Engineer
22 days agoPython · AWS · Azure
8 years of experience, Kubernetes, SLURM, large-scale cluster management, GPU hardware, high-performance computing, observability tools, incident management tools, AWS, Azure, GCP, OCI, Linux system administration, Ansible, Python, shell scripting, DNS, DHCP, storage systems, core networking, problem-solving, bare-metal infrastructure
Solutions Architect, Ethernet Networking - NVIS
22 days agoPython · Kubernetes · Docker
5 years of experience, Ethernet networking expertise, BGP, VxLAN, EVPN, Data center architecture, Network automation (Ansible, Salt, Python), Advanced network troubleshooting, Linux administration, Customer-facing experience, AI tools usage, Kubernetes, Docker, Networking simulation tools (NVIDIA Air, GNS3, EVE-NG), Network management tools (Grafana, Prometheus, Datadog)
Senior Site Reliability Engineer - Storage
22 days agoPython · Go · AWS
8 years of experience, HPC storage solutions, Enterprise NAS solutions, Distributed filesystems (Lustre, GPFS), Python/Bash/Golang, Cloud services (AWS, Azure, GCP), Monitoring stacks (Prometheus, Grafana, etc.), RDMA fabrics (InfiniBand, RoCE), HPC cluster management tools (Slurm, PBS, LSF), Containerization (Docker, Kubernetes)
Site Reliability Engineer
22 days agoAWS · Kubernetes
7 years of experience, Event-driven architecture, Deep AWS, Infrastructure as Code, Kubernetes, Observability and SLOs, Chaos engineering, Distributed systems debugging, Proven technical leadership, AI / MLOps infrastructure, Experience in payments industry
Site Reliability Engineer
22 days agoNode.js · Python · Java
Ansible, Terraform, Kubernetes, Linux, Windows, Python, Java, Golang, Node.js, Nginx, HAProxy, Docker, cloud-first mindset, security-first mindset
Cloud Engineer – Windows & Linux Platform Automation
22 days agoPython · SQL · AWS
4 years of experience, Microsoft platform management, Linux platform management, Public cloud services (Azure, AWS, GCP), Production networking, Production automation, SQL Server, PostgreSQL, MongoDB, Elasticsearch, Python, Terraform, PowerShell, Container scheduling engines (Mesos, Docker, Kubernetes), CI/CD solutions (GitHub Actions, Jenkins, CircleCI, ArgoCD), Level-3 support in ticketing systems
Reliability Monitoring Engineer
22 days ago6 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, SLOs, High-cardinality metrics, CI/CD integration, Linux internals, Container platforms, Observability cost optimization
Systems Reliability Engineer
22 days agoPython · Java · AWS
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Prometheus, Grafana, CI/CD pipelines, Distributed system design, Incident response, SLOs and error budgets, Chaos engineering, AWS, Azure, GCP
Staff Platform Engineer
23 days agoKubernetes
10 years of experience, Kubernetes, Infrastructure-as-code (Terraform), Database proficiency (Postgres), GitOps model (Argo), Security mindset, AI in engineering, Architectural judgment, Multi-tenant isolation design, Compliance-heavy deployments (FedRAMP), Excellent communication and collaboration
Cloud Engineer – Windows & Linux Platform Automation
23 days agoPython · SQL · AWS
4 years of experience, Microsoft platform management, Linux platform management, Public cloud services (Azure, AWS, GCP), Production networking, Configuration-management tools, SQL Server, PostgreSQL, MongoDB, Elasticsearch, Python, Terraform, PowerShell, Container scheduling engines (Mesos, Docker, Kubernetes), CI/CD solutions (GitHub Actions, Jenkins, CircleCI, ArgoCD), Level-3 support in ticketing systems
Production Engineer (IC4)
25 days agoPython
3 years of experience, Python, RPM packaging, Enterprise OS modernization, Configuration management, CI/CD pipelines, Tier-2 operational support, Service onboarding, Monitoring and logging, Independent bug ownership
Infrastructure Engineer, APAC
26 days agoPython · AWS · GCP
3 years of experience, Python, Terraform, AWS, GCP, Linux internals, Bash scripting, Ansible, Saltstack, Git, CI/CD pipelines, Containerization, Zabbix, Prometheus, Grafana, Opensearch, Elasticsearch