Remote Site Reliability Engineer jobs open to North America.

North America? Yes. $$$? When they say. Stack? Before the essay.

Can you be more specific?

66 openings. None older than 30 days.

Newest first · page 2

Site Reliability Engineer

15 days ago
salary not disclosedopen to WorldwideFull-time

Python · Azure · Kubernetes

3 years of experience, SLIs/SLOs definition, Multi-tenant SaaS platforms, Datadog, Grafana, Elastic Stack, High-availability architectures, Kubernetes, Python, Bash, Incident response, Cloud experience (Azure)

HostPapaHostPapa is a Canadian-based web hosting company offering B2B and B2C solutions, including shared, reseller, and VPS hosting services, with a focus on small businesses and a global presence.

Senior Site Reliability Engineer, BCM - DGX Cloud

15 days ago
$208,000 to $333,500 per yearopen to United StatesSeniorFull-time

Python · Kubernetes

8 years of experience, Fluency in Python, In-depth knowledge of Linux, Cluster networking proficiency, Experience with Kubernetes, High-performance computing experience, System administration experience with BCM

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Site Reliability Engineer

15 days ago
salary not disclosedopen to United StatesContract

AWS · Azure · Kubernetes

7 years of experience, AWS, Azure, Kubernetes, Cloud-native transformation, Cloud networking, Terraform, Security best practices

ASCENDINGASCENDING is a B2B technology services company specializing in AI-driven solutions and cloud transformation strategies, headquartered in an unspecified location, serving various industries including finance, healthcare, and education.

Cloud Infrastructure Engineer

16 days ago
$125,000 to $135,000 per yearopen to United StatesFull-time

Azure

4 years of experience, Azure administration, PowerShell scripting, Microsoft Entra ID, DevSecOps collaboration, Cloud Operations experience, Troubleshooting complex infrastructure issues, Automation solutions development, Networking fundamentals

AIP PublishingAIP Publishing is a Melville, NY-based B2B scholarly publisher specializing in peer-reviewed journals and resources in the physical sciences, serving a global community of researchers and academic institutions.

Senior Network Engineer - Operations

16 days ago
salary not disclosedopen to United StatesSeniorFull-time

Python

5 years of experience, Large Ethernet fabrics, Multi-Vendor & NOS (Juniper, Cisco, Arista), Scale-up vs scale-out architectures, Production networks at scale, Automation (Python/Go, Ansible), 100G+ environments, AI, GPU, or HPC exposure

TensorWaveTensorWave is a Las Vegas-based B2B cloud computing provider specializing in AI and high-performance computing infrastructure, utilizing AMD Instinct GPUs to deliver scalable solutions for enterprises and AI researchers.

Senior Support Engineer - Toronto

16 days ago
salary not disclosedopen to CanadaSeniorFull-time

Python

8 years of experience, API platform expertise, Automation in support operations, Advanced monitoring and alerting, Incident response leadership, Scripting (Python), Cloud infrastructure knowledge, Cross-functional communication

OpenAIOpenAI is a San Francisco-based AI research and deployment company specializing in generative AI models and cloud-based services, operating primarily in B2B markets while also reaching consumers through products like ChatGPT.

Senior Site Reliability Engineer

17 days ago
$93,700 to $138,700 per yearopen to CanadaSeniorFull-time

TypeScript · Python · Kubernetes

5 years of experience, Kubernetes (EKS, GKE), Terraform, Pulumi, GitOps (ArgoCD, Flux), Observability stack (Prometheus, Grafana, Datadog), SLI/SLO/error budget fluency, Production-quality code in Go, Python, or TypeScript, Incident response leadership, Service mesh (Istio, Cilium), Live-service game experience

2k2K is a Novato, California-based video game publisher specializing in a diverse range of genres including sports, action, and role-playing, primarily operating in the B2C market.

Lab Support Engineer 2, Hopkinton, MA

18 days ago
$83,045 to $107,470 per yearopen to United StatesFull-time

Python

Python, Bash, PowerShell, Dell PowerEdge, PowerStore, Linux administration, VMware, vCenter, datacenter operations, networking fundamentals, problem-solving skills

DellDell Technologies designs and develops hardware and software for infrastructure, enterprise solutions, and data management.

Associate Infrastructure Engineer

18 days ago
$158,300 to $190,000 per yearopen to United States + CanadaFull-time

Python · AWS · GCP

2 years of experience, Kubernetes, AWS or GCP, GitOps tools, Observability stack, Infrastructure-as-code tools, Node, Python or Go, Debugging distributed systems, On-call rotation participation, Proactive embrace of AI

WebflowWebflow is a San Francisco-based SaaS company providing a Website Experience Platform (WXP) that empowers B2B marketing teams to design, manage, and optimize custom websites for a global audience.

Staff Platform Engineer, Core Cloud Platform

18 days ago
$223,100 to $305,000 per yearopen to United States + Canada + United Kingdom + Singapore + India + Ireland + FinlandStaffFull-time

Python · Go · AWS

Kubernetes expertise, AWS cloud proficiency, Production SaaS systems experience, Networking and service mesh knowledge, Operational troubleshooting skills, Python and Golang proficiency, CI/CD practices advocacy, Technical roadmap design and leadership

AlphaSenseAlphaSense is a New York City-based B2B fintech platform specializing in AI-driven market intelligence and search solutions for financial institutions and top companies globally.

Senior Cloud / DevSecOps Engineer (R-00200)

18 days ago
salary not disclosedopen to United StatesSeniorFull-time

Python · AWS · Kubernetes

AWS GovCloud, AWS CloudFormation, Terraform, Python, Bash, PowerShell, Ansible, CI/CD, Kubernetes, Amazon EKS, Amazon ECS, Cloud Security, DoD RMF, DISA STIGs, Infrastructure as Code, Hybrid Cloud Engineering, Windows Server, RHEL, Technical Leadership

True Zero TechnologiesTrue Zero Technologies is a veteran-owned cybersecurity consulting firm headquartered in Fairfax, VA, specializing in B2B services for federal agencies and the public sector.

Platform Infrastructure Engineer

18 days ago
$100,000 to $150,000 per yearopen to United StatesFull-time

Python · Kubernetes

6 years of experience, OpenShift, Kubernetes, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Cluster monitoring and logging, Container image security, Service mesh (Istio, Linkerd), GitOps workflows (Argo CD, Flux)

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

Service Infrastructure Engineer

18 days ago
$100,000 to $150,000 per yearopen to United StatesFull-time

Python · Kubernetes

6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, Service mesh production experience

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

Cloud Operations Engineer

20 days ago
$110,000 to $127,000 per yearopen to United StatesFull-time

AWS · Azure

7 years of experience, Microsoft Azure, AWS, Linux administration, Office 365 migrations, Intune configuration, AWS GovCloud, Cloud security practices, System troubleshooting, Automation tools, Security tools knowledge, Analytical skills, Interpersonal communication, Team collaboration

CyberSheathCyberSheath is a managed security services provider specializing in cybersecurity compliance for the U.S. Defense Industrial Base (DIB), operating as a B2B entity focused on DoD contractors.

Systems Operations and Administrator

20 days ago
$112,000 to $178,250 per yearopen to United StatesFull-time

5 years of experience, Linux administration, Enterprise networking standards, Data-center infrastructure, Networking knowledge (TCP/IP, DNS, DHCP), Scripting and automation, Infrastructure monitoring tools, Capacity planning, Operational process establishment

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Sr. Staff Platform/Data Reliability Engineer, Databricks (R5537)

20 days ago
$180,000 to $270,000 per yearopen to United States + United Arab Emirates + UkraineStaffFull-time

12 years of experience, Databricks, CI/CD for data platforms, Observability, Incident management, Compute policy design, Regulated environment experience, Delta Lake, Infrastructure-as-code, Defense or aerospace industry experience

Shield AIShield AI is a San Diego-based defense technology company specializing in AI-powered autonomous drones and aircraft systems for military and commercial applications, operating primarily in the B2B sector.

SRE L1 Support/Cloud Platform Ops Engineer

20 days ago
$70,000 to $100,000 per yearopen to United StatesFull-time

2 years of experience, Basic Linux, Monitoring tools (Prometheus, Grafana, Nagios), ServiceNow/Jira, Physical data center tasks, Curiosity about automation, Structured data handling, Strong communication skills

Bitdeer GroupBitdeer Technologies Group is a Singapore-based B2B cryptocurrency mining and AI cloud solutions provider, specializing in cloud mining platforms and high-performance computing, with global operations including data centers in the U.S., Norway, and Bhutan.

Infrastructure Security Engineer

21 days ago
$140,000 to $160,000 per yearopen to United StatesFull-time

AWS · Kubernetes

5 years of experience, AWS, Kubernetes, CI/CD security integrations, Terraform, Helm, Flux, ArgoCD, Okta, Web Application Firewalls, monitoring and logging platforms, security policies as code, incident response, active Secret clearance

Sphinx DefenseSphinx Defense is a Washington, DC-based B2G SaaS provider specializing in advanced software solutions for satellite operations and national security, serving the U.S. Space Force and allied forces.

Service Reliability Engineer

22 days ago
$168,000 to $333,500 per yearopen to United StatesFull-time

Python · AWS · Azure

8 years of experience, Kubernetes, SLURM, large-scale cluster management, GPU hardware, high-performance computing, observability tools, incident management tools, AWS, Azure, GCP, OCI, Linux system administration, Ansible, Python, shell scripting, DNS, DHCP, storage systems, core networking, problem-solving, bare-metal infrastructure

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Solutions Architect, Ethernet Networking - NVIS

22 days ago
$124,000 to $195,500 per yearopen to North AmericaFull-time

Python · Kubernetes · Docker

5 years of experience, Ethernet networking expertise, BGP, VxLAN, EVPN, Data center architecture, Network automation (Ansible, Salt, Python), Advanced network troubleshooting, Linux administration, Customer-facing experience, AI tools usage, Kubernetes, Docker, Networking simulation tools (NVIDIA Air, GNS3, EVE-NG), Network management tools (Grafana, Prometheus, Datadog)

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Senior Site Reliability Engineer - Storage

22 days ago
$168,000 to $322,000 per yearopen to United StatesSeniorFull-time

Python · Go · AWS

8 years of experience, HPC storage solutions, Enterprise NAS solutions, Distributed filesystems (Lustre, GPFS), Python/Bash/Golang, Cloud services (AWS, Azure, GCP), Monitoring stacks (Prometheus, Grafana, etc.), RDMA fabrics (InfiniBand, RoCE), HPC cluster management tools (Slurm, PBS, LSF), Containerization (Docker, Kubernetes)

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Site Reliability Engineer

22 days ago
salary not disclosedopen to WorldwideFull-time

AWS · Kubernetes

7 years of experience, Event-driven architecture, Deep AWS, Infrastructure as Code, Kubernetes, Observability and SLOs, Chaos engineering, Distributed systems debugging, Proven technical leadership, AI / MLOps infrastructure, Experience in payments industry

YunoYuno is a global payment orchestration platform headquartered in [location], specializing in B2B payment infrastructure that enables businesses to integrate over 1,000 payment methods through a single API.

Site Reliability Engineer

22 days ago
salary not disclosedopen to United StatesFull-time

Node.js · Python · Java

Ansible, Terraform, Kubernetes, Linux, Windows, Python, Java, Golang, Node.js, Nginx, HAProxy, Docker, cloud-first mindset, security-first mindset

Ontrac SolutionsOntrac Solutions is a technology consulting firm specializing in AI-driven systems and digital transformation, headquartered remotely with a focus on B2B services for healthcare, retail, and wellness industries.

Cloud Engineer – Windows & Linux Platform Automation

22 days ago
salary not disclosedopen to United StatesFull-time

Python · SQL · AWS

4 years of experience, Microsoft platform management, Linux platform management, Public cloud services (Azure, AWS, GCP), Production networking, Production automation, SQL Server, PostgreSQL, MongoDB, Elasticsearch, Python, Terraform, PowerShell, Container scheduling engines (Mesos, Docker, Kubernetes), CI/CD solutions (GitHub Actions, Jenkins, CircleCI, ArgoCD), Level-3 support in ticketing systems

Ontrac SolutionsOntrac Solutions is a technology consulting firm specializing in AI-driven systems and digital transformation, headquartered remotely with a focus on B2B services for healthcare, retail, and wellness industries.

Reliability Monitoring Engineer

22 days ago
$100,000 to $150,000 per yearopen to United StatesFull-time

6 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, SLOs, High-cardinality metrics, CI/CD integration, Linux internals, Container platforms, Observability cost optimization

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

Systems Reliability Engineer

22 days ago
$100,000 to $150,000 per yearopen to United StatesFull-time

Python · Java · AWS

6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Prometheus, Grafana, CI/CD pipelines, Distributed system design, Incident response, SLOs and error budgets, Chaos engineering, AWS, Azure, GCP

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

Staff Platform Engineer

23 days ago
$250,000 to $285,000 per yearopen to United StatesStaffFull-time

Kubernetes

10 years of experience, Kubernetes, Infrastructure-as-code (Terraform), Database proficiency (Postgres), GitOps model (Argo), Security mindset, AI in engineering, Architectural judgment, Multi-tenant isolation design, Compliance-heavy deployments (FedRAMP), Excellent communication and collaboration

CortexCortex is a fully remote B2B software company specializing in an AI-powered Internal Developer Portal that enhances engineering productivity for teams across the US.

Cloud Engineer – Windows & Linux Platform Automation

23 days ago
salary not disclosedopen to United StatesFull-time

Python · SQL · AWS

4 years of experience, Microsoft platform management, Linux platform management, Public cloud services (Azure, AWS, GCP), Production networking, Configuration-management tools, SQL Server, PostgreSQL, MongoDB, Elasticsearch, Python, Terraform, PowerShell, Container scheduling engines (Mesos, Docker, Kubernetes), CI/CD solutions (GitHub Actions, Jenkins, CircleCI, ArgoCD), Level-3 support in ticketing systems

Ontrac SolutionsOntrac Solutions is a technology consulting firm specializing in AI-driven systems and digital transformation, headquartered remotely with a focus on B2B services for healthcare, retail, and wellness industries.

Production Engineer (IC4)

25 days ago
$100,000 to $130,000 per yearopen to United StatesFull-time

Python

3 years of experience, Python, RPM packaging, Enterprise OS modernization, Configuration management, CI/CD pipelines, Tier-2 operational support, Service onboarding, Monitoring and logging, Independent bug ownership

Ontrac SolutionsOntrac Solutions is a technology consulting firm specializing in AI-driven systems and digital transformation, headquartered remotely with a focus on B2B services for healthcare, retail, and wellness industries.

Infrastructure Engineer, APAC

26 days ago
$80,000 to $120,000 per yearopen to United StatesFull-time

Python · AWS · GCP

3 years of experience, Python, Terraform, AWS, GCP, Linux internals, Bash scripting, Ansible, Saltstack, Git, CI/CD pipelines, Containerization, Zabbix, Prometheus, Grafana, Opensearch, Elasticsearch

DTEX SystemsDTEX Systems is a cybersecurity B2B SaaS provider specializing in insider risk management, headquartered in an unspecified location, serving global enterprises and government sectors.