Remote Site Reliability Engineer jobs.

“Remote”? Depends where. $$$? When they say. Stack? Before the essay.

Where can you apply from?

98 openings. None older than 30 days.

Newest first · page 3

Senior Site Reliability Engineer

21 days ago
salary not disclosedopen to IndiaSeniorFull-time

TypeScript · Python · Kubernetes

5 years of experience, Kubernetes (EKS, GKE), Terraform, Pulumi, GitOps (ArgoCD, Flux), Datadog, Prometheus, Grafana, SLI/SLO/error budget fluency, Go, Python, TypeScript, Linux internals, TCP/IP networking, Incident response leadership, Live-service game experience, Service mesh (Istio, Cilium), FinOps, Cloud certifications

2k2K is a Novato, California-based video game publisher specializing in a diverse range of genres including sports, action, and role-playing, primarily operating in the B2C market.

Site Reliability Engineer

21 days ago
salary not disclosedopen to WorldwideFull-time

Python · Azure · Kubernetes

3 years of experience, SLIs/SLOs definition, Multi-tenant SaaS platforms, Datadog, Grafana, Elastic Stack, High-availability architectures, Kubernetes, Python, Bash, Incident response, Cloud experience (Azure)

HostPapaHostPapa is a Canadian-based web hosting company offering B2B and B2C solutions, including shared, reseller, and VPS hosting services, with a focus on small businesses and a global presence.

Site Reliability Engineer

21 days ago
€50,000 to €80,000 per yearopen to Spain + PortugalFull-time

Go · AWS · GCP

Kubernetes, AWS, Google Cloud, Golang, CI/CD optimization, Distributed systems management, Monitoring tools (Prometheus, Grafana), Troubleshooting complex infrastructure issues, Disaster recovery strategies, Cloud security practices

ArangoDBArangoDB is a San Francisco-based B2B multi-model database platform that integrates graph, document, and key-value data models, serving industries like AI, healthcare, and finance globally.

Senior Site Reliability Engineer, BCM - DGX Cloud

21 days ago
$208,000 to $333,500 per yearopen to United StatesSeniorFull-time

Python · Kubernetes

8 years of experience, Fluency in Python, In-depth knowledge of Linux, Cluster networking proficiency, Experience with Kubernetes, High-performance computing experience, System administration experience with BCM

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Senior Site Reliability Engineer in Test, SDET

21 days ago
salary not disclosedopen to ChinaSeniorFull-time

Python · Kubernetes

8 years of experience, GitLab CI, ArgoCD, Kubernetes, Linux, Python, SRE principles, SBOM tooling, AI/ML techniques, test environment management, chaos engineering

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Site Reliability Engineer

21 days ago
salary not disclosedopen to United StatesContract

AWS · Azure · Kubernetes

7 years of experience, AWS, Azure, Kubernetes, Cloud-native transformation, Cloud networking, Terraform, Security best practices

ASCENDINGASCENDING is a B2B technology services company specializing in AI-driven solutions and cloud transformation strategies, headquartered in an unspecified location, serving various industries including finance, healthcare, and education.

Site Reliability Engineer (Europe)

21 days ago
salary not disclosedopen to EuropeFull-time

Go · AWS · GCP

Golang, AWS, Google Cloud, Kubernetes, CI/CD pipelines, Prometheus, Grafana, ELK stack, Core networking, Security best practices, Automation tools

ArangoDBArangoDB is a San Francisco-based B2B multi-model database platform that integrates graph, document, and key-value data models, serving industries like AI, healthcare, and finance globally.

Cloud Infrastructure Engineer

22 days ago
$125,000 to $135,000 per yearopen to United StatesFull-time

Azure

4 years of experience, Azure administration, PowerShell scripting, Microsoft Entra ID, DevSecOps collaboration, Cloud Operations experience, Troubleshooting complex infrastructure issues, Automation solutions development, Networking fundamentals

AIP PublishingAIP Publishing is a Melville, NY-based B2B scholarly publisher specializing in peer-reviewed journals and resources in the physical sciences, serving a global community of researchers and academic institutions.

Senior Network Engineer - Operations

22 days ago
salary not disclosedopen to United StatesSeniorFull-time

Python

5 years of experience, Large Ethernet fabrics, Multi-Vendor & NOS (Juniper, Cisco, Arista), Scale-up vs scale-out architectures, Production networks at scale, Automation (Python/Go, Ansible), 100G+ environments, AI, GPU, or HPC exposure

TensorWaveTensorWave is a Las Vegas-based B2B cloud computing provider specializing in AI and high-performance computing infrastructure, utilizing AMD Instinct GPUs to deliver scalable solutions for enterprises and AI researchers.

Senior Support Engineer - Toronto

22 days ago
salary not disclosedopen to CanadaSeniorFull-time

Python

8 years of experience, API platform expertise, Automation in support operations, Advanced monitoring and alerting, Incident response leadership, Scripting (Python), Cloud infrastructure knowledge, Cross-functional communication

OpenAIOpenAI is a San Francisco-based AI research and deployment company specializing in generative AI models and cloud-based services, operating primarily in B2B markets while also reaching consumers through products like ChatGPT.

Senior Site Reliability Engineer

23 days ago
$93,700 to $138,700 per yearopen to CanadaSeniorFull-time

TypeScript · Python · Kubernetes

5 years of experience, Kubernetes (EKS, GKE), Terraform, Pulumi, GitOps (ArgoCD, Flux), Observability stack (Prometheus, Grafana, Datadog), SLI/SLO/error budget fluency, Production-quality code in Go, Python, or TypeScript, Incident response leadership, Service mesh (Istio, Cilium), Live-service game experience

2k2K is a Novato, California-based video game publisher specializing in a diverse range of genres including sports, action, and role-playing, primarily operating in the B2C market.

Lab Support Engineer 2, Hopkinton, MA

23 days ago
$83,045 to $107,470 per yearopen to United StatesFull-time

Python

Python, Bash, PowerShell, Dell PowerEdge, PowerStore, Linux administration, VMware, vCenter, datacenter operations, networking fundamentals, problem-solving skills

DellDell Technologies designs and develops hardware and software for infrastructure, enterprise solutions, and data management.

Staff Platform Engineer, Core Cloud Platform

24 days ago
$223,100 to $305,000 per yearopen to United States + Canada + United Kingdom + Singapore + India + Ireland + FinlandStaffFull-time

Python · Go · AWS

Kubernetes expertise, AWS cloud proficiency, Production SaaS systems experience, Networking and service mesh knowledge, Operational troubleshooting skills, Python and Golang proficiency, CI/CD practices advocacy, Technical roadmap design and leadership

AlphaSenseAlphaSense is a New York City-based B2B fintech platform specializing in AI-driven market intelligence and search solutions for financial institutions and top companies globally.

Cloud 2nd line ops Engineer

24 days ago
€43,200 to €45,600 per yearopen to PortugalFull-time

Azure

Terraform, Azure DevOps, Azure cloud services, Azure resources, Azure RBAC, Azure Networking, Problem-Solving, Effective Communication, Documentation, Customer centric approach, Adaptability

Irium PortugalIrium Portugal is a Lisbon-based B2B IT services provider specializing in managed services, digital transformation, and cybersecurity solutions for businesses across Portugal.

Senior Cloud / DevSecOps Engineer (R-00200)

24 days ago
salary not disclosedopen to United StatesSeniorFull-time

Python · AWS · Kubernetes

AWS GovCloud, AWS CloudFormation, Terraform, Python, Bash, PowerShell, Ansible, CI/CD, Kubernetes, Amazon EKS, Amazon ECS, Cloud Security, DoD RMF, DISA STIGs, Infrastructure as Code, Hybrid Cloud Engineering, Windows Server, RHEL, Technical Leadership

True Zero TechnologiesTrue Zero Technologies is a veteran-owned cybersecurity consulting firm headquartered in Fairfax, VA, specializing in B2B services for federal agencies and the public sector.

Platform Infrastructure Engineer

24 days ago
$100,000 to $150,000 per yearopen to United StatesFull-time

Python · Kubernetes

6 years of experience, OpenShift, Kubernetes, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Cluster monitoring and logging, Container image security, Service mesh (Istio, Linkerd), GitOps workflows (Argo CD, Flux)

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

Service Infrastructure Engineer

24 days ago
$100,000 to $150,000 per yearopen to United StatesFull-time

Python · Kubernetes

6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, Service mesh production experience

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

WFH Sr. Network and Systems Administrator (Firewall) - #35234

25 days ago
salary not disclosedopen to PhilippinesSeniorFull-time

Azure

5 years of experience, Firewall management, Active Directory, Network security, VPN configuration, Enterprise-scale software deployments, Advanced PowerShell scripting, Hospitality industry experience, Microsoft Azure, Intrusion detection/prevention, Automation frameworks

Manila RecruitmentManila Recruitment is a Philippines-based recruitment agency specializing in staffing and job placement services for the healthcare and finance sectors, operating primarily in the local B2B and B2C markets.

Cloud Operations Engineer

26 days ago
$110,000 to $127,000 per yearopen to United StatesFull-time

AWS · Azure

7 years of experience, Microsoft Azure, AWS, Linux administration, Office 365 migrations, Intune configuration, AWS GovCloud, Cloud security practices, System troubleshooting, Automation tools, Security tools knowledge, Analytical skills, Interpersonal communication, Team collaboration

CyberSheathCyberSheath is a managed security services provider specializing in cybersecurity compliance for the U.S. Defense Industrial Base (DIB), operating as a B2B entity focused on DoD contractors.

Senior Staff Site Reliability Engineer

26 days ago
$150,000 to $200,000 per yearopen to IndiaStaffFull-time

Python · AWS · Kubernetes

10 years of experience, Kubernetes, AI/ML infrastructure, OpenTelemetry, AWS CDK, Terraform, Python, Go, Distributed systems debugging, Agentic AI platforms, Automation systems, Technical strategy steering

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Systems Operations and Administrator

26 days ago
$112,000 to $178,250 per yearopen to United StatesFull-time

5 years of experience, Linux administration, Enterprise networking standards, Data-center infrastructure, Networking knowledge (TCP/IP, DNS, DHCP), Scripting and automation, Infrastructure monitoring tools, Capacity planning, Operational process establishment

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Sr. Staff Platform/Data Reliability Engineer, Databricks (R5537)

26 days ago
$180,000 to $270,000 per yearopen to United States + United Arab Emirates + UkraineStaffFull-time

12 years of experience, Databricks, CI/CD for data platforms, Observability, Incident management, Compute policy design, Regulated environment experience, Delta Lake, Infrastructure-as-code, Defense or aerospace industry experience

Shield AIShield AI is a San Diego-based defense technology company specializing in AI-powered autonomous drones and aircraft systems for military and commercial applications, operating primarily in the B2B sector.

SRE L1 Support/Cloud Platform Ops Engineer

26 days ago
$70,000 to $100,000 per yearopen to United StatesFull-time

2 years of experience, Basic Linux, Monitoring tools (Prometheus, Grafana, Nagios), ServiceNow/Jira, Physical data center tasks, Curiosity about automation, Structured data handling, Strong communication skills

Bitdeer GroupBitdeer Technologies Group is a Singapore-based B2B cryptocurrency mining and AI cloud solutions provider, specializing in cloud mining platforms and high-performance computing, with global operations including data centers in the U.S., Norway, and Bhutan.

Infrastructure Security Engineer

27 days ago
$140,000 to $160,000 per yearopen to United StatesFull-time

AWS · Kubernetes

5 years of experience, AWS, Kubernetes, CI/CD security integrations, Terraform, Helm, Flux, ArgoCD, Okta, Web Application Firewalls, monitoring and logging platforms, security policies as code, incident response, active Secret clearance

Sphinx DefenseSphinx Defense is a Washington, DC-based B2G SaaS provider specializing in advanced software solutions for satellite operations and national security, serving the U.S. Space Force and allied forces.

IT Systems Engineer

28 days ago
12,900 to 17,700 PLN per monthopen to PolandFull-time

Python · AWS · GCP

3 years of experience, Python, PowerShell, Infrastructure as Code (IaC), CI/CD tools, AWS, GCP, Ansible, Terraform, Zabbix, Prometheus, Grafana, Analytical thinking, Bilingual (Polish and English)

CD PROJEKT REDCD PROJEKT RED is a Warsaw-based video game developer specializing in story-driven RPGs, including The Witcher series and Cyberpunk 2077, operating primarily in the B2C gaming industry with a global target market.

Service Reliability Engineer

28 days ago
$168,000 to $333,500 per yearopen to United StatesFull-time

Python · AWS · Azure

8 years of experience, Kubernetes, SLURM, large-scale cluster management, GPU hardware, high-performance computing, observability tools, incident management tools, AWS, Azure, GCP, OCI, Linux system administration, Ansible, Python, shell scripting, DNS, DHCP, storage systems, core networking, problem-solving, bare-metal infrastructure

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Solutions Architect, Ethernet Networking - NVIS

28 days ago
$124,000 to $195,500 per yearopen to North AmericaFull-time

Python · Kubernetes · Docker

5 years of experience, Ethernet networking expertise, BGP, VxLAN, EVPN, Data center architecture, Network automation (Ansible, Salt, Python), Advanced network troubleshooting, Linux administration, Customer-facing experience, AI tools usage, Kubernetes, Docker, Networking simulation tools (NVIDIA Air, GNS3, EVE-NG), Network management tools (Grafana, Prometheus, Datadog)

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Senior Site Reliability Engineer - Storage

28 days ago
$168,000 to $322,000 per yearopen to United StatesSeniorFull-time

Python · Go · AWS

8 years of experience, HPC storage solutions, Enterprise NAS solutions, Distributed filesystems (Lustre, GPFS), Python/Bash/Golang, Cloud services (AWS, Azure, GCP), Monitoring stacks (Prometheus, Grafana, etc.), RDMA fabrics (InfiniBand, RoCE), HPC cluster management tools (Slurm, PBS, LSF), Containerization (Docker, Kubernetes)

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Site Reliability Engineer

28 days ago
salary not disclosedopen to WorldwideFull-time

AWS · Kubernetes

7 years of experience, Event-driven architecture, Deep AWS, Infrastructure as Code, Kubernetes, Observability and SLOs, Chaos engineering, Distributed systems debugging, Proven technical leadership, AI / MLOps infrastructure, Experience in payments industry

YunoYuno is a global payment orchestration platform headquartered in [location], specializing in B2B payment infrastructure that enables businesses to integrate over 1,000 payment methods through a single API.

Site Reliability Engineer

28 days ago
salary not disclosedopen to United StatesFull-time

Node.js · Python · Java

Ansible, Terraform, Kubernetes, Linux, Windows, Python, Java, Golang, Node.js, Nginx, HAProxy, Docker, cloud-first mindset, security-first mindset

Ontrac SolutionsOntrac Solutions is a technology consulting firm specializing in AI-driven systems and digital transformation, headquartered remotely with a focus on B2B services for healthcare, retail, and wellness industries.