Remote Site Reliability Engineer jobs.

“Remote”? Depends where. $$$? When they say. Stack? Before the essay.

Where can you apply from?

103 openings. None older than 30 days.

Newest first

Cloud Operations Engineer

1 day ago
$135,000 to $150,000 per yearopen to United StatesFull-time

Python · GCP · Kubernetes

3 years of experience, GCP, Cloud infrastructure engineering, Observability tools, Python, Docker, Kubernetes, AI tools, 24/7 on-call rotation

BranchBranch is a remote-first fintech company specializing in workforce payment solutions, headquartered in the mid-Atlantic region, targeting working Americans with B2B/B2C financial services.

Site Reliability Engineer

1 day ago
$80,000 to $120,000 per yearopen to SwedenFull-time

Python · AWS · Azure

Python, Docker, Kubernetes, Google Cloud Platform, AWS, Azure, Terraform, Cloudformation, SaltStack, Ansible, Grafana, Prometheus, ELK stack, Game development passion, Independent work, Collaborative spirit

SharkmobSharkmob is a Swedish game developer that creates AAA-quality PC and console games, including Exoborne and Bloodhunt.

Systems Observability Specialist

2 days ago
$135,000 to $155,000 per yearopen to United StatesFull-time

6 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, SLOs, High-cardinality metrics, CI/CD integration, Linux internals, Container platforms, Observability cost optimization

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

Kubernetes Service Engineer

2 days ago
$100,000 to $150,000 per yearopen to United StatesFull-time

Python · Kubernetes

6 years of experience, Istio, Linkerd, Envoy, mTLS, Kubernetes networking, Go, Python, Distributed tracing, Service mesh architecture, Zero-trust networking

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

Platform Reliability Engineer

2 days ago
$100,000 to $150,000 per yearopen to United StatesFull-time

Python · Java · Kubernetes

6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Observability tooling, CI/CD pipelines, Distributed system design, SLOs and error budgets, Chaos engineering, Cloud platforms

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

Site Observability Engineer

2 days ago
$100,000 to $150,000 per yearopen to United StatesFull-time

6 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, Distributed tracing, High-cardinality metrics, SLOs, CI/CD integration, Linux internals, Container platforms

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

Site Reliability Engineer

2 days ago
salary not disclosedopen to IndiaFull-time

AWS · Azure · GCP

Kubernetes, Cloud platforms (AWS, Azure, GCP), Infrastructure-as-code (Terraform, AWS CDK), Observability tools (OpenTelemetry, Prometheus, Grafana), CI/CD pipelines, AI/ML exposure, Scripting for automation, Container orchestration

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Senior Staff Site Reliability Engineer

2 days ago
salary not disclosedopen to IndiaStaffFull-time

Kubernetes

8 years of experience, Kubernetes-based platforms, AI inference services, Control planes, Platform APIs, GPU scheduling, Vector databases, GitOps, CI/CD, Observability tools, Self-service platforms, Inference-serving frameworks, Open-source contributions

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Senior Infrastructure Security Engineer

2 days ago
kr702,200 to kr1,053,400 per yearopen to DenmarkSeniorFull-time

Python · AWS · Azure

Cloud security (GCP, AWS, Azure), Microservices (Containers, Kubernetes), Automation (Bash, Python, Terraform, Ansible), Security standards (NIST, PCI-DSS, SOC2), AI/ML infrastructure security, Network security fundamentals, Security features (AuthN, AuthZ, PKI)

UnityUnity Technologies is a San Francisco-based gaming company providing a leading B2B SaaS platform for real-time 3D development, widely used by creators across various industries globally.

Staff Production Operations Engineer

2 days ago
$199,400 to $299,200 per yearopen to United StatesStaffFull-time

AWS · Kubernetes

10 years of experience, AWS, Kubernetes, AI-assisted tools, Distributed service architecture, Incident management, Automation frameworks, Observability, CI/CD, Full-stack engineering, Production operations

sonyinteractiveentertainmentglobalSony Interactive Entertainment is a San Mateo-based global video game and digital entertainment company, primarily B2C, known for the PlayStation brand and its innovative gaming hardware, software, and network services.

Senior Software Engineer, Platform & Infrastructure - Riot Technology

3 days ago
salary not disclosedopen to United StatesSeniorFull-time

Python · AWS · GCP

3 years of experience, Kubernetes, AWS or GCP, Infrastructure-as-code, CI/CD, GPU compute infrastructure, Python, Distributed systems, MLOps workflows, Multi-node orchestration

Riot GamesRiot Games is a Los Angeles-based video game developer and publisher specializing in competitive multiplayer esports titles, operating globally with a B2C business model.

Data Center Engineer

3 days ago
$126,810 to $153,900 per yearopen to United StatesFull-time

6 years of experience, Large-scale Data Center Infrastructure, Server and network equipment installation, Real-time requirements management, Root cause analysis, Automation of maintenance actions, Development of infrastructure standards, Cross-functional collaboration, Ability to lift 75 pounds

RobloxRoblox is a San Mateo-based gaming and entertainment platform that enables user-generated content and experiences, operating primarily as a B2C service with a global family-friendly audience.

Infrastructure Solutions Architect - OEM Deployment

3 days ago
$124,000 to $241,500 per yearopen to United StatesFull-time

Python

2 years of experience, Large-scale datacenter rollouts, NVIDIA Cloud Partner integration, TCP/IP networking expertise, Bash scripting, Ansible, Python programming, GPU and DPU technologies, Signal integrity principles, Thermal management, Power distribution, Cabling and rack layout

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Monitoring Platform Engineer (Mid-Level)

3 days ago
salary not disclosedopen to United StatesFull-time

3 years of experience, Splunk, DataDog, AppDynamics, Monitoring solutions, Attention to detail, Written communication

Makpar CorporationMakpar Corporation is a Centreville, VA-based B2G IT solutions provider specializing in professional and technical services for the U.S. federal government, focusing on IT modernization, cybersecurity, and cloud migration.

Monitoring Platform Engineer (Senior)

3 days ago
salary not disclosedopen to United StatesSeniorFull-time

6 years of experience, Splunk SOAR, Ansible, Enterprise monitoring tools, Secured environments, Strong troubleshooting skills, ITSM tools

Makpar CorporationMakpar Corporation is a Centreville, VA-based B2G IT solutions provider specializing in professional and technical services for the U.S. federal government, focusing on IT modernization, cybersecurity, and cloud migration.

Site Reliability Engineer SME - Senior

3 days ago
salary not disclosedopen to United StatesSeniorFull-time

6 years of experience, OpenTelemetry, Splunk SOAR, Ansible, Event-to-incident workflow design, Dashboard building, Operational reporting

Makpar CorporationMakpar Corporation is a Centreville, VA-based B2G IT solutions provider specializing in professional and technical services for the U.S. federal government, focusing on IT modernization, cybersecurity, and cloud migration.

Kubernetes & OpenShift Engineer

3 days ago
$135,000 to $155,000 per yearopen to United StatesFull-time

Python · Kubernetes

6 years of experience, Kubernetes internals, OpenShift internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Cluster monitoring and logging, Container image security

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

Observability Engineer

3 days ago
$100,000 to $160,000 per yearopen to United StatesFull-time

12 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, Distributed tracing, High-cardinality metrics, SLOs, CI/CD integration, Linux internals, eBPF-based observability

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

Site Reliability Engineer (SRE)

3 days ago
$100,000 to $180,000 per yearopen to United StatesFull-time

Python · AWS · Azure

10 years of experience, Kubernetes, Prometheus, Grafana, Python, Go, CI/CD pipelines, Chaos engineering, Distributed systems, SLOs and error budgets, Cloud platforms (AWS, Azure, GCP)

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

Site Reliability Engineer Technical Lead

3 days ago
$100,000 to $150,000 per yearopen to United StatesLeadFull-time

Python · AWS · Azure

8 years of experience, Deep SRE and Systems Expertise, AWS, GCP, Azure, Kubernetes, Python, Automation and Tooling, Dynatrace, Splunk, ELK Stack, Observability and Analysis, Exceptional leadership, Communication skills

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

Azure Infrastructure Engineer

3 days ago
$100,000 to $120,000 per yearopen to United StatesFull-time

Python · Azure · Kubernetes

6 years of experience, Azure core services, Infrastructure-as-code (Terraform, Bicep, ARM), Azure Kubernetes Service (AKS), Azure DevOps or GitHub Actions, PowerShell, Bash, Python scripting, Cloud security principles, Monitoring and observability strategies, Hybrid cloud or multi-cloud experience, FinOps practices, Regulated environments (HIPAA, PCI-DSS, SOC 2, FedRAMP)

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

Virtual Platform Engineer

3 days ago
$125,000 to $145,000 per yearopen to United StatesFull-time

AWS · Kubernetes

6 years of experience, vSphere, vSAN, NSX-T, PowerCLI, Tanzu Kubernetes Grid, disaster recovery patterns, troubleshooting across compute, network, and storage layers, VMware Cloud on AWS

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

Senior Site Reliability Engineer, DGX Cloud

4 days ago
salary not disclosedopen to SwitzerlandSeniorFull-time

Python · Kubernetes

10 years of experience, Kubernetes administration, GPU workloads optimization, Infrastructure automation (Terraform, Ansible), High-level programming (Python, Go), Linux operating systems, SRE principles (SLOs, SLIs), Observability stacks (OpenTelemetry, Prometheus), GPU-accelerated clusters with KubeVirt, Generative-AI techniques, Workflow orchestration (Temporal, Airflow)

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Senior Site Reliability Engineer

5 days ago
salary not disclosedopen to IndiaSeniorFull-time

TypeScript · Python · Kubernetes

5 years of experience, Kubernetes (EKS, GKE), Terraform, Pulumi, GitOps (ArgoCD, Flux), Datadog, Prometheus, Grafana, SLI/SLO/error budget fluency, Go, Python, TypeScript, Linux internals, TCP/IP networking, Incident response leadership, Live-service game experience, Service mesh (Istio, Cilium), FinOps, Cloud certifications

2k2K is a Novato, California-based video game publisher specializing in a diverse range of genres including sports, action, and role-playing, primarily operating in the B2C market.

Observability Engineer

5 days ago
$89,000 to $110,000 per yearopen to United StatesFull-time

Python · Java

6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, distributed tracing, structured logging, Go, Python, Java, high-cardinality metrics, SLOs, error budgets, SRE principles, CI/CD integration, Linux internals, networking, container platforms, Thanos, Mimir, Cortex, Loki, Tempo, eBPF, cost optimization, regulated environments

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

OpenShift Platform Engineer

5 days ago
$86,000 to $103,000 per yearopen to United StatesFull-time

Python · Kubernetes

6 years of experience, OpenShift, Kubernetes internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Container image security, Cluster monitoring tools, Service mesh (OpenShift Service Mesh, Istio, Linkerd), GitOps workflows (Argo CD, Flux)

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

Service Mesh Engineer

5 days ago
$95,000 to $115,000 per yearopen to United StatesFull-time

Python · Kubernetes

6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, Observability

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

Site Reliability Engineer

5 days ago
salary not disclosedopen to WorldwideFull-time

Python · Azure · Kubernetes

3 years of experience, SLIs/SLOs definition, Multi-tenant SaaS platforms, Datadog, Grafana, Elastic Stack, High-availability architectures, Kubernetes, Python, Bash, Incident response, Cloud experience (Azure)

HostPapaHostPapa is a Canadian-based web hosting company offering B2B and B2C solutions, including shared, reseller, and VPS hosting services, with a focus on small businesses and a global presence.

Site Reliability Engineer

5 days ago
€50,000 to €80,000 per yearopen to Spain + PortugalFull-time

Go · AWS · GCP

Kubernetes, AWS, Google Cloud, Golang, CI/CD optimization, Distributed systems management, Monitoring tools (Prometheus, Grafana), Troubleshooting complex infrastructure issues, Disaster recovery strategies, Cloud security practices

ArangoDBArangoDB is a San Francisco-based B2B multi-model database platform that integrates graph, document, and key-value data models, serving industries like AI, healthcare, and finance globally.

Senior Site Reliability Engineer, BCM - DGX Cloud

5 days ago
$208,000 to $333,500 per yearopen to United StatesSeniorFull-time

Python · Kubernetes

8 years of experience, Fluency in Python, In-depth knowledge of Linux, Cluster networking proficiency, Experience with Kubernetes, High-performance computing experience, System administration experience with BCM

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.