Remote Site Reliability Engineer jobs.

“Remote”? Depends where. $$$? When they say. Stack? Before the essay.

Where can you apply from?

107 openings. None older than 30 days.

Newest first

Cloud Operations Engineer

about 9 hours ago
$135,000 to $150,000 per yearopen to United StatesFull-time

Python · GCP · Kubernetes

3 years of experience, GCP, Cloud infrastructure engineering, Observability tools, Python, Docker, Kubernetes, AI tools, 24/7 on-call rotation

BranchBranch is a remote-first fintech company specializing in workforce payment solutions, headquartered in the mid-Atlantic region, targeting working Americans with B2B/B2C financial services.

Site Reliability Engineer

1 day ago
salary not disclosedopen to IndiaFull-time

AWS · Azure · GCP

Kubernetes, Cloud platforms (AWS, Azure, GCP), Infrastructure-as-code (Terraform, AWS CDK), Observability tools (OpenTelemetry, Prometheus, Grafana), CI/CD pipelines, AI/ML exposure, Scripting for automation, Container orchestration

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Senior Staff Site Reliability Engineer

1 day ago
salary not disclosedopen to IndiaStaffFull-time

Kubernetes

8 years of experience, Kubernetes-based platforms, AI inference services, Control planes, Platform APIs, GPU scheduling, Vector databases, GitOps, CI/CD, Observability tools, Self-service platforms, Inference-serving frameworks, Open-source contributions

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Senior Infrastructure Security Engineer

1 day ago
kr702,200 to kr1,053,400 per yearopen to DenmarkSeniorFull-time

Python · AWS · Azure

Cloud security (GCP, AWS, Azure), Microservices (Containers, Kubernetes), Automation (Bash, Python, Terraform, Ansible), Security standards (NIST, PCI-DSS, SOC2), AI/ML infrastructure security, Network security fundamentals, Security features (AuthN, AuthZ, PKI)

UnityUnity Technologies is a San Francisco-based gaming company providing a leading B2B SaaS platform for real-time 3D development, widely used by creators across various industries globally.

Staff Production Operations Engineer

1 day ago
$199,400 to $299,200 per yearopen to United StatesStaffFull-time

AWS · Kubernetes

10 years of experience, AWS, Kubernetes, AI-assisted tools, Distributed service architecture, Incident management, Automation frameworks, Observability, CI/CD, Full-stack engineering, Production operations

sonyinteractiveentertainmentglobalSony Interactive Entertainment is a San Mateo-based global video game and digital entertainment company, primarily B2C, known for the PlayStation brand and its innovative gaming hardware, software, and network services.

Senior Software Engineer, Platform & Infrastructure - Riot Technology

1 day ago
salary not disclosedopen to United StatesSeniorFull-time

Python · AWS · GCP

3 years of experience, Kubernetes, AWS or GCP, Infrastructure-as-code, CI/CD, GPU compute infrastructure, Python, Distributed systems, MLOps workflows, Multi-node orchestration

Riot GamesRiot Games is a Los Angeles-based video game developer and publisher specializing in competitive multiplayer esports titles, operating globally with a B2C business model.

Data Center Engineer

1 day ago
$126,810 to $153,900 per yearopen to United StatesFull-time

6 years of experience, Large-scale Data Center Infrastructure, Server and network equipment installation, Real-time requirements management, Root cause analysis, Automation of maintenance actions, Development of infrastructure standards, Cross-functional collaboration, Ability to lift 75 pounds

RobloxRoblox is a San Mateo-based gaming and entertainment platform that enables user-generated content and experiences, operating primarily as a B2C service with a global family-friendly audience.

Infrastructure Solutions Architect - OEM Deployment

2 days ago
$124,000 to $241,500 per yearopen to United StatesFull-time

Python

2 years of experience, Large-scale datacenter rollouts, NVIDIA Cloud Partner integration, TCP/IP networking expertise, Bash scripting, Ansible, Python programming, GPU and DPU technologies, Signal integrity principles, Thermal management, Power distribution, Cabling and rack layout

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Monitoring Platform Engineer (Mid-Level)

2 days ago
salary not disclosedopen to United StatesFull-time

3 years of experience, Splunk, DataDog, AppDynamics, Monitoring solutions, Attention to detail, Written communication

Makpar CorporationMakpar Corporation is a Centreville, VA-based B2G IT solutions provider specializing in professional and technical services for the U.S. federal government, focusing on IT modernization, cybersecurity, and cloud migration.

Monitoring Platform Engineer (Senior)

2 days ago
salary not disclosedopen to United StatesSeniorFull-time

6 years of experience, Splunk SOAR, Ansible, Enterprise monitoring tools, Secured environments, Strong troubleshooting skills, ITSM tools

Makpar CorporationMakpar Corporation is a Centreville, VA-based B2G IT solutions provider specializing in professional and technical services for the U.S. federal government, focusing on IT modernization, cybersecurity, and cloud migration.

Site Reliability Engineer SME - Senior

2 days ago
salary not disclosedopen to United StatesSeniorFull-time

6 years of experience, OpenTelemetry, Splunk SOAR, Ansible, Event-to-incident workflow design, Dashboard building, Operational reporting

Makpar CorporationMakpar Corporation is a Centreville, VA-based B2G IT solutions provider specializing in professional and technical services for the U.S. federal government, focusing on IT modernization, cybersecurity, and cloud migration.

Kubernetes & OpenShift Engineer

2 days ago
$135,000 to $155,000 per yearopen to United StatesFull-time

Python · Kubernetes

6 years of experience, Kubernetes internals, OpenShift internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Cluster monitoring and logging, Container image security

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

Observability Engineer

2 days ago
$100,000 to $160,000 per yearopen to United StatesFull-time

12 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, Distributed tracing, High-cardinality metrics, SLOs, CI/CD integration, Linux internals, eBPF-based observability

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

Site Reliability Engineer (SRE)

2 days ago
$100,000 to $180,000 per yearopen to United StatesFull-time

Python · AWS · Azure

10 years of experience, Kubernetes, Prometheus, Grafana, Python, Go, CI/CD pipelines, Chaos engineering, Distributed systems, SLOs and error budgets, Cloud platforms (AWS, Azure, GCP)

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

Site Reliability Engineer Technical Lead

2 days ago
$100,000 to $150,000 per yearopen to United StatesLeadFull-time

Python · AWS · Azure

8 years of experience, Deep SRE and Systems Expertise, AWS, GCP, Azure, Kubernetes, Python, Automation and Tooling, Dynatrace, Splunk, ELK Stack, Observability and Analysis, Exceptional leadership, Communication skills

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

Azure Infrastructure Engineer

2 days ago
$100,000 to $120,000 per yearopen to United StatesFull-time

Python · Azure · Kubernetes

6 years of experience, Azure core services, Infrastructure-as-code (Terraform, Bicep, ARM), Azure Kubernetes Service (AKS), Azure DevOps or GitHub Actions, PowerShell, Bash, Python scripting, Cloud security principles, Monitoring and observability strategies, Hybrid cloud or multi-cloud experience, FinOps practices, Regulated environments (HIPAA, PCI-DSS, SOC 2, FedRAMP)

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

Virtual Platform Engineer

2 days ago
$125,000 to $145,000 per yearopen to United StatesFull-time

AWS · Kubernetes

6 years of experience, vSphere, vSAN, NSX-T, PowerCLI, Tanzu Kubernetes Grid, disaster recovery patterns, troubleshooting across compute, network, and storage layers, VMware Cloud on AWS

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

Senior Site Reliability Engineer, DGX Cloud

3 days ago
salary not disclosedopen to SwitzerlandSeniorFull-time

Python · Kubernetes

10 years of experience, Kubernetes administration, GPU workloads optimization, Infrastructure automation (Terraform, Ansible), High-level programming (Python, Go), Linux operating systems, SRE principles (SLOs, SLIs), Observability stacks (OpenTelemetry, Prometheus), GPU-accelerated clusters with KubeVirt, Generative-AI techniques, Workflow orchestration (Temporal, Airflow)

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Senior Site Reliability Engineer

4 days ago
salary not disclosedopen to IndiaSeniorFull-time

TypeScript · Python · Kubernetes

5 years of experience, Kubernetes (EKS, GKE), Terraform, Pulumi, GitOps (ArgoCD, Flux), Datadog, Prometheus, Grafana, SLI/SLO/error budget fluency, Go, Python, TypeScript, Linux internals, TCP/IP networking, Incident response leadership, Live-service game experience, Service mesh (Istio, Cilium), FinOps, Cloud certifications

2k2K is a Novato, California-based video game publisher specializing in a diverse range of genres including sports, action, and role-playing, primarily operating in the B2C market.

Observability Engineer

4 days ago
$89,000 to $110,000 per yearopen to United StatesFull-time

Python · Java

6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, distributed tracing, structured logging, Go, Python, Java, high-cardinality metrics, SLOs, error budgets, SRE principles, CI/CD integration, Linux internals, networking, container platforms, Thanos, Mimir, Cortex, Loki, Tempo, eBPF, cost optimization, regulated environments

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

OpenShift Platform Engineer

4 days ago
$86,000 to $103,000 per yearopen to United StatesFull-time

Python · Kubernetes

6 years of experience, OpenShift, Kubernetes internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Container image security, Cluster monitoring tools, Service mesh (OpenShift Service Mesh, Istio, Linkerd), GitOps workflows (Argo CD, Flux)

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

Service Mesh Engineer

4 days ago
$95,000 to $115,000 per yearopen to United StatesFull-time

Python · Kubernetes

6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, Observability

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

Site Reliability Engineer

4 days ago
salary not disclosedopen to WorldwideFull-time

Python · Azure · Kubernetes

3 years of experience, SLIs/SLOs definition, Multi-tenant SaaS platforms, Datadog, Grafana, Elastic Stack, High-availability architectures, Kubernetes, Python, Bash, Incident response, Cloud experience (Azure)

HostPapaHostPapa is a Canadian-based web hosting company offering B2B and B2C solutions, including shared, reseller, and VPS hosting services, with a focus on small businesses and a global presence.

Site Reliability Engineer

4 days ago
€50,000 to €80,000 per yearopen to Spain + PortugalFull-time

Go · AWS · GCP

Kubernetes, AWS, Google Cloud, Golang, CI/CD optimization, Distributed systems management, Monitoring tools (Prometheus, Grafana), Troubleshooting complex infrastructure issues, Disaster recovery strategies, Cloud security practices

ArangoDBArangoDB is a San Francisco-based B2B multi-model database platform that integrates graph, document, and key-value data models, serving industries like AI, healthcare, and finance globally.

Senior Site Reliability Engineer, BCM - DGX Cloud

4 days ago
$208,000 to $333,500 per yearopen to United StatesSeniorFull-time

Python · Kubernetes

8 years of experience, Fluency in Python, In-depth knowledge of Linux, Cluster networking proficiency, Experience with Kubernetes, High-performance computing experience, System administration experience with BCM

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Senior Site Reliability Engineer in Test, SDET

4 days ago
salary not disclosedopen to ChinaSeniorFull-time

Python · Kubernetes

8 years of experience, GitLab CI, ArgoCD, Kubernetes, Linux, Python, SRE principles, SBOM tooling, AI/ML techniques, test environment management, chaos engineering

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Site Reliability Engineer

4 days ago
salary not disclosedopen to United StatesContract

AWS · Azure · Kubernetes

7 years of experience, AWS, Azure, Kubernetes, Cloud-native transformation, Cloud networking, Terraform, Security best practices

ASCENDINGASCENDING is a B2B technology services company specializing in AI-driven solutions and cloud transformation strategies, headquartered in an unspecified location, serving various industries including finance, healthcare, and education.

Cloud Infrastructure Engineer

5 days ago
$125,000 to $135,000 per yearopen to United StatesFull-time

Azure

4 years of experience, Azure administration, PowerShell scripting, Microsoft Entra ID, DevSecOps collaboration, Cloud Operations experience, Troubleshooting complex infrastructure issues, Automation solutions development, Networking fundamentals

AIP PublishingAIP Publishing is a Melville, NY-based B2B scholarly publisher specializing in peer-reviewed journals and resources in the physical sciences, serving a global community of researchers and academic institutions.

Senior Network Engineer - Operations

5 days ago
salary not disclosedopen to United StatesSeniorFull-time

Python

5 years of experience, Large Ethernet fabrics, Multi-Vendor & NOS (Juniper, Cisco, Arista), Scale-up vs scale-out architectures, Production networks at scale, Automation (Python/Go, Ansible), 100G+ environments, AI, GPU, or HPC exposure

TensorWaveTensorWave is a Las Vegas-based B2B cloud computing provider specializing in AI and high-performance computing infrastructure, utilizing AMD Instinct GPUs to deliver scalable solutions for enterprises and AI researchers.

Lead Site Reliability Engineer (Performance & Scalability) | Contract | Remote US

5 days ago
salary not disclosedopen to United StatesLeadContract

Site Reliability Engineering, Performance engineering, Capacity planning, Observability, Production distributed systems, SLOs and SLIs, Load and stress testing, Incident management, Graceful degradation design, Technical performance translation

Tech HoldingTech Holding is a Los Angeles-based B2B consulting firm specializing in professional services and managed solutions for cloud transformation and technology integration across various industries.