Remote Site Reliability Engineer jobs.

“Remote”? Depends where. $$$? When they say. Stack? Before the essay.

Where can you apply from?

96 openings. None older than 30 days.

Newest first · page 2

Senior R&D Site Reliability Engineer

8 days ago
salary not disclosedopen to ChinaSeniorFull-time

Python · AWS

Terraform, AWS architecture, CI/CD (Gitlab runner), Linux system management, Python scripting, Observability tools (Prometheus, Grafana, ELK Stack), Automation tools development, Analytical and problem-solving skills, English proficiency (IELTS 6.5), Mandarin Chinese (plus)

VirtuosVirtuos is a video game development company specializing in full-cycle game development and art production for console, PC, and mobile titles.

Senior Site Reliability Engineer

8 days ago
salary not disclosedopen to IsraelSeniorFull-time

Python · Kubernetes

8 years of experience, Database engineering, MySQL, MSSQL, Oracle, Query optimization, Database-as-a-Service, Python, Go, Kubernetes, Hybrid database replication, Observability tools

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Senior Site Reliability Engineering - Storage

8 days ago
salary not disclosedopen to IsraelSeniorFull-time

Python · Kubernetes · Docker

12 years of experience, NAS, SAN, Object Storage, SRE concepts, Infrastructure as Code, Terraform, Ansible, Docker, Kubernetes, Python, Go, Shell, AI/ML storage, data analytics, debugging distributed systems, mentoring engineers

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Site Reliability Engineer Technical Lead

8 days ago
$100,000 to $150,000 per yearopen to United StatesLeadFull-time

Python · AWS · Azure

8 years of experience, Deep SRE and Systems Expertise, AWS, GCP, Azure, Kubernetes, Python, Automation and Tooling, Dynatrace, Splunk, ELK Stack, Observability and Analysis, Exceptional leadership, Communication skills

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

NCX Senior Engineer

9 days ago
$224,000 to $356,500 per yearopen to United StatesSeniorFull-time

Kubernetes

8 years of experience, GPU infrastructure management, NVIDIA technologies, Kubernetes expertise, Infrastructure observability tools, Automation for lifecycle management, Linux-based distributed systems, Collaboration with cloud partners, Failure modes in distributed AI workloads

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Site Reliability Engineer

11 days ago
$80,000 to $120,000 per yearopen to SwedenFull-time

Python · AWS · Azure

Python, Docker, Kubernetes, Google Cloud Platform, AWS, Azure, Terraform, Cloudformation, SaltStack, Ansible, Grafana, Prometheus, ELK stack, Game development passion, Independent work, Collaborative spirit

SharkmobSharkmob is a Swedish game developer that creates AAA-quality PC and console games, including Exoborne and Bloodhunt.

Site Reliability Engineer

12 days ago
salary not disclosedopen to IndiaFull-time

AWS · Azure · GCP

Kubernetes, Cloud platforms (AWS, Azure, GCP), Infrastructure-as-code (Terraform, AWS CDK), Observability tools (OpenTelemetry, Prometheus, Grafana), CI/CD pipelines, AI/ML exposure, Scripting for automation, Container orchestration

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Senior Staff Site Reliability Engineer

12 days ago
salary not disclosedopen to IndiaStaffFull-time

Kubernetes

8 years of experience, Kubernetes-based platforms, AI inference services, Control planes, Platform APIs, GPU scheduling, Vector databases, GitOps, CI/CD, Observability tools, Self-service platforms, Inference-serving frameworks, Open-source contributions

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Senior Infrastructure Security Engineer

12 days ago
kr702,200 to kr1,053,400 per yearopen to DenmarkSeniorFull-time

Python · AWS · Azure

Cloud security (GCP, AWS, Azure), Microservices (Containers, Kubernetes), Automation (Bash, Python, Terraform, Ansible), Security standards (NIST, PCI-DSS, SOC2), AI/ML infrastructure security, Network security fundamentals, Security features (AuthN, AuthZ, PKI)

UnityUnity Technologies is a San Francisco-based gaming company providing a leading B2B SaaS platform for real-time 3D development, widely used by creators across various industries globally.

Staff Production Operations Engineer

12 days ago
$199,400 to $299,200 per yearopen to United StatesStaffFull-time

AWS · Kubernetes

10 years of experience, AWS, Kubernetes, AI-assisted tools, Distributed service architecture, Incident management, Automation frameworks, Observability, CI/CD, Full-stack engineering, Production operations

sonyinteractiveentertainmentglobalSony Interactive Entertainment is a San Mateo-based global video game and digital entertainment company, primarily B2C, known for the PlayStation brand and its innovative gaming hardware, software, and network services.

Senior Software Engineer, Platform & Infrastructure - Riot Technology

13 days ago
salary not disclosedopen to United StatesSeniorFull-time

Python · AWS · GCP

3 years of experience, Kubernetes, AWS or GCP, Infrastructure-as-code, CI/CD, GPU compute infrastructure, Python, Distributed systems, MLOps workflows, Multi-node orchestration

Riot GamesRiot Games is a Los Angeles-based video game developer and publisher specializing in competitive multiplayer esports titles, operating globally with a B2C business model.

Infrastructure Solutions Architect - OEM Deployment

13 days ago
$124,000 to $241,500 per yearopen to United StatesFull-time

Python

2 years of experience, Large-scale datacenter rollouts, NVIDIA Cloud Partner integration, TCP/IP networking expertise, Bash scripting, Ansible, Python programming, GPU and DPU technologies, Signal integrity principles, Thermal management, Power distribution, Cabling and rack layout

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Monitoring Platform Engineer (Mid-Level)

13 days ago
salary not disclosedopen to United StatesFull-time

3 years of experience, Splunk, DataDog, AppDynamics, Monitoring solutions, Attention to detail, Written communication

Makpar CorporationMakpar Corporation is a Centreville, VA-based B2G IT solutions provider specializing in professional and technical services for the U.S. federal government, focusing on IT modernization, cybersecurity, and cloud migration.

Monitoring Platform Engineer (Senior)

13 days ago
salary not disclosedopen to United StatesSeniorFull-time

6 years of experience, Splunk SOAR, Ansible, Enterprise monitoring tools, Secured environments, Strong troubleshooting skills, ITSM tools

Makpar CorporationMakpar Corporation is a Centreville, VA-based B2G IT solutions provider specializing in professional and technical services for the U.S. federal government, focusing on IT modernization, cybersecurity, and cloud migration.

Senior Site Reliability Engineer, DGX Cloud

14 days ago
salary not disclosedopen to SwitzerlandSeniorFull-time

Python · Kubernetes

10 years of experience, Kubernetes administration, GPU workloads optimization, Infrastructure automation (Terraform, Ansible), High-level programming (Python, Go), Linux operating systems, SRE principles (SLOs, SLIs), Observability stacks (OpenTelemetry, Prometheus), GPU-accelerated clusters with KubeVirt, Generative-AI techniques, Workflow orchestration (Temporal, Airflow)

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Senior Site Reliability Engineer

15 days ago
salary not disclosedopen to IndiaSeniorFull-time

TypeScript · Python · Kubernetes

5 years of experience, Kubernetes (EKS, GKE), Terraform, Pulumi, GitOps (ArgoCD, Flux), Datadog, Prometheus, Grafana, SLI/SLO/error budget fluency, Go, Python, TypeScript, Linux internals, TCP/IP networking, Incident response leadership, Live-service game experience, Service mesh (Istio, Cilium), FinOps, Cloud certifications

2k2K is a Novato, California-based video game publisher specializing in a diverse range of genres including sports, action, and role-playing, primarily operating in the B2C market.

Site Reliability Engineer

15 days ago
salary not disclosedopen to WorldwideFull-time

Python · Azure · Kubernetes

3 years of experience, SLIs/SLOs definition, Multi-tenant SaaS platforms, Datadog, Grafana, Elastic Stack, High-availability architectures, Kubernetes, Python, Bash, Incident response, Cloud experience (Azure)

HostPapaHostPapa is a Canadian-based web hosting company offering B2B and B2C solutions, including shared, reseller, and VPS hosting services, with a focus on small businesses and a global presence.

Site Reliability Engineer

15 days ago
€50,000 to €80,000 per yearopen to Spain + PortugalFull-time

Go · AWS · GCP

Kubernetes, AWS, Google Cloud, Golang, CI/CD optimization, Distributed systems management, Monitoring tools (Prometheus, Grafana), Troubleshooting complex infrastructure issues, Disaster recovery strategies, Cloud security practices

ArangoDBArangoDB is a San Francisco-based B2B multi-model database platform that integrates graph, document, and key-value data models, serving industries like AI, healthcare, and finance globally.

Senior Site Reliability Engineer, BCM - DGX Cloud

15 days ago
$208,000 to $333,500 per yearopen to United StatesSeniorFull-time

Python · Kubernetes

8 years of experience, Fluency in Python, In-depth knowledge of Linux, Cluster networking proficiency, Experience with Kubernetes, High-performance computing experience, System administration experience with BCM

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Senior Site Reliability Engineer in Test, SDET

15 days ago
salary not disclosedopen to ChinaSeniorFull-time

Python · Kubernetes

8 years of experience, GitLab CI, ArgoCD, Kubernetes, Linux, Python, SRE principles, SBOM tooling, AI/ML techniques, test environment management, chaos engineering

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Site Reliability Engineer

15 days ago
salary not disclosedopen to United StatesContract

AWS · Azure · Kubernetes

7 years of experience, AWS, Azure, Kubernetes, Cloud-native transformation, Cloud networking, Terraform, Security best practices

ASCENDINGASCENDING is a B2B technology services company specializing in AI-driven solutions and cloud transformation strategies, headquartered in an unspecified location, serving various industries including finance, healthcare, and education.

Site Reliability Engineer (Europe)

15 days ago
salary not disclosedopen to EuropeFull-time

Go · AWS · GCP

Golang, AWS, Google Cloud, Kubernetes, CI/CD pipelines, Prometheus, Grafana, ELK stack, Core networking, Security best practices, Automation tools

ArangoDBArangoDB is a San Francisco-based B2B multi-model database platform that integrates graph, document, and key-value data models, serving industries like AI, healthcare, and finance globally.

Cloud Infrastructure Engineer

16 days ago
$125,000 to $135,000 per yearopen to United StatesFull-time

Azure

4 years of experience, Azure administration, PowerShell scripting, Microsoft Entra ID, DevSecOps collaboration, Cloud Operations experience, Troubleshooting complex infrastructure issues, Automation solutions development, Networking fundamentals

AIP PublishingAIP Publishing is a Melville, NY-based B2B scholarly publisher specializing in peer-reviewed journals and resources in the physical sciences, serving a global community of researchers and academic institutions.

Senior Network Engineer - Operations

16 days ago
salary not disclosedopen to United StatesSeniorFull-time

Python

5 years of experience, Large Ethernet fabrics, Multi-Vendor & NOS (Juniper, Cisco, Arista), Scale-up vs scale-out architectures, Production networks at scale, Automation (Python/Go, Ansible), 100G+ environments, AI, GPU, or HPC exposure

TensorWaveTensorWave is a Las Vegas-based B2B cloud computing provider specializing in AI and high-performance computing infrastructure, utilizing AMD Instinct GPUs to deliver scalable solutions for enterprises and AI researchers.

Senior Support Engineer - Toronto

16 days ago
salary not disclosedopen to CanadaSeniorFull-time

Python

8 years of experience, API platform expertise, Automation in support operations, Advanced monitoring and alerting, Incident response leadership, Scripting (Python), Cloud infrastructure knowledge, Cross-functional communication

OpenAIOpenAI is a San Francisco-based AI research and deployment company specializing in generative AI models and cloud-based services, operating primarily in B2B markets while also reaching consumers through products like ChatGPT.

Senior Site Reliability Engineer

17 days ago
$93,700 to $138,700 per yearopen to CanadaSeniorFull-time

TypeScript · Python · Kubernetes

5 years of experience, Kubernetes (EKS, GKE), Terraform, Pulumi, GitOps (ArgoCD, Flux), Observability stack (Prometheus, Grafana, Datadog), SLI/SLO/error budget fluency, Production-quality code in Go, Python, or TypeScript, Incident response leadership, Service mesh (Istio, Cilium), Live-service game experience

2k2K is a Novato, California-based video game publisher specializing in a diverse range of genres including sports, action, and role-playing, primarily operating in the B2C market.

Lab Support Engineer 2, Hopkinton, MA

18 days ago
$83,045 to $107,470 per yearopen to United StatesFull-time

Python

Python, Bash, PowerShell, Dell PowerEdge, PowerStore, Linux administration, VMware, vCenter, datacenter operations, networking fundamentals, problem-solving skills

DellDell Technologies designs and develops hardware and software for infrastructure, enterprise solutions, and data management.

Associate Infrastructure Engineer

18 days ago
$158,300 to $190,000 per yearopen to United States + CanadaFull-time

Python · AWS · GCP

2 years of experience, Kubernetes, AWS or GCP, GitOps tools, Observability stack, Infrastructure-as-code tools, Node, Python or Go, Debugging distributed systems, On-call rotation participation, Proactive embrace of AI

WebflowWebflow is a San Francisco-based SaaS company providing a Website Experience Platform (WXP) that empowers B2B marketing teams to design, manage, and optimize custom websites for a global audience.

Staff Platform Engineer, Core Cloud Platform

18 days ago
$223,100 to $305,000 per yearopen to United States + Canada + United Kingdom + Singapore + India + Ireland + FinlandStaffFull-time

Python · Go · AWS

Kubernetes expertise, AWS cloud proficiency, Production SaaS systems experience, Networking and service mesh knowledge, Operational troubleshooting skills, Python and Golang proficiency, CI/CD practices advocacy, Technical roadmap design and leadership

AlphaSenseAlphaSense is a New York City-based B2B fintech platform specializing in AI-driven market intelligence and search solutions for financial institutions and top companies globally.

Cloud 2nd line ops Engineer

18 days ago
€43,200 to €45,600 per yearopen to PortugalFull-time

Azure

Terraform, Azure DevOps, Azure cloud services, Azure resources, Azure RBAC, Azure Networking, Problem-Solving, Effective Communication, Documentation, Customer centric approach, Adaptability

Irium PortugalIrium Portugal is a Lisbon-based B2B IT services provider specializing in managed services, digital transformation, and cybersecurity solutions for businesses across Portugal.