Remote Site Reliability Engineer jobs.
“Remote”? Depends where. $$$? When they say. Stack? Before the essay.
Where can you apply from?
96 openings. None older than 30 days.
Newest first · page 2
Senior R&D Site Reliability Engineer
8 days agoPython · AWS
Terraform, AWS architecture, CI/CD (Gitlab runner), Linux system management, Python scripting, Observability tools (Prometheus, Grafana, ELK Stack), Automation tools development, Analytical and problem-solving skills, English proficiency (IELTS 6.5), Mandarin Chinese (plus)
Senior Site Reliability Engineer
8 days agoPython · Kubernetes
8 years of experience, Database engineering, MySQL, MSSQL, Oracle, Query optimization, Database-as-a-Service, Python, Go, Kubernetes, Hybrid database replication, Observability tools
Senior Site Reliability Engineering - Storage
8 days agoPython · Kubernetes · Docker
12 years of experience, NAS, SAN, Object Storage, SRE concepts, Infrastructure as Code, Terraform, Ansible, Docker, Kubernetes, Python, Go, Shell, AI/ML storage, data analytics, debugging distributed systems, mentoring engineers
Site Reliability Engineer Technical Lead
8 days agoPython · AWS · Azure
8 years of experience, Deep SRE and Systems Expertise, AWS, GCP, Azure, Kubernetes, Python, Automation and Tooling, Dynatrace, Splunk, ELK Stack, Observability and Analysis, Exceptional leadership, Communication skills
NCX Senior Engineer
9 days agoKubernetes
8 years of experience, GPU infrastructure management, NVIDIA technologies, Kubernetes expertise, Infrastructure observability tools, Automation for lifecycle management, Linux-based distributed systems, Collaboration with cloud partners, Failure modes in distributed AI workloads
Site Reliability Engineer
11 days agoPython · AWS · Azure
Python, Docker, Kubernetes, Google Cloud Platform, AWS, Azure, Terraform, Cloudformation, SaltStack, Ansible, Grafana, Prometheus, ELK stack, Game development passion, Independent work, Collaborative spirit
Site Reliability Engineer
12 days agoAWS · Azure · GCP
Kubernetes, Cloud platforms (AWS, Azure, GCP), Infrastructure-as-code (Terraform, AWS CDK), Observability tools (OpenTelemetry, Prometheus, Grafana), CI/CD pipelines, AI/ML exposure, Scripting for automation, Container orchestration
Senior Staff Site Reliability Engineer
12 days agoKubernetes
8 years of experience, Kubernetes-based platforms, AI inference services, Control planes, Platform APIs, GPU scheduling, Vector databases, GitOps, CI/CD, Observability tools, Self-service platforms, Inference-serving frameworks, Open-source contributions
Senior Infrastructure Security Engineer
12 days agoPython · AWS · Azure
Cloud security (GCP, AWS, Azure), Microservices (Containers, Kubernetes), Automation (Bash, Python, Terraform, Ansible), Security standards (NIST, PCI-DSS, SOC2), AI/ML infrastructure security, Network security fundamentals, Security features (AuthN, AuthZ, PKI)
Staff Production Operations Engineer
12 days agoAWS · Kubernetes
10 years of experience, AWS, Kubernetes, AI-assisted tools, Distributed service architecture, Incident management, Automation frameworks, Observability, CI/CD, Full-stack engineering, Production operations
Senior Software Engineer, Platform & Infrastructure - Riot Technology
13 days agoPython · AWS · GCP
3 years of experience, Kubernetes, AWS or GCP, Infrastructure-as-code, CI/CD, GPU compute infrastructure, Python, Distributed systems, MLOps workflows, Multi-node orchestration
Infrastructure Solutions Architect - OEM Deployment
13 days agoPython
2 years of experience, Large-scale datacenter rollouts, NVIDIA Cloud Partner integration, TCP/IP networking expertise, Bash scripting, Ansible, Python programming, GPU and DPU technologies, Signal integrity principles, Thermal management, Power distribution, Cabling and rack layout
Monitoring Platform Engineer (Mid-Level)
13 days ago3 years of experience, Splunk, DataDog, AppDynamics, Monitoring solutions, Attention to detail, Written communication
Monitoring Platform Engineer (Senior)
13 days ago6 years of experience, Splunk SOAR, Ansible, Enterprise monitoring tools, Secured environments, Strong troubleshooting skills, ITSM tools
Senior Site Reliability Engineer, DGX Cloud
14 days agoPython · Kubernetes
10 years of experience, Kubernetes administration, GPU workloads optimization, Infrastructure automation (Terraform, Ansible), High-level programming (Python, Go), Linux operating systems, SRE principles (SLOs, SLIs), Observability stacks (OpenTelemetry, Prometheus), GPU-accelerated clusters with KubeVirt, Generative-AI techniques, Workflow orchestration (Temporal, Airflow)
Senior Site Reliability Engineer
15 days agoTypeScript · Python · Kubernetes
5 years of experience, Kubernetes (EKS, GKE), Terraform, Pulumi, GitOps (ArgoCD, Flux), Datadog, Prometheus, Grafana, SLI/SLO/error budget fluency, Go, Python, TypeScript, Linux internals, TCP/IP networking, Incident response leadership, Live-service game experience, Service mesh (Istio, Cilium), FinOps, Cloud certifications
Site Reliability Engineer
15 days agoPython · Azure · Kubernetes
3 years of experience, SLIs/SLOs definition, Multi-tenant SaaS platforms, Datadog, Grafana, Elastic Stack, High-availability architectures, Kubernetes, Python, Bash, Incident response, Cloud experience (Azure)
Site Reliability Engineer
15 days agoGo · AWS · GCP
Kubernetes, AWS, Google Cloud, Golang, CI/CD optimization, Distributed systems management, Monitoring tools (Prometheus, Grafana), Troubleshooting complex infrastructure issues, Disaster recovery strategies, Cloud security practices
Senior Site Reliability Engineer, BCM - DGX Cloud
15 days agoPython · Kubernetes
8 years of experience, Fluency in Python, In-depth knowledge of Linux, Cluster networking proficiency, Experience with Kubernetes, High-performance computing experience, System administration experience with BCM
Senior Site Reliability Engineer in Test, SDET
15 days agoPython · Kubernetes
8 years of experience, GitLab CI, ArgoCD, Kubernetes, Linux, Python, SRE principles, SBOM tooling, AI/ML techniques, test environment management, chaos engineering
Site Reliability Engineer
15 days agoAWS · Azure · Kubernetes
7 years of experience, AWS, Azure, Kubernetes, Cloud-native transformation, Cloud networking, Terraform, Security best practices
Site Reliability Engineer (Europe)
15 days agoGo · AWS · GCP
Golang, AWS, Google Cloud, Kubernetes, CI/CD pipelines, Prometheus, Grafana, ELK stack, Core networking, Security best practices, Automation tools
Cloud Infrastructure Engineer
16 days agoAzure
4 years of experience, Azure administration, PowerShell scripting, Microsoft Entra ID, DevSecOps collaboration, Cloud Operations experience, Troubleshooting complex infrastructure issues, Automation solutions development, Networking fundamentals
Senior Network Engineer - Operations
16 days agoPython
5 years of experience, Large Ethernet fabrics, Multi-Vendor & NOS (Juniper, Cisco, Arista), Scale-up vs scale-out architectures, Production networks at scale, Automation (Python/Go, Ansible), 100G+ environments, AI, GPU, or HPC exposure
Senior Support Engineer - Toronto
16 days agoPython
8 years of experience, API platform expertise, Automation in support operations, Advanced monitoring and alerting, Incident response leadership, Scripting (Python), Cloud infrastructure knowledge, Cross-functional communication
Senior Site Reliability Engineer
17 days agoTypeScript · Python · Kubernetes
5 years of experience, Kubernetes (EKS, GKE), Terraform, Pulumi, GitOps (ArgoCD, Flux), Observability stack (Prometheus, Grafana, Datadog), SLI/SLO/error budget fluency, Production-quality code in Go, Python, or TypeScript, Incident response leadership, Service mesh (Istio, Cilium), Live-service game experience
Lab Support Engineer 2, Hopkinton, MA
18 days agoPython
Python, Bash, PowerShell, Dell PowerEdge, PowerStore, Linux administration, VMware, vCenter, datacenter operations, networking fundamentals, problem-solving skills
Associate Infrastructure Engineer
18 days agoPython · AWS · GCP
2 years of experience, Kubernetes, AWS or GCP, GitOps tools, Observability stack, Infrastructure-as-code tools, Node, Python or Go, Debugging distributed systems, On-call rotation participation, Proactive embrace of AI
Staff Platform Engineer, Core Cloud Platform
18 days agoPython · Go · AWS
Kubernetes expertise, AWS cloud proficiency, Production SaaS systems experience, Networking and service mesh knowledge, Operational troubleshooting skills, Python and Golang proficiency, CI/CD practices advocacy, Technical roadmap design and leadership
Cloud 2nd line ops Engineer
18 days agoAzure
Terraform, Azure DevOps, Azure cloud services, Azure resources, Azure RBAC, Azure Networking, Problem-Solving, Effective Communication, Documentation, Customer centric approach, Adaptability