Remote Site Reliability Engineer jobs.
“Remote”? Depends where. $$$? When they say. Stack? Before the essay.
Where can you apply from?
107 openings. None older than 30 days.
Newest first
Cloud Operations Engineer
about 9 hours agoPython · GCP · Kubernetes
3 years of experience, GCP, Cloud infrastructure engineering, Observability tools, Python, Docker, Kubernetes, AI tools, 24/7 on-call rotation
Site Reliability Engineer
1 day agoAWS · Azure · GCP
Kubernetes, Cloud platforms (AWS, Azure, GCP), Infrastructure-as-code (Terraform, AWS CDK), Observability tools (OpenTelemetry, Prometheus, Grafana), CI/CD pipelines, AI/ML exposure, Scripting for automation, Container orchestration
Senior Staff Site Reliability Engineer
1 day agoKubernetes
8 years of experience, Kubernetes-based platforms, AI inference services, Control planes, Platform APIs, GPU scheduling, Vector databases, GitOps, CI/CD, Observability tools, Self-service platforms, Inference-serving frameworks, Open-source contributions
Senior Infrastructure Security Engineer
1 day agoPython · AWS · Azure
Cloud security (GCP, AWS, Azure), Microservices (Containers, Kubernetes), Automation (Bash, Python, Terraform, Ansible), Security standards (NIST, PCI-DSS, SOC2), AI/ML infrastructure security, Network security fundamentals, Security features (AuthN, AuthZ, PKI)
Staff Production Operations Engineer
1 day agoAWS · Kubernetes
10 years of experience, AWS, Kubernetes, AI-assisted tools, Distributed service architecture, Incident management, Automation frameworks, Observability, CI/CD, Full-stack engineering, Production operations
Senior Software Engineer, Platform & Infrastructure - Riot Technology
1 day agoPython · AWS · GCP
3 years of experience, Kubernetes, AWS or GCP, Infrastructure-as-code, CI/CD, GPU compute infrastructure, Python, Distributed systems, MLOps workflows, Multi-node orchestration
Data Center Engineer
1 day ago6 years of experience, Large-scale Data Center Infrastructure, Server and network equipment installation, Real-time requirements management, Root cause analysis, Automation of maintenance actions, Development of infrastructure standards, Cross-functional collaboration, Ability to lift 75 pounds
Infrastructure Solutions Architect - OEM Deployment
2 days agoPython
2 years of experience, Large-scale datacenter rollouts, NVIDIA Cloud Partner integration, TCP/IP networking expertise, Bash scripting, Ansible, Python programming, GPU and DPU technologies, Signal integrity principles, Thermal management, Power distribution, Cabling and rack layout
Monitoring Platform Engineer (Mid-Level)
2 days ago3 years of experience, Splunk, DataDog, AppDynamics, Monitoring solutions, Attention to detail, Written communication
Monitoring Platform Engineer (Senior)
2 days ago6 years of experience, Splunk SOAR, Ansible, Enterprise monitoring tools, Secured environments, Strong troubleshooting skills, ITSM tools
Site Reliability Engineer SME - Senior
2 days ago6 years of experience, OpenTelemetry, Splunk SOAR, Ansible, Event-to-incident workflow design, Dashboard building, Operational reporting
Kubernetes & OpenShift Engineer
2 days agoPython · Kubernetes
6 years of experience, Kubernetes internals, OpenShift internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Cluster monitoring and logging, Container image security
Observability Engineer
2 days ago12 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, Distributed tracing, High-cardinality metrics, SLOs, CI/CD integration, Linux internals, eBPF-based observability
Site Reliability Engineer (SRE)
2 days agoPython · AWS · Azure
10 years of experience, Kubernetes, Prometheus, Grafana, Python, Go, CI/CD pipelines, Chaos engineering, Distributed systems, SLOs and error budgets, Cloud platforms (AWS, Azure, GCP)
Site Reliability Engineer Technical Lead
2 days agoPython · AWS · Azure
8 years of experience, Deep SRE and Systems Expertise, AWS, GCP, Azure, Kubernetes, Python, Automation and Tooling, Dynatrace, Splunk, ELK Stack, Observability and Analysis, Exceptional leadership, Communication skills
Azure Infrastructure Engineer
2 days agoPython · Azure · Kubernetes
6 years of experience, Azure core services, Infrastructure-as-code (Terraform, Bicep, ARM), Azure Kubernetes Service (AKS), Azure DevOps or GitHub Actions, PowerShell, Bash, Python scripting, Cloud security principles, Monitoring and observability strategies, Hybrid cloud or multi-cloud experience, FinOps practices, Regulated environments (HIPAA, PCI-DSS, SOC 2, FedRAMP)
Virtual Platform Engineer
2 days agoAWS · Kubernetes
6 years of experience, vSphere, vSAN, NSX-T, PowerCLI, Tanzu Kubernetes Grid, disaster recovery patterns, troubleshooting across compute, network, and storage layers, VMware Cloud on AWS
Senior Site Reliability Engineer, DGX Cloud
3 days agoPython · Kubernetes
10 years of experience, Kubernetes administration, GPU workloads optimization, Infrastructure automation (Terraform, Ansible), High-level programming (Python, Go), Linux operating systems, SRE principles (SLOs, SLIs), Observability stacks (OpenTelemetry, Prometheus), GPU-accelerated clusters with KubeVirt, Generative-AI techniques, Workflow orchestration (Temporal, Airflow)
Senior Site Reliability Engineer
4 days agoTypeScript · Python · Kubernetes
5 years of experience, Kubernetes (EKS, GKE), Terraform, Pulumi, GitOps (ArgoCD, Flux), Datadog, Prometheus, Grafana, SLI/SLO/error budget fluency, Go, Python, TypeScript, Linux internals, TCP/IP networking, Incident response leadership, Live-service game experience, Service mesh (Istio, Cilium), FinOps, Cloud certifications
Observability Engineer
4 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, distributed tracing, structured logging, Go, Python, Java, high-cardinality metrics, SLOs, error budgets, SRE principles, CI/CD integration, Linux internals, networking, container platforms, Thanos, Mimir, Cortex, Loki, Tempo, eBPF, cost optimization, regulated environments
OpenShift Platform Engineer
4 days agoPython · Kubernetes
6 years of experience, OpenShift, Kubernetes internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Container image security, Cluster monitoring tools, Service mesh (OpenShift Service Mesh, Istio, Linkerd), GitOps workflows (Argo CD, Flux)
Service Mesh Engineer
4 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, Observability
Site Reliability Engineer
4 days agoPython · Azure · Kubernetes
3 years of experience, SLIs/SLOs definition, Multi-tenant SaaS platforms, Datadog, Grafana, Elastic Stack, High-availability architectures, Kubernetes, Python, Bash, Incident response, Cloud experience (Azure)
Site Reliability Engineer
4 days agoGo · AWS · GCP
Kubernetes, AWS, Google Cloud, Golang, CI/CD optimization, Distributed systems management, Monitoring tools (Prometheus, Grafana), Troubleshooting complex infrastructure issues, Disaster recovery strategies, Cloud security practices
Senior Site Reliability Engineer, BCM - DGX Cloud
4 days agoPython · Kubernetes
8 years of experience, Fluency in Python, In-depth knowledge of Linux, Cluster networking proficiency, Experience with Kubernetes, High-performance computing experience, System administration experience with BCM
Senior Site Reliability Engineer in Test, SDET
4 days agoPython · Kubernetes
8 years of experience, GitLab CI, ArgoCD, Kubernetes, Linux, Python, SRE principles, SBOM tooling, AI/ML techniques, test environment management, chaos engineering
Site Reliability Engineer
4 days agoAWS · Azure · Kubernetes
7 years of experience, AWS, Azure, Kubernetes, Cloud-native transformation, Cloud networking, Terraform, Security best practices
Cloud Infrastructure Engineer
5 days agoAzure
4 years of experience, Azure administration, PowerShell scripting, Microsoft Entra ID, DevSecOps collaboration, Cloud Operations experience, Troubleshooting complex infrastructure issues, Automation solutions development, Networking fundamentals
Senior Network Engineer - Operations
5 days agoPython
5 years of experience, Large Ethernet fabrics, Multi-Vendor & NOS (Juniper, Cisco, Arista), Scale-up vs scale-out architectures, Production networks at scale, Automation (Python/Go, Ansible), 100G+ environments, AI, GPU, or HPC exposure
Lead Site Reliability Engineer (Performance & Scalability) | Contract | Remote US
5 days agoSite Reliability Engineering, Performance engineering, Capacity planning, Observability, Production distributed systems, SLOs and SLIs, Load and stress testing, Incident management, Graceful degradation design, Technical performance translation