Remote Site Reliability Engineer jobs.
“Remote”? Depends where. $$$? When they say. Stack? Before the essay.
Where can you apply from?
99 openings. None older than 30 days.
Newest first
Cloud Systems Engineer
about 9 hours agoPython · AWS · Azure
5 years of experience, AWS, Azure, Microsoft 365, Terraform, PowerShell, Bash, Python, IAM, Security compliance, Linux, Windows Server, Defense sector experience, Automation of workflows, Documentation skills, AI platform administration
Senior Site Reliability Engineer | SRE (Remote, EU/CET)
about 9 hours agoAWS
AWS, Infrastructure as code, Containerized workloads, Observability, Cloud security, Vulnerability management, SOC 2, ISO 27001, Automation, AI-assisted development, Production incident management
Staff Site Reliability Engineer
about 9 hours agoPython · AWS · Kubernetes
7 years of experience, AWS, Kubernetes, Terraform, SLO frameworks, Incident response programs, Observability platforms, Python, Go, AI tools, Healthcare compliance (HIPAA), Mentorship
Senior Database Reliability Engineer
about 9 hours agoGo · SQL · AWS
6 years of experience, Kubernetes, AWS, RDS (MySQL/Postgres), Golang, SQL-based RDBMS, Observability tools, Distributed systems design patterns, AI tooling, Microservices architecture, Engineering best practices
Senior R&D Site Reliability Engineer
2 days agoPython · AWS
Terraform, AWS architecture, CI/CD (Gitlab runner), Linux system management, Python scripting, Observability tools (Prometheus, Grafana, ELK Stack), Automation tools development, Analytical and problem-solving skills, English proficiency (IELTS 6.5), Mandarin Chinese (plus)
Senior Site Reliability Engineer
2 days agoPython · Kubernetes
8 years of experience, Database engineering, MySQL, MSSQL, Oracle, Query optimization, Database-as-a-Service, Python, Go, Kubernetes, Hybrid database replication, Observability tools
Senior Site Reliability Engineering - Storage
2 days agoPython · Kubernetes · Docker
12 years of experience, NAS, SAN, Object Storage, SRE concepts, Infrastructure as Code, Terraform, Ansible, Docker, Kubernetes, Python, Go, Shell, AI/ML storage, data analytics, debugging distributed systems, mentoring engineers
Cloud Infrastructure Network Engineer
2 days agoKubernetes
6 years of experience, Cloud networking experience, VPC/VNet design, Hybrid connectivity, Infrastructure-as-code (Terraform), Cloud security and network controls, Multi-cloud networking, Kubernetes networking, SD-WAN familiarity, eBPF-based networking tools
OpenShift Administrator
2 days agoPython · Kubernetes
6 years of experience, OpenShift, Kubernetes internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Container image security, Cluster monitoring tools, Service mesh (OpenShift Service Mesh, Istio, Linkerd), GitOps workflows (Argo CD, Flux)
Service Mesh Architect
2 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, Cilium, SPIFFE/SPIRE, Zero-trust networking
Site Reliability Engineer Technical Lead
2 days agoPython · AWS · Azure
8 years of experience, Deep SRE and Systems Expertise, AWS, GCP, Azure, Kubernetes, Python, Automation and Tooling, Dynatrace, Splunk, ELK Stack, Observability and Analysis, Exceptional leadership, Communication skills
NCX Senior Engineer
3 days agoKubernetes
8 years of experience, GPU infrastructure management, NVIDIA technologies, Kubernetes expertise, Infrastructure observability tools, Automation for lifecycle management, Linux-based distributed systems, Collaboration with cloud partners, Failure modes in distributed AI workloads
Telemetry Engineer
5 days agoPython · Java
6 years of experience, Prometheus, Grafana, OpenTelemetry, Datadog, New Relic, Splunk, Go, Python, Java, SRE principles, incident management, distributed tracing, structured logging, Linux, containers, CI/CD
VMware Infrastructure Engineer
5 days agoAWS · Kubernetes
6 years of experience, vSphere, vSAN, NSX-T, PowerCLI, Tanzu Kubernetes Grid, disaster recovery, VMware Cloud on AWS, Aria Operations
DevOps & SRE Engineer
5 days agoPython · Java · Kubernetes
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Observability tooling, CI/CD pipelines, SLOs and error budgets, Chaos engineering, Cloud platforms
Site Reliability Engineer
5 days agoPython · AWS · Azure
Python, Docker, Kubernetes, Google Cloud Platform, AWS, Azure, Terraform, Cloudformation, SaltStack, Ansible, Grafana, Prometheus, ELK stack, Game development passion, Independent work, Collaborative spirit
Systems Observability Specialist
6 days ago6 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, SLOs, High-cardinality metrics, CI/CD integration, Linux internals, Container platforms, Observability cost optimization
Kubernetes Service Engineer
6 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, mTLS, Kubernetes networking, Go, Python, Distributed tracing, Service mesh architecture, Zero-trust networking
Platform Reliability Engineer
6 days agoPython · Java · Kubernetes
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Observability tooling, CI/CD pipelines, Distributed system design, SLOs and error budgets, Chaos engineering, Cloud platforms
Site Observability Engineer
6 days ago6 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, Distributed tracing, High-cardinality metrics, SLOs, CI/CD integration, Linux internals, Container platforms
Site Reliability Engineer
6 days agoAWS · Azure · GCP
Kubernetes, Cloud platforms (AWS, Azure, GCP), Infrastructure-as-code (Terraform, AWS CDK), Observability tools (OpenTelemetry, Prometheus, Grafana), CI/CD pipelines, AI/ML exposure, Scripting for automation, Container orchestration
Senior Staff Site Reliability Engineer
6 days agoKubernetes
8 years of experience, Kubernetes-based platforms, AI inference services, Control planes, Platform APIs, GPU scheduling, Vector databases, GitOps, CI/CD, Observability tools, Self-service platforms, Inference-serving frameworks, Open-source contributions
Senior Infrastructure Security Engineer
6 days agoPython · AWS · Azure
Cloud security (GCP, AWS, Azure), Microservices (Containers, Kubernetes), Automation (Bash, Python, Terraform, Ansible), Security standards (NIST, PCI-DSS, SOC2), AI/ML infrastructure security, Network security fundamentals, Security features (AuthN, AuthZ, PKI)
Staff Production Operations Engineer
6 days agoAWS · Kubernetes
10 years of experience, AWS, Kubernetes, AI-assisted tools, Distributed service architecture, Incident management, Automation frameworks, Observability, CI/CD, Full-stack engineering, Production operations
Senior Software Engineer, Platform & Infrastructure - Riot Technology
7 days agoPython · AWS · GCP
3 years of experience, Kubernetes, AWS or GCP, Infrastructure-as-code, CI/CD, GPU compute infrastructure, Python, Distributed systems, MLOps workflows, Multi-node orchestration
Data Center Engineer
7 days ago6 years of experience, Large-scale Data Center Infrastructure, Server and network equipment installation, Real-time requirements management, Root cause analysis, Automation of maintenance actions, Development of infrastructure standards, Cross-functional collaboration, Ability to lift 75 pounds
Infrastructure Solutions Architect - OEM Deployment
7 days agoPython
2 years of experience, Large-scale datacenter rollouts, NVIDIA Cloud Partner integration, TCP/IP networking expertise, Bash scripting, Ansible, Python programming, GPU and DPU technologies, Signal integrity principles, Thermal management, Power distribution, Cabling and rack layout
Monitoring Platform Engineer (Mid-Level)
7 days ago3 years of experience, Splunk, DataDog, AppDynamics, Monitoring solutions, Attention to detail, Written communication
Monitoring Platform Engineer (Senior)
7 days ago6 years of experience, Splunk SOAR, Ansible, Enterprise monitoring tools, Secured environments, Strong troubleshooting skills, ITSM tools
Site Reliability Engineer SME - Senior
7 days ago6 years of experience, OpenTelemetry, Splunk SOAR, Ansible, Event-to-incident workflow design, Dashboard building, Operational reporting