Remote Site Reliability Engineer jobs.
“Remote”? Depends where. $$$? When they say. Stack? Before the essay.
Where can you apply from?
104 openings. None older than 30 days.
Newest first
Customer Reliability Engineer, Infrastructure
about 9 hours agoPython · Kubernetes
5 years of experience, Kubernetes, Cloud infrastructure management, Production distributed systems, Linux, Customer issue handling, Strong communication, DevOps or CI/CD, Python scripting, Troubleshooting
Cloud Systems Engineer
1 day agoPython · AWS · Azure
5 years of experience, AWS, Azure, Microsoft 365, Terraform, PowerShell, Bash, Python, IAM, Security compliance, Linux, Windows Server, Defense sector experience, Automation of workflows, Documentation skills, AI platform administration
Senior Site Reliability Engineer | SRE (Remote, EU/CET)
1 day agoAWS
AWS, Infrastructure as code, Containerized workloads, Observability, Cloud security, Vulnerability management, SOC 2, ISO 27001, Automation, AI-assisted development, Production incident management
Staff Site Reliability Engineer
1 day agoPython · AWS · Kubernetes
7 years of experience, AWS, Kubernetes, Terraform, SLO frameworks, Incident response programs, Observability platforms, Python, Go, AI tools, Healthcare compliance (HIPAA), Mentorship
Senior Database Reliability Engineer
1 day agoGo · SQL · AWS
6 years of experience, Kubernetes, AWS, RDS (MySQL/Postgres), Golang, SQL-based RDBMS, Observability tools, Distributed systems design patterns, AI tooling, Microservices architecture, Engineering best practices
Senior Staff SRE – Compute Platform
1 day agoPython · Kubernetes
10 years of experience, Kubernetes administration, Bare-metal infrastructure, Python or Go, Infrastructure as Code, SRE and observability, HPC or AI infrastructure, VMware vSphere, Generative AI applications, Secure operational platforms, High-impact infrastructure projects
Senior Site Reliability Engineer
1 day agoPython
12 years of experience, eBPF, XDP, Terraform, Go, Python, Linux OS, DNS, LDAP, containerization architecture, distributed-systems infrastructure, network protocols
Kafka Engineer
2 days agoKubernetes
7 years of experience, Kafka internals, Kafka security, Kafka Connect, Infrastructure-as-code, Observability tooling, DevOps practices, Kubernetes experience, Streaming frameworks
Monitoring Engineer
2 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, distributed tracing, structured logging, Go, Python, Java, high-cardinality metrics, SLOs, error budgets, SRE principles, CI/CD integration, Linux internals, networking, container platforms
Observability Engineer
2 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, Go, Python, Java, High-cardinality metrics, SLOs, CI/CD integration, Linux internals, eBPF tooling
OpenShift Platform Engineer
2 days agoKubernetes
8 years of experience, OpenShift, Kubernetes internals, GitOps, CI/CD pipelines, Linux administration, Infrastructure-as-code, Container security, Disaster recovery strategies, Multi-tenant platform design, Monitoring and observability tools
Reliability Engineer
2 days agoPython · Java · Kubernetes
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Observability tooling, CI/CD pipelines, Distributed system design, SLOs and error budgets, Chaos engineering, Cloud platforms
Apache Kafka Developer
2 days agoKubernetes
6 years of experience, Kafka internals, Kafka security, Kafka Connect, Infrastructure-as-code, Observability tooling, Confluent Certified Administrator, Kafka on Kubernetes, Managed Kafka services, Stream processing frameworks, Data governance
Senior R&D Site Reliability Engineer
3 days agoPython · AWS
Terraform, AWS architecture, CI/CD (Gitlab runner), Linux system management, Python scripting, Observability tools (Prometheus, Grafana, ELK Stack), Automation tools development, Analytical and problem-solving skills, English proficiency (IELTS 6.5), Mandarin Chinese (plus)
Senior Site Reliability Engineer
3 days agoPython · Kubernetes
8 years of experience, Database engineering, MySQL, MSSQL, Oracle, Query optimization, Database-as-a-Service, Python, Go, Kubernetes, Hybrid database replication, Observability tools
Senior Site Reliability Engineering - Storage
3 days agoPython · Kubernetes · Docker
12 years of experience, NAS, SAN, Object Storage, SRE concepts, Infrastructure as Code, Terraform, Ansible, Docker, Kubernetes, Python, Go, Shell, AI/ML storage, data analytics, debugging distributed systems, mentoring engineers
Cloud Infrastructure Network Engineer
3 days agoKubernetes
6 years of experience, Cloud networking experience, VPC/VNet design, Hybrid connectivity, Infrastructure-as-code (Terraform), Cloud security and network controls, Multi-cloud networking, Kubernetes networking, SD-WAN familiarity, eBPF-based networking tools
OpenShift Administrator
3 days agoPython · Kubernetes
6 years of experience, OpenShift, Kubernetes internals, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Container image security, Cluster monitoring tools, Service mesh (OpenShift Service Mesh, Istio, Linkerd), GitOps workflows (Argo CD, Flux)
Service Mesh Architect
3 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, Cilium, SPIFFE/SPIRE, Zero-trust networking
Site Reliability Engineer Technical Lead
3 days agoPython · AWS · Azure
8 years of experience, Deep SRE and Systems Expertise, AWS, GCP, Azure, Kubernetes, Python, Automation and Tooling, Dynatrace, Splunk, ELK Stack, Observability and Analysis, Exceptional leadership, Communication skills
NCX Senior Engineer
4 days agoKubernetes
8 years of experience, GPU infrastructure management, NVIDIA technologies, Kubernetes expertise, Infrastructure observability tools, Automation for lifecycle management, Linux-based distributed systems, Collaboration with cloud partners, Failure modes in distributed AI workloads
Telemetry Engineer
6 days agoPython · Java
6 years of experience, Prometheus, Grafana, OpenTelemetry, Datadog, New Relic, Splunk, Go, Python, Java, SRE principles, incident management, distributed tracing, structured logging, Linux, containers, CI/CD
VMware Infrastructure Engineer
6 days agoAWS · Kubernetes
6 years of experience, vSphere, vSAN, NSX-T, PowerCLI, Tanzu Kubernetes Grid, disaster recovery, VMware Cloud on AWS, Aria Operations
DevOps & SRE Engineer
6 days agoPython · Java · Kubernetes
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Observability tooling, CI/CD pipelines, SLOs and error budgets, Chaos engineering, Cloud platforms
Site Reliability Engineer
6 days agoPython · AWS · Azure
Python, Docker, Kubernetes, Google Cloud Platform, AWS, Azure, Terraform, Cloudformation, SaltStack, Ansible, Grafana, Prometheus, ELK stack, Game development passion, Independent work, Collaborative spirit
Systems Observability Specialist
7 days ago6 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, SLOs, High-cardinality metrics, CI/CD integration, Linux internals, Container platforms, Observability cost optimization
Kubernetes Service Engineer
7 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, mTLS, Kubernetes networking, Go, Python, Distributed tracing, Service mesh architecture, Zero-trust networking
Platform Reliability Engineer
7 days agoPython · Java · Kubernetes
6 years of experience, Python, Go, Java, Kubernetes, Linux at scale, Observability tooling, CI/CD pipelines, Distributed system design, SLOs and error budgets, Chaos engineering, Cloud platforms
Site Observability Engineer
7 days ago6 years of experience, Prometheus, Grafana, Datadog, OpenTelemetry, Distributed tracing, High-cardinality metrics, SLOs, CI/CD integration, Linux internals, Container platforms
Site Reliability Engineer
7 days agoAWS · Azure · GCP
Kubernetes, Cloud platforms (AWS, Azure, GCP), Infrastructure-as-code (Terraform, AWS CDK), Observability tools (OpenTelemetry, Prometheus, Grafana), CI/CD pipelines, AI/ML exposure, Scripting for automation, Container orchestration