Remote Site Reliability Engineer jobs open to India.
India? Yes. $$$? When they say. Stack? Before the essay.
11 openings. None older than 30 days.
Newest first
Senior Staff SRE – Compute Platform
3 days agoPython · Kubernetes
10 years of experience, Kubernetes administration, Bare-metal infrastructure, Python or Go, Infrastructure as Code, SRE and observability, HPC or AI infrastructure, VMware vSphere, Generative AI applications, Secure operational platforms, High-impact infrastructure projects
Site Reliability Engineer
9 days agoAWS · Azure · GCP
Kubernetes, Cloud platforms (AWS, Azure, GCP), Infrastructure-as-code (Terraform, AWS CDK), Observability tools (OpenTelemetry, Prometheus, Grafana), CI/CD pipelines, AI/ML exposure, Scripting for automation, Container orchestration
Senior Staff Site Reliability Engineer
9 days agoKubernetes
8 years of experience, Kubernetes-based platforms, AI inference services, Control planes, Platform APIs, GPU scheduling, Vector databases, GitOps, CI/CD, Observability tools, Self-service platforms, Inference-serving frameworks, Open-source contributions
Senior Site Reliability Engineer
12 days agoTypeScript · Python · Kubernetes
5 years of experience, Kubernetes (EKS, GKE), Terraform, Pulumi, GitOps (ArgoCD, Flux), Datadog, Prometheus, Grafana, SLI/SLO/error budget fluency, Go, Python, TypeScript, Linux internals, TCP/IP networking, Incident response leadership, Live-service game experience, Service mesh (Istio, Cilium), FinOps, Cloud certifications
Site Reliability Engineer
12 days agoPython · Azure · Kubernetes
3 years of experience, SLIs/SLOs definition, Multi-tenant SaaS platforms, Datadog, Grafana, Elastic Stack, High-availability architectures, Kubernetes, Python, Bash, Incident response, Cloud experience (Azure)
Staff Platform Engineer, Core Cloud Platform
15 days agoPython · Go · AWS
Kubernetes expertise, AWS cloud proficiency, Production SaaS systems experience, Networking and service mesh knowledge, Operational troubleshooting skills, Python and Golang proficiency, CI/CD practices advocacy, Technical roadmap design and leadership
Senior Staff Site Reliability Engineer
17 days agoPython · AWS · Kubernetes
10 years of experience, Kubernetes, AI/ML infrastructure, OpenTelemetry, AWS CDK, Terraform, Python, Go, Distributed systems debugging, Agentic AI platforms, Automation systems, Technical strategy steering
Site Reliability Engineer
19 days agoAWS · Kubernetes
7 years of experience, Event-driven architecture, Deep AWS, Infrastructure as Code, Kubernetes, Observability and SLOs, Chaos engineering, Distributed systems debugging, Proven technical leadership, AI / MLOps infrastructure, Experience in payments industry
Principal Site Reliability Engineer
22 days agoPython
10 years of experience, Openshift, Nutanix AHV, VMware vSphere, RedHat OpenShift, Python, Go, Ansible, GitHub, 24x7 operations
Senior HPC Cluster Engineer - AI, ML
23 days agoPython · Docker
5 years of experience, AI/HPC advanced job schedulers, Slurm, Centos/RHEL, Ubuntu Linux, Cluster configuration management, Docker, Python, MPI, NVIDIA GPUs, CUDA Programming, InfiniBand, Lustre
Operations Support Engineer
24 days ago5 years of experience, ITIL-aligned frameworks, Second-level support for mission-critical systems, Observability tools, Linux/Unix environments, Container orchestration, CI/CD pipelines, Middleware and integration technologies, Database performance monitoring, System hardening and patching, English fluency