Senior Site Reliability Engineer - India
JumpCloud · full-time · posted 29 Sep
You came here asking one thing. Salary?
$$$ → ???
They didn't say.
*The posting gives you no number to evaluate. That absence is part of the offer.
Second question — can you apply from where you sit?
Open to India
Compare it with where you sit.
→ inferred locally from your browser clock. nothing is stored.
The actual job
Senior Site Reliability Engineer - India
JumpCloud
- Seniority
- Senior
- Top skills
- Python · Go · AWS
- Engagement
- Full-time
- Posted
- 5 days ago
What the posting asks for
- 8+ years SRE/DevOps/platform
- Kubernetes (EKS)
- Terraform
- AWS/GCP
- Python or Golang
Employer text
The posting, in its own words
About JumpCloud®
JumpCloud is Intelligent, Secure IT.
What You’ll Be Doing:
-
Architect, scale, and continuously improve the reliability, availability, and performance of JumpCloud’s multi-region microservices, APIs, and authentication infrastructure (AWS/GCP).
-
Architect, build, and maintain Disaster Recovery (DR) process, multi-region failover automation, and business continuity strategies to ensure rapid recovery against strict RTO and RPO objectives.
-
Lead the design and enforcement of SLIs, SLOs, and Error Budget frameworks across multi-disciplinary engineering teams.
-
Drive end-to-end observability strategy using Datadog, implementing actionable Golden Signals monitoring to drastically reduce MTTD/MTTR and eliminate alert fatigue.
-
Lead on-call escalation, major incident management, and drive strict adherence to 99.99% availability SLAs.
-
Facilitate blameless post-incident reviews, executing systemic root-cause remediations to prevent recurring failure modes.
-
Architect, manage, and scale production Kubernetes (EKS) clusters, implementing advanced GitOps workflows (Argo CD, Kargo) and deployment patterns.
-
Design and maintain modular, enterprise-grade Infrastructure-as-Code using Terraform across multi-account, multi-region cloud environments.
-
Design, build, and maintain interactive FinOps and cost-optimization dashboards to provide engineering and leadership teams with actionable insights into multi-cloud spend, unit economics, and resource utilization.
-
Eliminate complex operational toil by writing production-grade Python or Go tooling, platform automation, and custom integrations.
-
Champion AI-assisted software development workflows (Cursor, Claude Code, GitHub Copilot) to accelerate automation, runbook creation, and incident triage across the team.
-
Author operational runbooks, architecture decision records, and mentor mid-level/junior engineers to raise the overall technical bar.
We’re Looking For:
-
8+ years of professional software engineering experience in SRE, DevOps, or Platform Engineering operating 24/7 mission-critical, highly available distributed systems.
-
Bachelor's degree in Computer Science, Software Engineering, or equivalent technical discipline.
-
Strong Python/Go Capabilities: Advanced software engineering skill set for writing internal SRE platforms, tools, and API integrations.
-
Deep Kubernetes Expertise: Hands-on experience with production EKS/GKE cluster lifecycles, ingress/egress, networking, RBAC, and GitOps tooling (Argo CD).
-
Advanced IaC & AWS/GCP: Deep Terraform proficiency (module architecture, state management refactoring) across complex multi-account AWS environments (IAM, VPCs, Transit Gateway, ALB/NLB, Route53).
-
FinOps & Cost Optimization Leadership: Demonstrated experience driving cloud cost-efficiency strategies, resource right-sizing, cost-allocation tagging, workload optimization, and building FinOps dashboards to embed financial accountability into engineering workflows.
-
Disaster Recovery & High Availability: Proven background in designing and testing multi-region Disaster Recovery architectures, automating failover systems, and monitoring recovery health via DR dashboards.
-
Observability & Reliability Architecture: Track record of defining SLI/SLOs, managing PagerDuty schedules, and optimizing production observability platforms.
-
Experience designing and operating enterprise service meshes (Istio, Linkerd, or similar) and production ingress/proxy systems (HAProxy, NGINX, or similar).
-
Technical Mentorship: Demonstrated ability to lead technical discussions, write architectural design docs/RFCs, and mentor engineering peers.
-
Strong problem-solving, communication, and collaboration skills with a passion for solving complex distributed systems challenges at scale.
-
A strong team player who helps us live by our core values: building connections, thinking big, and getting 1% better every day.
Preferred Qualifications:
-
Basic understanding of chaos engineering principles or testing resilience in staging/production.
-
Experience with secrets management architectures (Vault, AWS Secrets Manager, External Secrets Operator, Cert-Manager).
-
Background in DevSecOps practices, service meshes (Istio), and automated vulnerability remediation within cloud infrastructure code.
-
Background supporting identity services, IAM, enterprise directory platforms, or security-focused SaaS solutions.