Remote Site Reliability Engineer jobs open to United States.
United States? Yes. $$$? When they say. Stack? Before the essay.
52 openings. None older than 30 days.
Newest first · page 2
Monitoring Engineer
12 days agoPython · Java
6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, distributed tracing, structured logging, Go, Python, Java, high-cardinality metrics, SLOs, error budgets, SRE principles, CI/CD integration, Linux internals, networking, container platforms, Thanos, Mimir, Cortex, Loki, Tempo, eBPF-based tooling, observability cost optimization, regulated environments
Engineering / Platform Professionals
12 days ago2 years of experience, Technical document analysis, Cloud infrastructure experience, Attention to detail, Ability to interpret technical concepts, Experience reviewing engineering reports, Familiarity with technical architecture documentation
Senior Site Reliability Engineer
16 days agoPython · Kubernetes
6 years of experience, Cloud-native infrastructure, Kubernetes deployment, Terraform, Ansible, CI/CD tooling, Distributed systems troubleshooting, Python, Bash, CUDA-based GPU programs, Security in sensitive environments
Senior Software Engineer, Site Reliability & Security
16 days agoTypeScript · Python · Go
5 years of experience, AWS or GCP, Kubernetes, Docker, Golang, Typescript, Python, Infrastructure as Code, GitOps, Grafana, Prometheus, UNIX shell
Senior Staff Site Reliability Engineer - Compute Core Engineering
16 days ago12 years of experience, eBPF, DPU, Containerization, Distributed Systems, Terraform, Linux Kernel Internals, Network Protocols, Microservices Architecture, Infrastructure as Code
Customer Reliability Engineer, Infrastructure
17 days agoPython · Kubernetes
5 years of experience, Kubernetes, Cloud infrastructure management, Production distributed systems, Linux, Customer issue handling, Strong communication, DevOps or CI/CD, Python scripting, Troubleshooting
Staff Site Reliability Engineer
17 days agoAWS
10 years of experience, AWS, Terraform, Puppet, Ansible, Linux administration, Containerization, PCI DSS compliance, Infrastructure as code, Scripting for automation, Technical leadership
Cloud Systems Engineer
18 days agoPython · AWS · Azure
5 years of experience, AWS, Azure, Microsoft 365, Terraform, PowerShell, Bash, Python, IAM, Security compliance, Linux, Windows Server, Defense sector experience, Automation of workflows, Documentation skills, AI platform administration
NCX Senior Engineer
21 days agoKubernetes
8 years of experience, GPU infrastructure management, NVIDIA technologies, Kubernetes expertise, Infrastructure observability tools, Automation for lifecycle management, Linux-based distributed systems, Collaboration with cloud partners, Failure modes in distributed AI workloads
Staff Production Operations Engineer
24 days agoAWS · Kubernetes
10 years of experience, AWS, Kubernetes, AI-assisted tools, Distributed service architecture, Incident management, Automation frameworks, Observability, CI/CD, Full-stack engineering, Production operations
Senior Software Engineer, Platform & Infrastructure - Riot Technology
24 days agoPython · AWS · GCP
3 years of experience, Kubernetes, AWS or GCP, Infrastructure-as-code, CI/CD, GPU compute infrastructure, Python, Distributed systems, MLOps workflows, Multi-node orchestration
Site Reliability Engineer
27 days agoPython · Azure · Kubernetes
3 years of experience, SLIs/SLOs definition, Multi-tenant SaaS platforms, Datadog, Grafana, Elastic Stack, High-availability architectures, Kubernetes, Python, Bash, Incident response, Cloud experience (Azure)
Senior Site Reliability Engineer, BCM - DGX Cloud
27 days agoPython · Kubernetes
8 years of experience, Fluency in Python, In-depth knowledge of Linux, Cluster networking proficiency, Experience with Kubernetes, High-performance computing experience, System administration experience with BCM
Site Reliability Engineer
27 days agoAWS · Azure · Kubernetes
7 years of experience, AWS, Azure, Kubernetes, Cloud-native transformation, Cloud networking, Terraform, Security best practices
Cloud Infrastructure Engineer
28 days agoAzure
4 years of experience, Azure administration, PowerShell scripting, Microsoft Entra ID, DevSecOps collaboration, Cloud Operations experience, Troubleshooting complex infrastructure issues, Automation solutions development, Networking fundamentals
Senior Network Engineer - Operations
28 days agoPython
5 years of experience, Large Ethernet fabrics, Multi-Vendor & NOS (Juniper, Cisco, Arista), Scale-up vs scale-out architectures, Production networks at scale, Automation (Python/Go, Ansible), 100G+ environments, AI, GPU, or HPC exposure
Lab Support Engineer 2, Hopkinton, MA
29 days agoPython
Python, Bash, PowerShell, Dell PowerEdge, PowerStore, Linux administration, VMware, vCenter, datacenter operations, networking fundamentals, problem-solving skills
Sr Engineer, SRE TechOps CICD (Remote)
29 days agoPython · Kubernetes
What You’ll Need: Must be eligible for CJIS clearance (requires U.S. citizenship or Green Card/permanent resident status).; What You’ll Need: + years of experience working in large-scale production SRE or infrastructure environments.; What You’ll Need: + years of experience leveraging and integrating AI-assisted workflows to increase engineering efficiency.; What You’ll Need: Bachelor's degree in computer science or another highly technical, scientific discipline, or equivalent work experience.; What You’ll Need: On-premise and cloud expertise deploying and operating CI/CD tools (Bazel, Jenkins, GitLab CI, GitHub Actions), IaC provisioning (Ansible, Chef, Puppet, Salt, Terraform), source code management (Bitbucket, GitHub, GitLab), and monitoring/observability platforms (Datadog, Grafana, Humio/LogScale, Honeycomb, New Relic, Prometheus, Splunk).; What You’ll Need: Experience creating, deploying, operating, and scaling applications on Kubernetes.; What You’ll Need: Extensive experience deploying and managing data infrastructure at scale (Cassandra, Postgres, MySQL, MongoDB, OpenSearch, Kafka, Redis/Valkey).; What You’ll Need: Proficiency in common scripting languages (Python, Go, Bash, PowerShell).; What You’ll Need: Experience with storage technologies (SAN, NAS, NFS, Object Storage).; What You’ll Need: Experience architecting and deploying big data systems.; What You’ll Need: Security-first mindset with a working understanding of cybersecurity principles.; What You’ll Need: Proven ability to make well-informed, timely decisions under ambiguity.
Staff Platform Engineer, Core Cloud Platform
30 days agoPython · Go · AWS
Kubernetes expertise, AWS cloud proficiency, Production SaaS systems experience, Networking and service mesh knowledge, Operational troubleshooting skills, Python and Golang proficiency, CI/CD practices advocacy, Technical roadmap design and leadership
Senior Cloud / DevSecOps Engineer (R-00200)
30 days agoPython · AWS · Kubernetes
AWS GovCloud, AWS CloudFormation, Terraform, Python, Bash, PowerShell, Ansible, CI/CD, Kubernetes, Amazon EKS, Amazon ECS, Cloud Security, DoD RMF, DISA STIGs, Infrastructure as Code, Hybrid Cloud Engineering, Windows Server, RHEL, Technical Leadership
Platform Infrastructure Engineer
30 days agoPython · Kubernetes
6 years of experience, OpenShift, Kubernetes, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Cluster monitoring and logging, Container image security, Service mesh (Istio, Linkerd), GitOps workflows (Argo CD, Flux)
Service Infrastructure Engineer
30 days agoPython · Kubernetes
6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, Service mesh production experience