Remote Site Reliability Engineer jobs open to United States.

United States? Yes. $$$? When they say. Stack? Before the essay.

52 openings. None older than 30 days.

Newest first · page 2

Monitoring Engineer

12 days ago
$75,000 to $85,000 per yearopen to United StatesFull-time

Python · Java

6 years of experience, Prometheus, Grafana, Datadog, New Relic, Splunk, OpenTelemetry, distributed tracing, structured logging, Go, Python, Java, high-cardinality metrics, SLOs, error budgets, SRE principles, CI/CD integration, Linux internals, networking, container platforms, Thanos, Mimir, Cortex, Loki, Tempo, eBPF-based tooling, observability cost optimization, regulated environments

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

Engineering / Platform Professionals

12 days ago
$80 to $160 per houropen to EuropeContract

2 years of experience, Technical document analysis, Cloud infrastructure experience, Attention to detail, Ability to interpret technical concepts, Experience reviewing engineering reports, Familiarity with technical architecture documentation

Weekday AIWeekday AI is a Bengaluru-based recruitment SaaS platform specializing in AI-driven talent sourcing for B2B clients in the recruitment and HR tech industry, primarily serving the US and Indian markets.

Senior Site Reliability Engineer

16 days ago
$142,800 to $178,500 per yearopen to United States + CanadaSeniorFull-time

Python · Kubernetes

6 years of experience, Cloud-native infrastructure, Kubernetes deployment, Terraform, Ansible, CI/CD tooling, Distributed systems troubleshooting, Python, Bash, CUDA-based GPU programs, Security in sensitive environments

PlanetPlanet Labs is a San Francisco-based B2B company specializing in satellite imagery and earth data analytics for applications in agriculture, environmental monitoring, and disaster response.

Senior Software Engineer, Site Reliability & Security

16 days ago
$100,000 to $150,000 per yearopen to United StatesSeniorFull-time

TypeScript · Python · Go

5 years of experience, AWS or GCP, Kubernetes, Docker, Golang, Typescript, Python, Infrastructure as Code, GitOps, Grafana, Prometheus, UNIX shell

WellSaid LabsWellSaid Labs is a Bellevue, Washington-based SaaS company specializing in AI voice generation for enterprise applications, targeting sectors like e-learning, podcasts, and social media.

Senior Staff Site Reliability Engineer - Compute Core Engineering

16 days ago
$200,000 to $322,000 per yearopen to North AmericaStaffFull-time

12 years of experience, eBPF, DPU, Containerization, Distributed Systems, Terraform, Linux Kernel Internals, Network Protocols, Microservices Architecture, Infrastructure as Code

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Customer Reliability Engineer, Infrastructure

17 days ago
$125,000 to $130,000 per yearopen to United StatesFull-time

Python · Kubernetes

5 years of experience, Kubernetes, Cloud infrastructure management, Production distributed systems, Linux, Customer issue handling, Strong communication, DevOps or CI/CD, Python scripting, Troubleshooting

AstronomerAstronomer.io is a B2B managed platform provider specializing in Apache Airflow for data management and workflow automation, operating primarily in the US market.

Staff Site Reliability Engineer

17 days ago
$175,000 to $230,000 per yearopen to United StatesStaffFull-time

AWS

10 years of experience, AWS, Terraform, Puppet, Ansible, Linux administration, Containerization, PCI DSS compliance, Infrastructure as code, Scripting for automation, Technical leadership

PayJunctionPayJunction is a B2B SaaS payment processing company headquartered in Santa Barbara, specializing in flexible payment solutions for businesses and developers across various industries.

Cloud Systems Engineer

18 days ago
$160,000 to $190,000 per yearopen to United StatesFull-time

Python · AWS · Azure

5 years of experience, AWS, Azure, Microsoft 365, Terraform, PowerShell, Bash, Python, IAM, Security compliance, Linux, Windows Server, Defense sector experience, Automation of workflows, Documentation skills, AI platform administration

DEFCON AIDEFCON AI is a defense technology company specializing in AI-powered SaaS solutions for military logistics and optimization, headquartered in the U.S. and primarily serving the Department of Defense.

NCX Senior Engineer

21 days ago
$224,000 to $356,500 per yearopen to United StatesSeniorFull-time

Kubernetes

8 years of experience, GPU infrastructure management, NVIDIA technologies, Kubernetes expertise, Infrastructure observability tools, Automation for lifecycle management, Linux-based distributed systems, Collaboration with cloud partners, Failure modes in distributed AI workloads

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Staff Production Operations Engineer

24 days ago
$199,400 to $299,200 per yearopen to United StatesStaffFull-time

AWS · Kubernetes

10 years of experience, AWS, Kubernetes, AI-assisted tools, Distributed service architecture, Incident management, Automation frameworks, Observability, CI/CD, Full-stack engineering, Production operations

sonyinteractiveentertainmentglobalSony Interactive Entertainment is a San Mateo-based global video game and digital entertainment company, primarily B2C, known for the PlayStation brand and its innovative gaming hardware, software, and network services.

Senior Software Engineer, Platform & Infrastructure - Riot Technology

24 days ago
salary not disclosedopen to United StatesSeniorFull-time

Python · AWS · GCP

3 years of experience, Kubernetes, AWS or GCP, Infrastructure-as-code, CI/CD, GPU compute infrastructure, Python, Distributed systems, MLOps workflows, Multi-node orchestration

Riot GamesRiot Games is a Los Angeles-based video game developer and publisher specializing in competitive multiplayer esports titles, operating globally with a B2C business model.

Site Reliability Engineer

27 days ago
salary not disclosedopen to WorldwideFull-time

Python · Azure · Kubernetes

3 years of experience, SLIs/SLOs definition, Multi-tenant SaaS platforms, Datadog, Grafana, Elastic Stack, High-availability architectures, Kubernetes, Python, Bash, Incident response, Cloud experience (Azure)

HostPapaHostPapa is a Canadian-based web hosting company offering B2B and B2C solutions, including shared, reseller, and VPS hosting services, with a focus on small businesses and a global presence.

Senior Site Reliability Engineer, BCM - DGX Cloud

27 days ago
$208,000 to $333,500 per yearopen to United StatesSeniorFull-time

Python · Kubernetes

8 years of experience, Fluency in Python, In-depth knowledge of Linux, Cluster networking proficiency, Experience with Kubernetes, High-performance computing experience, System administration experience with BCM

NVIDIANVIDIA Corporation is a Santa Clara-based technology company specializing in designing GPUs and AI solutions for gaming, professional visualization, and cloud services, operating in both B2B and B2C markets globally.

Site Reliability Engineer

27 days ago
salary not disclosedopen to United StatesContract

AWS · Azure · Kubernetes

7 years of experience, AWS, Azure, Kubernetes, Cloud-native transformation, Cloud networking, Terraform, Security best practices

ASCENDINGASCENDING is a B2B technology services company specializing in AI-driven solutions and cloud transformation strategies, headquartered in an unspecified location, serving various industries including finance, healthcare, and education.

Cloud Infrastructure Engineer

28 days ago
$125,000 to $135,000 per yearopen to United StatesFull-time

Azure

4 years of experience, Azure administration, PowerShell scripting, Microsoft Entra ID, DevSecOps collaboration, Cloud Operations experience, Troubleshooting complex infrastructure issues, Automation solutions development, Networking fundamentals

AIP PublishingAIP Publishing is a Melville, NY-based B2B scholarly publisher specializing in peer-reviewed journals and resources in the physical sciences, serving a global community of researchers and academic institutions.

Senior Network Engineer - Operations

28 days ago
salary not disclosedopen to United StatesSeniorFull-time

Python

5 years of experience, Large Ethernet fabrics, Multi-Vendor & NOS (Juniper, Cisco, Arista), Scale-up vs scale-out architectures, Production networks at scale, Automation (Python/Go, Ansible), 100G+ environments, AI, GPU, or HPC exposure

TensorWaveTensorWave is a Las Vegas-based B2B cloud computing provider specializing in AI and high-performance computing infrastructure, utilizing AMD Instinct GPUs to deliver scalable solutions for enterprises and AI researchers.

Lab Support Engineer 2, Hopkinton, MA

29 days ago
$83,045 to $107,470 per yearopen to United StatesFull-time

Python

Python, Bash, PowerShell, Dell PowerEdge, PowerStore, Linux administration, VMware, vCenter, datacenter operations, networking fundamentals, problem-solving skills

DellDell Technologies designs and develops hardware and software for infrastructure, enterprise solutions, and data management.

Sr Engineer, SRE TechOps CICD (Remote)

29 days ago
USD 140,000–215,000 per yearopen to United StatesSeniorFull-time

Python · Kubernetes

What You’ll Need: Must be eligible for CJIS clearance (requires U.S. citizenship or Green Card/permanent resident status).; What You’ll Need: + years of experience working in large-scale production SRE or infrastructure environments.; What You’ll Need: + years of experience leveraging and integrating AI-assisted workflows to increase engineering efficiency.; What You’ll Need: Bachelor's degree in computer science or another highly technical, scientific discipline, or equivalent work experience.; What You’ll Need: On-premise and cloud expertise deploying and operating CI/CD tools (Bazel, Jenkins, GitLab CI, GitHub Actions), IaC provisioning (Ansible, Chef, Puppet, Salt, Terraform), source code management (Bitbucket, GitHub, GitLab), and monitoring/observability platforms (Datadog, Grafana, Humio/LogScale, Honeycomb, New Relic, Prometheus, Splunk).; What You’ll Need: Experience creating, deploying, operating, and scaling applications on Kubernetes.; What You’ll Need: Extensive experience deploying and managing data infrastructure at scale (Cassandra, Postgres, MySQL, MongoDB, OpenSearch, Kafka, Redis/Valkey).; What You’ll Need: Proficiency in common scripting languages (Python, Go, Bash, PowerShell).; What You’ll Need: Experience with storage technologies (SAN, NAS, NFS, Object Storage).; What You’ll Need: Experience architecting and deploying big data systems.; What You’ll Need: Security-first mindset with a working understanding of cybersecurity principles.; What You’ll Need: Proven ability to make well-informed, timely decisions under ambiguity.

CrowdStrikeCrowdStrike is an Austin-based B2B cybersecurity leader providing advanced cloud-native solutions for endpoint security and threat intelligence to a global market.

Staff Platform Engineer, Core Cloud Platform

30 days ago
$223,100 to $305,000 per yearopen to United States + Canada + United Kingdom + Singapore + India + Ireland + FinlandStaffFull-time

Python · Go · AWS

Kubernetes expertise, AWS cloud proficiency, Production SaaS systems experience, Networking and service mesh knowledge, Operational troubleshooting skills, Python and Golang proficiency, CI/CD practices advocacy, Technical roadmap design and leadership

AlphaSenseAlphaSense is a New York City-based B2B fintech platform specializing in AI-driven market intelligence and search solutions for financial institutions and top companies globally.

Senior Cloud / DevSecOps Engineer (R-00200)

30 days ago
salary not disclosedopen to United StatesSeniorFull-time

Python · AWS · Kubernetes

AWS GovCloud, AWS CloudFormation, Terraform, Python, Bash, PowerShell, Ansible, CI/CD, Kubernetes, Amazon EKS, Amazon ECS, Cloud Security, DoD RMF, DISA STIGs, Infrastructure as Code, Hybrid Cloud Engineering, Windows Server, RHEL, Technical Leadership

True Zero TechnologiesTrue Zero Technologies is a veteran-owned cybersecurity consulting firm headquartered in Fairfax, VA, specializing in B2B services for federal agencies and the public sector.

Platform Infrastructure Engineer

30 days ago
$100,000 to $150,000 per yearopen to United StatesFull-time

Python · Kubernetes

6 years of experience, OpenShift, Kubernetes, Linux administration, Infrastructure-as-code (Ansible, Terraform, Helm), CI/CD pipelines (Tekton, Jenkins, Argo CD), Scripting (Bash, Python, Go), Cluster monitoring and logging, Container image security, Service mesh (Istio, Linkerd), GitOps workflows (Argo CD, Flux)

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.

Service Infrastructure Engineer

30 days ago
$100,000 to $150,000 per yearopen to United StatesFull-time

Python · Kubernetes

6 years of experience, Istio, Linkerd, Envoy, Kubernetes, mTLS, PKI, Go, Python, Distributed tracing, Service mesh production experience

Bright Vision TechnologiesBright Vision Technologies is a New Jersey-based IT staffing firm specializing in placing technical professionals in software development roles across the U.S. government and enterprise sectors.