CareerMoonshot

Site Reliability Engineer, Sr. Site Reliability Engineer

Prove · Remote

📍 United States (Remote)💰 $130,000 - 150,000via greenhousePosted 2026-09-14
Apply on company site ↗
Career Moonshot pulls this listing straight from the employer's hiring system — no recruiter middleman, no reposts. Applying takes you directly to Prove.
About Prove  As the world moves to a mobile-first economy, businesses need to modernize how they acquire, engage with and enable consumers. Prove’s phone-centric identity tokenization and passive cryptographic authentication solutions reduce friction, enhance security and privacy across all digital channels, and accelerate revenues while reducing operating expenses and fraud losses. Over 1,000 enterprise customers use Prove’s platform to process 20 billion customer requests annually across industries, including banking, lending, healthcare, gaming, crypto, e-commerce, marketplaces, and payments. For the latest updates from Prove, follow us on LinkedIn. Prove is driving the future of digital identity. We are looking for Provers who know how to make an impact. We’re talking self-starting professionals who thrive in a fast-paced environment, process information quickly, and make intelligent decisions. The work is challenging and requires not only smart but natural curiosity and tenacity. Teamwork is also important to us – we work together and play together.    Prove has big plans, and we’re excited about the future. If this sounds like the place for you – come join our team!  Title: Site Reliability Engineer, Senior Site Reliability Engineer Department: Platform Engineering Reports To: Director, Platform Engineering FLSA Status: Exempt Location: US Remote Job Summary We are seeking Mid to Senior level Site Reliability Engineers to join our Platform Engineering team. In this role, you will be instrumental in designing, implementing, maintaining and deploying highly available complex, scalable and reliable systems leveraging automation, effective monitoring and infrastructure-as code. Working closely with our application engineering teams to ensure our services meet the highest standards of reliability, performance, and security. Key Responsibilities for Senior level The Site Reliability Engineering teams at Prove are responsible for driving maximum uptime for existing and developing products. Qualified candidates will be well versed in the difference between methods and ownership of outcomes and be able to demonstrate and document their relevant experience.  Observability Leadership Design and implement comprehensive observability solutions across our infrastructure and within applications Establish metrics, logging, and tracing systems that enable quick identification and resolution of issues Create alerting thresholds and automated responses based on service level objectives (SLOs) Provide actionable insights into service to service communications Infrastructure Management Design, build, and maintain scalable cloud infrastructure on AWS Implement infrastructure-as-code using tools such as Terraform Automate routine operational tasks to reduce toil and improve efficiency Ensure infrastructure security compliance and implement least-privilege access controls Design and implement infrastructure-as-code deployments for container based applications Scale containers based on custom metrics for applications and critical observability infrastructure Incident Response Conduct thorough post-incident reviews and implement preventative measures Use observability data to perform root cause analysis and system improvements Participate in a 24/7 on call rotation to achieve 99.999% system availability.  Required Qualifications for Senior level 5+ years of experience in Site Reliability Engineering, Platform Engineering or equivalent experience. Software Engineering roles with a strong infrastructure and production engineering aspect also qualify. Expert knowledge of observability platforms and practices (OpenTelemetry, Prometheus, Grafana, Jaeger, ELK stack / Splunk, etc) Experience with Kubernetes and container orchestration Strong experience with infrastructure-as-code tools (Terraform, Spacelift, Pulumi) Proficiency in at least one programming language ( Go, Python ) Deep understanding of cloud platforms, preferably AWS Bachelor's degree in Computer Science, Engineering, or equivalent practical experience Key Responsibilities for Site Reliability Engineer The Site Reliability Engineering teams at Prove are responsible for driving maximum uptime for existing and developing products. Qualified candidates will be well versed in the difference between methods and ownership of outcomes and be able to demonstrate and document their relevant experience.  Preferred Qualifications Experience with distributed systems and microservice architectures Experience working in a high compliance environment Hand-on experience instrumenting code with OpenTelemetry Familiarity with service mesh technologies  Contributions to open-source projects Experience in the identity verification or financial technology industry Application development experience Optimize Improve new and existing systems by increasing reliability, performance, and scalability Automate routine operational tasks to reduce toil and improve efficiency Ensure infrastructure security compliance and implement least-privilege access controls Implement efficient infrastructure that balances rapid development and cost Embrace technological changes and development practices while maintaining reliability Respond Participate in a 24/7 on-call rotation Conduct thorough post-incident reviews and implement preventative measures Use observability data to identify system improvements Run Implement infrastructure as code in a myriad of high compliance development, production, and other environments Scale developer experiences by being the standard bearer of an opinionated platform approach Required Qualifications 3+ years of experience in Site Reliability or Platform Engineering teams Deep understanding of cloud platforms,  particularly AWS Strong experience with Kubernetes and container orchestration Experience withTerrafo

More Remote jobs

Remote jobs · Browse all locations