CareerMoonshot

Principal Site Reliability Engineer - Paze

Earlywarning · Chicago, IL

📍 Chicagovia workdayFirst listed here 2026-09-24
Apply on company site ↗
Career Moonshot pulls this listing straight from the employer's hiring system — no recruiter middleman, no reposts. Applying takes you directly to Earlywarning.
At Early Warning, we’ve powered and protected the U.S. financial system for over thirty years with cutting-edge solutions like Zelle®, Paze®, and so much more. As a trusted name in payments, we partner with thousands of institutions to increase access to financial services and protect transactions for hundreds of millions of consumers and small businesses. Positions located in Scottsdale, San Francisco, Chicago, or New York follow a hybrid work model to allow for a more collaborative working environment. Candidates responding to this posting must independently possess the eligibility to work in the United States, for any employer, at the date of hire. This position is ineligible for employment Visa sponsorship. Role Summary   The Principal Site Reliability Engineer applies software engineering and systems engineering practices to improve the reliability, resilience, scalability, and operational health of production services. The role partners with Software Engineering and other technology teams to ensure reliability, observability, recoverability, performance, and operational readiness are engineered into systems throughout their lifecycle.   The role   operates   at enterprise scope,   establishing   technical direction and applying evidence-driven engineering, technical rigor, sound judgment, automation, and broad   systems   expertise   across organizational boundaries.   Core Responsibilities   Use software engineering, automation, and DevOps principles and practices to continually improve how services are built, tested, deployed,   observed , operated, and recovered.   Use data, evidence, experimentation, and rigorous engineering analysis   appropriate to   the level to   identify   reliability risks, test assumptions, and guide technical decisions.   Define, implement, or improve SLIs, SLOs, error budgets, and other service-health measures appropriate to the scope of responsibility.   Improve observability through metrics, logging, tracing, monitoring, alerting, dashboards, and service-health instrumentation.   Drive continuous improvement across CI/CD, observability, deployment practices, Infrastructure as Code, automation, testing, incident response, capacity management, resilience, and operational readiness.   Identify   recurring or systemic production issues and translate operational experience into improvements in code, architecture, automation, tooling, and engineering practices.   Partner with Software Engineering teams to incorporate reliability, resiliency, scalability, performance, observability, recoverability, and operational readiness throughout the   development   lifecycle.   Participate in or lead incident response, troubleshooting, service restoration, and blameless post-incident learning   appropriate to   the level.   Provides enterprise-level technical leadership for critical production incidents and   establishes   or influences engineering practices that improve incident response, escalation, service   restoration   and sustainable on-call operations across the organization.   Reduce operational toil and unnecessary manual intervention through software, automation, reusable patterns, and better engineering practices.   Leveling Intent   Principal   represents   domain-level technical leadership and organizational impact. Deep individual   expertise   is expected, but Principal-level impact comes from   identifying   systemic risk,   establishing   technical direction,   influencing engineering practices and architecture, and multiplying the capability of the broader engineering organization.   Level Expectations   Acts as an enterprise force multiplier, raising the effectiveness and technical capability of engineers and teams across the organization while building sustainable organizational capability rather than individual dependency.   Demonstrates software engineering, systems thinking, troubleshooting, and   production   reliability capabilities appropriate to the level.   Applies evidence-driven reasoning and technical rigor to distinguish observed facts from assumptions and make defensible engineering recommendations.   Shares knowledge and contributes to sustainable engineering capability rather than creating dependency on individual   expertise .   Operates with significant autonomy across the organization's most consequential reliability challenges.   Establishes enterprise technical direction, develops senior technical leaders, and   demonstrates   impact well beyond systems personally touched.   Minimum Qualifications   Typically  15 + years of relevant professional experience in Software Engineering, Site Reliability Engineering, Systems Engineering, Cloud/Platform Engineering, DevOps, Infrastructure Engineering, Architecture where applicable, or a comparable technical discipline.   Experience with software development or scripting using one or more modern programming languages.   Experience with software engineering principles, distributed systems, production troubleshooting, automation, and observability   appropriate to   the level.   Experience with public cloud technologies and architectures, preferably AWS, along with infrastructure, networking, Linux/Unix, and modern application architectures   appropriate to   the level.   Demonstrated analytical, problem-solving, communication, and collaboration skills   appropriate to   the scope of the role.   Preferred Qualifications   Hands-on experience with AWS is preferred, or comparable experience with   another major cloud platform   such as Microsoft Azure, Google Cloud Platform (GCP), or Oracle Cloud Infrastructure (OCI).   Experience developing, deploying,   operating , or improving   highly available   production software or distributed systems.   Experience with CI/CD, Infrastructure as Code, containers or orchestration, observability, monitoring, alerting, and software-delivery automation.   Ex

More Chicago, IL jobs

Chicago, IL jobs · Browse all locations