CareerMoonshot

Site Reliability Engineering Lead

LexisNexis Risk Solutions · Georgia

📍 Alpharetta, GAvia workdayFirst listed here 2026-09-24
Apply on company site ↗
Career Moonshot pulls this listing straight from the employer's hiring system — no recruiter middleman, no reposts. Applying takes you directly to LexisNexis Risk Solutions.
Ready to lead the reliability, scalability, and operational excellence of mission-critical platforms while shaping the future of Site Reliability Engineering? Would you like to mentor high-performing engineers, drive cloud modernization, and influence enterprise-wide engineering practices in a highly collaborative environment? About the Business LexisNexis Risk Solutions is the essential partner in the assessment of risk. Within our Insurance vertical, we provide customers with solutions and decision tools that combine public and industry specific content with advanced technology and analytics to assist them in evaluating and predicting risk and enhancing operational efficiency. Our insurance risk solutions help drive better data-driven decisions across the insurance policy lifecycle, all while reducing risk. You can learn more about LexisNexis Risk at   https://risk.lexisnexis.com/insurance About our Team The ICS (Insurance Core Services) team is responsible for establishing and driving reliability, observability, automation, and operational excellence standards across Insurance technology platforms. The team partners closely with application, infrastructure, database, and cloud engineering teams to improve platform availability, scalability, performance, and resilience. ICS leads strategic initiatives including SLO/SLI implementation, observability platform adoption, cloud modernization, operational readiness reviews, performance engineering, and reliability automation. The team also develops reusable engineering frameworks, standards, and best practices that enable product teams to build and operate highly reliable cloud-native services at scale. The SRE Lead will play a key role in shaping reliability strategy, mentoring engineers, driving cross-functional initiatives, and partnering with business and technology stakeholders to improve service reliability and operational maturity across the organization. About the Role As a Site Reliability Engineering Lead, you will provide technical leadership and strategic direction for Site Reliability Engineering initiatives across multiple product portfolios. You will lead a team of SREs responsible for ensuring the reliability, scalability, security, performance, and operational excellence of mission-critical applications and platforms. The SRE Lead will partner closely with engineering, architecture, security, operations, and business stakeholders to drive cloud modernization, operational maturity, observability excellence, automation, and continuous improvement. This role combines hands-on technical expertise with people leadership, mentoring, strategic planning, and cross-functional collaboration. Responsibilities Lead and mentor a team of Site Reliability Engineers, fostering a culture of ownership, operational excellence, collaboration, and continuous learning. Define and drive SRE strategy, standards, best practices, and operational frameworks across engineering organizations. Partner with product and platform teams to improve application reliability, scalability, security, performance, and resilience. Establish and maintain service level objectives (SLOs), service level indicators (SLIs), and error budgets. Lead major incident management, root cause analysis, problem management, and post-incident review processes. Drive cloud modernization initiatives and support application migrations to Azure, AWS, and containerized environments. Champion automation and Infrastructure as Code (IaC) practices using tools such as Terraform, GitHub, GitLab, Jenkins, and Ansible. Develop and implement observability strategies utilizing metrics, logs, traces, alerting, and dashboards. Collaborate with security and compliance teams to ensure platform adherence to enterprise security and regulatory requirements. Lead architecture reviews and provide guidance on cloud-native and highly resilient application designs. Drive capacity planning, performance optimization, cost management, and operational efficiency initiatives. Establish engineering guardrails, governance controls, and deployment standards for production environments. Support organizational transformation toward DevOps and SRE practices. Manage operational risk and ensure business continuity and disaster recovery preparedness. Collaborate with stakeholders to prioritize reliability improvements and platform investments. Build and maintain strong relationships with product owners, engineering leaders, vendors, and business partners . Leadership Responsibilities Lead, coach, mentor, and develop a high-performing team of Site Reliability Engineers. Conduct resource planning and support hiring, onboarding, and career development activities. Establish team objectives aligned with business and technology strategies. Promote accountability, innovation, and operational excellence within the team. Act as a trusted advisor and subject matter expert for reliability engineering across the organization. Drive cross-team collaboration and alignment on strategic initiatives. Essential Skills and Attributes Strong leadership experience managing technical engineering teams. Deep expertise in Site Reliability Engineering, DevOps, Cloud Engineering, or Platform Engineering disciplines. Extensive experience with Azure and/or AWS cloud platforms. Strong understanding of Kubernetes, AKS, EKS, containerization, Docker, and cloud-native architectures. Expertise with Infrastructure as Code tools such as Terraform and Ansible. Strong background in observability platforms such as Grafana, Prometheus, OpenTelemetry, Splunk, Dynatrace, Datadog, or similar technologies. Experience managing large-scale production environments with stringent availability requirements. Strong understanding of security, compliance, networking, and cloud governance principles. Experience designing highly available, fault-tolerant, and resilient systems. Strong proficiency in at least on

More Georgia jobs

Georgia jobs · Browse all locations