CareerMoonshot

Lead Support Engineer

WPP Media · San Francisco Bay Area

📍 Los Angeles, United States; San Francisco, United States💰 $75,000via greenhousePosted 2026-09-22
Apply on company site ↗
Career Moonshot pulls this listing straight from the employer's hiring system — no recruiter middleman, no reposts. Applying takes you directly to WPP Media.
About WPP Media WPP is the trusted growth partner for the world’s leading brands. With exceptional talent, trusted data and intelligence, and world-class partnerships – all united by our pioneering agentic marketing platform, WPP Open – we help clients navigate change, capture opportunity, and deliver transformational growth.  WPP Media is WPP's AI-driven media operating unit, bringing together media, data, and partnerships to deliver creative personalisation at scale. Connected through WPP Open and powered by Open Intelligence, clients see exactly where, how, and why their media investment is working. For more information, visit wppmedia.com . Role Summary and Impact This role is within WPP Media, where you will be instrumental in owning production stability, observability, and system health for a mission-critical global campaign governance and compliance platform. As Lead Support Engineer, you will provide advanced L2+ and L3-oriented support for a system running natively on Google Cloud Platform (GCP). You will investigate complex production issues, implement minor code-level fixes in Python, lead root cause analysis, and improve the reliability of distributed systems. Working closely with the EMEA-based Engineering Tech Lead, Product Owner, and Quality Assurance partners, you will connect production insights with technical roadmap priorities. The opportunity combines hands-on Site Reliability Engineering (SRE) with technical squad leadership. A foundational monitoring and alerting setup is already in place, giving you the platform to evaluate the current approach, define a clear observability direction, and strengthen system health monitoring across the environment. You will automate runbooks, reduce manual operational effort, and help shape the future support squad for a dedicated product used across global advertising campaigns. Key Responsibilities Own production stability, system health, and observability for a mission-critical campaign governance and compliance platform running on GCP. Diagnose and resolve complex, intermittent, and high-priority incidents across application, database, infrastructure, and networking layers. Read, debug, and implement minor fixes and patches directly within the existing Python production codebase. Define and advance the observability strategy using GCP Cloud Logging, Cloud Monitoring, Error Reporting, Prometheus, PromQL, and suitable service level indicators and objectives. Lead incident response, post-incident reviews, and end-to-end root cause analysis, partnering with the EMEA Engineering team on permanent remediation. Monitor execution flows, performance, data refreshes, automated checks, and recurring compliance reporting to ensure reliable system operation. Identify opportunities to automate manual support activity, develop and maintain runbooks, and reduce operational toil. Provide technical leadership to the L2+ support squad through coaching, knowledge sharing, prioritization, and effective handovers. Partner with Product, Engineering, and Quality Assurance teams to assess operational risk, improve release and change management, and influence technical roadmap decisions. Skills and Experience Advanced education in computer science, software engineering, information technology, or a related technical discipline, or equivalent practical experience. Strong, current Python proficiency, including the ability to read, debug, and implement fixes directly within a shared production codebase. Deep, hands-on experience supporting enterprise systems on Google Cloud Platform, including Compute Engine, Google Kubernetes Engine, Cloud SQL, BigQuery, Pub/Sub, Firestore, Cloud Functions, Identity and Access Management, and virtual private cloud networking. Complex production troubleshooting and root cause analysis experience across distributed systems, including leadership of post-incident reviews. Strong experience with Docker and Kubernetes, particularly deploying, managing, and troubleshooting applications running in Google Kubernetes Engine. High proficiency with GCP Cloud Logging, Cloud Monitoring, Error Reporting, Prometheus, and PromQL, including establishing service level indicators and service level objectives. Experience using Terraform for infrastructure as code, tracing infrastructure-level issues, and identifying improvements to infrastructure management. Strong Python and Bash scripting skills for automating support tasks, parsing logs, developing operational tools, and reducing manual effort. Solid querying and troubleshooting experience across Firestore, BigQuery, PostgreSQL, and MySQL. Experience leading, mentoring, or serving as a technical lead for a squad or technical team, along with strong knowledge of Incident, Problem, and Change Management practices and the software development lifecycle. Life at WPP Media & Benefits Our passion for shaping the next era of media is powered by our commitment to Be Extraordinary, investing in our employees to inspire transformational creativity. We also Lead Optimistically, firmly believing in and Championing Growth and Development for every individual. This commitment allows WPP Media employees to leverage the extensive global WPP Media & WPP networks to pursue their passions, build vital professional connections, and learn at the cutting edge of marketing and advertising. We Create an Open environment built on trust and respect, where everyone feels they belong and has opportunities to progress. This inclusive culture is fostered through a variety of employee resource groups and frequent in-office events showcasing team wins, sharing thought leadership, and celebrating holidays and milestone events. Our comprehensive benefits package reflects this commitment, including competitive medical, group retirement plans, vision, and dental insurance, significant paid time off, preferential partner discounts, and employee mental health aware

More San Francisco Bay Area jobs

San Francisco Bay Area jobs · Browse all locations