Vice President – Director, AI Ops, Incident Management, Problem Management and Level 2 Application Support
OneMain Holdings, Inc. · Charlotte, NC
📍 Charlotte, NCvia workdayFirst listed here 2026-09-25
Apply on company site ↗
Career Moonshot pulls this listing straight from the employer's hiring system — no recruiter middleman, no reposts. Applying takes you directly to OneMain Holdings, Inc..
We are seeking an experienced and dynamic Vice President – Director responsible for AI Operations, Incident Management, Problem Management and Level 2 Application Support. This senior leader will be accountable for the strategy, execution, and continuous improvement of enterprise Incident Management, Problem Management, and Level 2 Application Support capabilities with a desired outcome of minimizing incidents and maximizing service availability.
The ideal candidate will possess a strong technology operational leadership background and a passion for building modern, proactive operations organizations that leverage automation, observability, artificial intelligence, and data-driven decision-making.
This leader will drive the transformation from reactive support models toward predictive and preventative operations through the adoption of AI Ops, intelligent automation, and continuous service improvement practices.
This role will partner closely with Engineering, Infrastructure, Architecture, Cybersecurity, Product, and Business stakeholders to improve operational resilience, reduce customer-impacting incidents, accelerate restoration times, eliminate recurring issues, and enhance the overall customer and team member experience.
Responsibilities
Develop Strategy and Vision
· Define and execute the enterprise strategy and roadmap for Incident Management, Problem Management, and Level 2 Application Support.
· Establish a long-term vision for operational excellence, resiliency, service restoration, and issue prevention.
· Develop and maintain organizational goals, key performance indicators (KPIs), and operational maturity targets.
· Ensure alignment between operational priorities, business objectives, customer experience goals and technology strategy.
Lead Incident Management
· Provide executive leadership and governance for major incident management across the enterprise.
· Establish and continuously improve incident response processes, escalation procedures, communication standards, and operational playbooks.
· Ensure rapid and effective coordination across Technology teams during critical incidents.
· Drive improvements in Mean Time to Detect (MTTD), Mean Time to Restore Service (MTTR), customer impact measurement, and incident communications.
· Partner with Observability, Monitoring, Infrastructure, and Application teams to improve detection, diagnosis, and recovery capabilities.
· Lead executive communications during significant customer or business impacting events.
Lead Problem Management and Continuous Improvement
· Own the enterprise Problem Management practice and associated governance.
· Establish rigorous root cause analysis standards and ensure corrective actions are identified, prioritized, and completed.
· Analyze operational trends, recurring incidents, and systemic risks to identify opportunities for stability improvements.
· Drive initiatives that eliminate recurring issues, reduce operational toil, and improve overall platform reliability.
· Develop reporting and executive dashboards that measure problem trends and remediation effectiveness.
· Foster a blameless culture focused on prevention rather than reaction.
Lead Level 2 Application Support
· Provide leadership for Level 2 Application Support teams responsible for diagnosing, troubleshooting, and restoring application services.
· Establish support models, staffing strategies, and operational procedures aligned to business priorities and customer outcomes.
· Ensure Standard Operating Procedures (SOPs), knowledge articles, runbooks, and troubleshooting guides remain current and effective.
· Partner closely with Level 3 Engineering teams to improve supportability, logging, instrumentation, diagnostics, and operational readiness.
· Ensure support teams are prepared to support new technologies, platforms, and application releases.
· Drive consistency and excellence across all application support functions.
Advance AI Ops and Intelligent Automation
· Develop and execute an enterprise AI Ops strategy that transforms incident detection, diagnosis, correlation, prediction, and remediation.
· Leverage machine learning, event correlation, predictive analytics, and automation technologies to identify operational issues before customer impact occurs.
· Drive adoption of intelligent alerting, automated triage, automated root cause identification, and self-healing capabilities.
· Partner with Observability, Engineering, and Data teams to build operational intelligence platforms that reduce noise, improve signal quality, and accelerate resolution.
· Identify opportunities to automate repetitive support, incident response, and problem management activities through workflows, orchestration, scripting, and AI-enabled solutions.
· Establish measurable targets for automation adoption, operational efficiency, incident reduction, and support productivity improvements.
· Continuously evaluate emerging AI, automation, and operational technologies to enhance operational resilience and effectiveness.
Collaboration and Stakeholder Engagement
· Build strong partnerships across Engineering, Infrastructure, Architecture, Cybersecurity, Product, Risk, Compliance and Operations organizations.
· Influence operational priorities and investment decisions through data-driven recommendations.
· Lead governance forums focused on operational stability, incident trends, problem remediation, and service performance.
· Present operational performance, risk assessments, and strategic recommendations to senior executives and stakeholders.
· Serve as a champion for operational excellence and customer-centric decision making throughout
More Charlotte, NC jobs
Charlotte, NC jobs · Browse all locations