CareerMoonshot

Infrastructure Engineer Lead – Cloud AI

American Electric Power · Ohio

📍 Gahanna, OHvia workdayFirst listed here 2026-09-23
Apply on company site ↗
Career Moonshot pulls this listing straight from the employer's hiring system — no recruiter middleman, no reposts. Applying takes you directly to American Electric Power.
Job Posting End Date 10-03-2026 Please note the job posting will close on the day before the posting end date. Job Summary This hands-on role builds and operates secure, scalable infrastructure for AI and machine learning workloads across AWS and approved on-premises environments. Responsibilities include cloud foundations, networking, containers, compute, identity, cost management, and observability. The engineer partners with AI, architecture, cybersecurity, network, and application teams to deliver secure, cost-effective, production-ready infrastructure while supporting core cloud engineering services. Job Description What you’ll do: Essential Job Functions & Tasks AI Cloud Foundations Design, build and maintain the AWS account structure, landing zones, and reference patterns that host AI and machine learning workloads, leveraging AWS Control Tower, Landing Zone Accelerator, and Terraform. Engineer the compute, storage, and networking foundations required for AI workloads, including GPU and accelerated instance families, high-throughput storage, and data access paths to enterprise data platforms. Enable and operate AI platform services such as Amazon Bedrock and SageMaker, including private connectivity, model access provisioning, logging, and quota management. Build reusable infrastructure-as-code modules, templates and pipelines so that AI teams can deploy quickly within approved guardrails. Design, build and maintain on-premises AI infrastructure for edge and specialized use cases. Support AI in disconnected or intermittently connected environments, including local compute, model distribution, patching, monitoring, backup, and recovery. Create deployment standards, automation, runbooks, and support processes for on-premises and edge AI. Integrate on-premises AI with enterprise identity, security, networking, monitoring, and governance where feasible. Controls, Security & Governance Partner with Cybersecurity, Enterprise Architecture and Compliance to align AI infrastructure with AEP security standards, regulatory obligations, and responsible AI guardrails. Review and remediate configuration drift, vulnerabilities and audit findings across the AI cloud estate; support evidence requests and control attestations. Contribute to onboarding and intake processes for new AI use cases, ensuring workloads are provisioned into the right accounts with the right controls from day one. FinOps & Cost Management Establish and operate FinOps practices for AI workloads, including tagging standards, showback/chargeback, budgets, anomaly detection, and forecasting. Analyze and optimize spend on GPU compute, inference and token consumption, storage, and data transfer; recommend commitment strategies such as Savings Plans and Reserved Instances. Provide cost transparency and consumption reporting to business stakeholders and technology leadership, and identify optimization opportunities before they become budget issues. Monitoring & Observability Partner with monitoring team to create logging, alerting and dashboards for AI infrastructure and workloads. Define service-level objectives and operational thresholds for AI platforms, including model endpoint availability, latency, throughput, and error rates. Support incident response, root cause analysis and problem management for AI-related infrastructure events; drive preventive actions and automation to reduce recurrence. Core Cloud Engineering & Platform Services Perform standard cloud engineering functions across the AEP AWS environment including account provisioning, environment builds, platform upgrades, patching, automation and lifecycle management. Design, deploy and operate Kubernetes environments (Amazon EKS, ROSA/OpenShift) including cluster architecture, autoscaling, node group and GPU scheduling, ingress, service mesh, RBAC, and cluster security hardening. Engineer and support platform services including load balancing, traffic management, API gateways, and integration patterns across cloud, on-premises, edge, and disconnected environments. Apply networking fundamentals, VPC design, subnetting, routing, Transit Gateway, Direct Connect, VPN, firewalls, TLS and certificate management to deliver secure, performant connectivity for AI and general workloads. Prepare cost estimates, justifications, alternative solutions and technical recommendations; produce technical documentation, runbooks and standards. Collaborate with Project Managers, Architects, Solution Engineers, Business Analysts and vendor partners to deliver consistent, reliable solutions that leverage AEP's technology standards, architectures and best practices. Adhere to and advocate for change, incident and problem management processes; participate in on-call rotation and after-hours support as required. Provide training, mentoring and technical work direction to other engineers on the team. Required Skills & Experience Demonstrated hands-on engineering experience in AWS, including IAM, VPC networking, compute, storage, encryption/KMS, and account/organization structure. Strong Kubernetes expertise including cluster design, operations, troubleshooting and security in EKS, ROSA or OpenShift. Solid networking fundamentals, including routing, DNS, load balancing, traffic management, firewalls, hybrid connectivity, and network design for disconnected or intermittently connected environments. Experience with API gateways, integration patterns, and exposing and securing services across environments. Proficiency with infrastructure-as-code and automation (Terraform required; Ansible, Python or PowerShell preferred) and CI/CD pipelines. Working knowledge of cloud security principles, identity and access management, and compliance requirements. Experience implementing monitoring and cost management for cloud environments, with the ability to establish local monitoring, logging, patching, backup, and recovery processes for on-premises

More Ohio jobs

Ohio jobs · Browse all locations