Software Development Senior Specialist
NTT AMERICA · Charlotte, NC
📍 Charlotte, US-NCvia phenomPosted 2026-09-09
Apply on company site ↗
Career Moonshot pulls this listing straight from the employer's hiring system — no recruiter middleman, no reposts. Applying takes you directly to NTT AMERICA.
Req ID: 388176
NTT DATA strives to hire exceptional, innovative and passionate individuals who want to grow with us. If you want to be part of an inclusive, adaptable, and forward-thinking organization, apply now.
We are currently seeking a Real-Time Inference Engineering Lead (FTE / Hybrid) to join our team in Charlotte , North Carolina (US-NC) , United States (US) .
Job Duties and Responsibilities:
Platform Context
The Cortex Predictive AI Platform accelerates predictive AI modernization and enterprise adoption across the full model lifecycle: governed data and features; model build, training, and validation; deployment and inference; and ongoing monitoring and operations. The Real-Time Services portfolio provides standardized, scalable model-serving capabilities for applications requiring reliable, performant online inference.
Position Summary
The Real-Time Inference Engineering Lead will design, build, and industrialize low-latency, resilient model-serving services for real-time predictive AI use cases. This role provides technical leadership for online inference architecture, deployment patterns, API services, capacity controls, observability, and operational practices across public cloud and on-premises environments.
The successful candidate will establish reusable patterns that enable application, data science, and ML engineering teams to deploy and operate predictive models safely and efficiently at scale. This is a hands-on engineering role requiring strong experience with model serving, Kubernetes, APIs, performance optimization, reliability engineering, CI/CD, and production operations.
Key Responsibilities
Define the target architecture and engineering standards for real-time predictive model-serving services across cloud and on-premises environments.
Design, build, test, deploy, and operate scalable online inference services that meet latency, throughput, availability, resiliency, and security requirements.
Establish reusable model-serving patterns for synchronous APIs, asynchronous inference, batch-adjacent processing, and event-driven real-time use cases where appropriate.
Build standardized deployment approaches for predictive models, including model packaging, versioning, release promotion, canary deployment, rollback, and retirement.
Design and implement secure API patterns for inference services, including authentication, authorization, traffic management, rate limiting, auditability, and integration with enterprise systems.
Engineer Kubernetes-based serving platforms using GKE, OpenShift, and related container orchestration capabilities.
Implement autoscaling, resource allocation, quota management, capacity planning, and workload-isolation controls for variable inference demand.
Conduct performance engineering, load testing, stress testing, and failure testing to validate service behavior under expected and peak production workloads.
Identify and implement latency-optimization opportunities across model initialization, feature retrieval, network paths, API handling, runtime configuration, and infrastructure utilization.
Define and implement monitoring, telemetry, dashboards, alerts, SLIs, SLOs, and error-budget practices for real-time inference services.
Partner with ML platform, data engineering, application engineering, security, and operations teams to integrate model services with governed data, feature, network, and identity capabilities.
Implement CI/CD and automated validation for model-serving services, infrastructure configuration, APIs, performance benchmarks, and release-readiness checks.
Build operational runbooks, incident-response procedures, support models, and production-readiness artifacts for real-time services.
Drive reliability improvements through root-cause analysis, capacity reviews, resiliency testing, disaster-recovery planning, and continuous operational improvement.
Mentor engineers and establish reusable technical documentation, reference implementations, and knowledge-transfer materials for real-time inference capabilities.
Required Qualifications
8+ years of software engineering, platform engineering, cloud engineering, SRE, or infrastructure engineering experience.
4+ years of experience designing, building, or operating production APIs, distributed systems, platform services, or real-time data and ML workloads.
Demonstrated experience leading technical design and engineering delivery for highly available, performance-sensitive production services.
Strong experience with online inference architecture, model-serving frameworks, or predictive-model deployment patterns.
Hands-on experience designing and operating RESTful, gRPC, or event-driven APIs.
4+ years satrong experience with Kubernetes and container platforms in production, including GKE, OpenShift, or comparable environments.
Experience with autoscaling, resource management, capacity planning, performance testing, and load testing for distributed services.
Experience implementing observability, monitoring, dashboards, alerts, SLIs, SLOs, and incident-management practices.
Experience with CI/CD, Git-based development, automated testing, deployment automation, and production-release practices.
Strong understanding of resiliency, high availability, fault tolerance, disaster recovery, and operational support for critical services.
Ability to work effectively with data science, ML engineering, platform engineering, application teams, security, and business stakeholders.
Required Skills / Knowledge
Online inference and low-latency model-serving architecture.
Model deployment, versioning, routing, rollout, rollback, and lifecycle management.
REST APIs, gRPC, API gateways, authentication, authorization, traffic management, and API observability.
Kubernetes, GKE, OpenShift, containers, service meshes, ingress, workload scheduling, and autoscaling.
Performance engineering, load testing, stress testin
More Charlotte, NC jobs
Charlotte, NC jobs · Browse all locations