CareerMoonshot

Staff Infrastructure Software Engineer, Fleet & Automation

nscaleoperationsukltd · San Francisco Bay Area

📍 Houston; New York; San Francisco; Seattlevia greenhousePosted 2026-09-21
Apply on company site ↗
Career Moonshot pulls this listing straight from the employer's hiring system — no recruiter middleman, no reposts. Applying takes you directly to nscaleoperationsukltd.
About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers.  Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you’ll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you’ll be contributing to building the technology that powers the future. About the Role We're hiring a Staff Software Engineer to build the software, automation, and control-plane capabilities that manage Nscale's fleet of AI infrastructure at scale. Your work will improve the acceptance, performance, and scalability of our AI and high-performance computing environments — driving higher availability, faster capacity delivery, and lower operational load as Nscale grows into one of the world's leading neo-cloud providers. This is a senior individual-contributor role for an engineer who enjoys solving hard infrastructure problems at the intersection of software, GPUs, networking, and large-scale operations. You will have the autonomy to investigate problems, learn quickly, innovate, and deliver improvements wherever they create meaningful impact for the team and the platform. You will work closely with teams across Nscale — including Deployment, AI Infrastructure Support, Data Centre Operations, Platform, SRE, Network, and hardware engineering — to translate operational challenges into reliable, scalable software. You will not need to own every component to make a difference: strong engineers identify opportunities, build a compelling case for a solution, and work with the right partners to deliver it. NOTE: We are hiring for various senior experience levels. The final leveling for the role will be based on your overall work experience, experience in AI Infra domain and interview feedback. What You'll Be Doing Lead the architecture, roadmap, and implementation of workflow automation and fleet-management systems, balancing scalability, reliability, and maintainability. Build and operate production-grade software, services, APIs, and automation that manage the lifecycle of GPU compute and supporting network infrastructure. Own end-to-end workflows for fleet inventory, provisioning, configuration, hardware and firmware lifecycle management, validation, health monitoring, remediation, capacity, and reliability at scale. Investigate complex production issues across hardware, GPUs, operating systems, networks, schedulers, and services; turn findings into durable software improvements rather than recurring manual work. Build safe, observable, and auditable control-plane workflows that give operators clear visibility and reliable ways to act. Establish engineering standards for reliability, observability, testing, CI/CD, security, incident response, and operational readiness. Use SLOs, telemetry, alerting, and postmortems to drive continuous improvement. Partner with Deployment, AI Infrastructure Support, Data Centre Operations, Platform, SRE, Network, and hardware teams to translate operational needs into robust, scalable automation. Influence the evolution of adjacent systems and services through sound technical judgment, clear communication, and practical solutions. Assess the impact of new hardware programmes on the software stack and ensure fleet-management capabilities are ready to support them. Lead technical design reviews and incident deep-dives; mentor other engineers and raise the engineering bar across the organization. Use AI-assisted development tools to increase delivery leverage while maintaining a high bar for correctness, security, and operational safety. About You 8+ years of experience building and operating large-scale infrastructure applications, platform services, cloud systems, or equivalent production systems. A Bachelor's degree in Computer Science, Computer Engineering, a relevant technical field, or equivalent practical experience. A strong software-engineering foundation in Python and/or Go, Java, C++, or similar languages, including API design, testing, code review, and production debugging. Deep understanding of Linux, distributed systems, networking fundamentals, and systems performance; you are comfortable working across stateful and stateless services. Experience designing and operating reliable automation or control-plane systems for complex infrastructure, large fleets, cloud platforms, or hardware lifecycle management. Proven ability to take ambiguous technical problems from architecture through implementation and production operation, while influencing peers and stakeholders without relying on formal authority. Hands-on experience with observability, monitoring, metrics, logs, tracing, alerting, incident response, capacity planning, and performance analysis. Strong communication skills and sound technical judgment. You can explain trade-offs clearly, build alignment, and move work forward in a fast-changing environment. A curious, pragmatic, high-ownership mindset. You enjoy finding the underlying cause of difficult problems and building the simplest durable solution. Strong Candidates Will Have Direct experience with AI, GPU, HPC, or large-scale cloud infrastructure, including NVIDIA GPUs, CUDA, NVLink/NVSwitch, NCCL, and workload schedulers such as Slurm and Kubernetes. Experience with high-performance datacentre networking, including InfiniBand, RoCE, Ethernet fabrics, routing, congestion control, topology-aware systems, or GPU Direct RDMA.

More San Francisco Bay Area jobs

San Francisco Bay Area jobs · Browse all locations