CareerMoonshot

Senior Principal AI Engineer

CERENCE OPERATING · Remote

📍 Remote - USAvia workdayFirst listed here 2026-08-08
Apply on company site ↗
Career Moonshot pulls this listing straight from the employer's hiring system — no recruiter middleman, no reposts. Applying takes you directly to CERENCE OPERATING.
A Moving Experience. What You Will Work On    Design and   operate   distributed training systems for large neural networks   (autoregressive, diffusion ,   State Space Models   etc. )   across GPU clusters   Optimise   multi ‑ node, multi ‑ GPU execution to   maximize   throughput and   utilization   Diagnose   &   resolve bottlenecks across compute, memory, and netwo rk   Improve training stability and fault tolerance at scale   Partner with research and applied ML teams to productioni z e   large ‑ model   training pipelines      Core Responsibilities    Distributed Training Infrastructure    Build and   optimize   GPU cluster orchestration using:   Slurm     Kubernetes    Ray    RunAI     Ensure efficient scheduling, isolation, and fairness across training workloads   Communication & Networking    Optimize   and debug distributed communication using:    NCCL    RDMA    InfiniBand    NVLink     Minimize   networking bottlenecks that dominate   end ‑ to ‑ end   training time    Training Frameworks    Scale large-model training using:    PyTorch   Distributed    Megatron ‑ LM     DeepSpeed     Own   multi ‑ node   launch configurations, failure recovery, and performance tuning    Memory & Performance Optimization    Apply advanced memory optimization techniques:    Activation checkpointing    ZeRO   (Stage 1–3) and offload strategies    Balance compute, memory, and communication to push model size and batch scale   What Success Looks Like     GPU   utilization   consistently stays high (>80–90%)    Training scales cleanly from single node to dozens or hundreds of GPUs   Communication overhead is minimized and predictable    Large training jobs run stably for days or weeks without failure    New models can be trained faster, larger, and more reliably than before      Required Experience & Skills     Strongly Required    Deep   hands ‑ on   experience with distributed systems or ML systems   Experience running   large ‑ scale   workloads on GPU clusters    Production experience with   PyTorch   distributed training    Strong understanding of parallelism strategies (data, tensor, pipeline parallelism)    Low ‑ level   understanding of GPU communication and networking     Critical Technical Skills    GPU orchestration:   Slurm , Kubernetes, Ray,   RunAI     Communication libraries: NCCL, RDMA, InfiniBand,   NVLink     Training frameworks:   PyTorch   Distributed,   Megatron ‑ LM ,   DeepSpeed   Memory   optimi s ation : activation checkpointing,   ZeRO   offload techniques    Common Problems   You’ll   Be Solving    Many teams fail at scale because:    GPU   utilization   is low despite large clusters    Networking and communication dominate training time    Training jobs crash or become unstable at large scale    You will be explicitly focused on   eliminating   these failure modes.      Ideal Background     This role is a strong fit for individuals who have worked as:    ML Systems Engineer    Distributed Systems Engineer    AI Infrastructure Engineer    HPC Engineer transitioning into ML    Experience working with large language models or foundation models is a strong plus, but deep systems   expertise   is valued over pure model architecture experience.      Why This Role Matters     Without robust distributed training infrastructure, progress on large   models   stalls. This role directly enables:    Larger models    Faster iteration cycles    More reliable research-to-production pipelines    You will be building the foundation that makes   large ‑ scale   AI possible.   Cerence Inc. (Nasdaq: CRNC and  www.cerence.com ) is the global industry leader in creating unique, moving experiences for the automotive world. Spun out from Nuance in October 2019, Cerence is a new, independent company that has quickly gained traction as a leader in the automotive voice assistant space, working with all of the world’s leading automakers – from Ford and Fiat Chrysler to Daimler, Audi and BMW to Geely and SAIC – to transform how a car feels, responds and learns. Its track record is built on more than 20 years of industry experience and leadership and more than 500 million cars on the road today across more than 70 languages.    As Cerence looks to the future and continues an ambitious growth agenda, we need someone to join the team and help build the future of voice and AI in cars. This is an exciting opportunity to join Cerence’s passionate, dedicated, global team and be a part of meaningful innovation in a rapidly growing industry.   EQUAL OPPORTUNITY EMPLOYER Cerence is firmly committed to Equal Employment Opportunity (EEO) and to compliance with all federal, state and local laws that prohibit employment discrimination on the basis of age, race, color, gender, gender identity, gender expression, sex, sex stereotyping, pregnancy, national origin, ancestry, religion, physical or mental disability, medical condition, marital status, citizenship status, sexual orientation, protected military or veteran status, genetic information and other protected classifications. Cerence Equal Employment Opportunity Policy Statement. All prospective and current Employees need to remain vigilant when it comes to executing security policies in the workplace. This includes: - Following workplace security protocols and training programs to familiarize with the ways to maintain a safe workplace. - Following security procedures to report any suspicious activity. - Having respect for corporate security procedures to allow those procedures to be effective. - Adhering to company's compliance and regulations. - Encouraging to follow a zero tolerance for workplace violence. - Basic knowledge of information security and data privacy requirements (e.g., how to protect data & how to be handling this data). - Demonstrative

More Remote jobs

Remote jobs · Browse all locations