CareerMoonshot

Sr. Software Engineer- AI/ML, AWS Neuron Apps

Amazon · Seattle, WA

📍 Seattle, Washington, USAvia amazonPosted 2025-10-30
Apply on company site ↗
Career Moonshot pulls this listing straight from the employer's hiring system — no recruiter middleman, no reposts. Applying takes you directly to Amazon.
Shape the Future of AI Accelerators at AWS Neuron Join the team behind AWS Neuron — the software stack that powers AWS's purpose-built AI accelerators, Inferentia and Trainium. As a Senior Software Engineer on our Machine Learning Applications team, you will optimize the world's most demanding AI models at a scale few engineers ever get to work on. What You'll Do • Build and scale distributed inference solutions for leading large language models, including GPT, Kimi, and Qwen • Partner directly with silicon architects and compiler engineers to shape the next generation of AI acceleration • Write custom kernels that optimize LLM computation graphs, improving latency and cost for billions of inference requests worldwide • Optimize state-of-the-art language, vision, and multimodal generative AI models for Neuron hardware Key job responsibilities You will drive the Evolution of Distributed AI at AWS Neuron Technical Impact You'll Drive: • Spearhead distributed inference architecture for PyTorch • Engineer breakthrough performance optimizations for AWS Trainium and Inferentia • Develop kernels to improve model efficiency on Amazon AI Accelerators • Transform complex tensor operations into highly optimized hardware implementations What Makes This Role Unique: • Direct influence on AWS's AI infrastructure used by thousands of ML applications • Full-stack optimization from high-level frameworks to hardware-specific primitives • Develop tools that define industry standards for ML deployment • Collaboration with both open-source ML communities and hardware architecture teams Your Technical Arsenal Should Include: • Deep expertise in Transformer architecture, Python and Pytorch internals • Strong understanding of distributed systems and ML optimization • Passion for performance tuning and system architecture A day in the life Work/Life Balance Our team puts a high value on work-life balance. It isn’t about how many hours you spend at home or at work; it’s about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives. Mentorship & Career Growth Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we’re building an environment that celebrates knowledge sharing and mentorship. We care about your career growth and strive to assign projects based on what will help each team member develop into a better-rounded professional and enable them to take on more complex tasks in the future. About the team At AWS Neuron, we're revolutionizing how the world's most sophisticated AI models run at scale through Amazon's next-generation AI accelerators. Operating at the unique intersection of ML frameworks and custom silicon, our team drives innovation from silicon architecture to production software deployment. We pioneer distributed inference solutions for PyTorch and JAX using XLA, optimize industry-leading LLMs like GPT and Llama, and collaborate directly with silicon architects to influence the future of AI hardware. Our systems handle millions of inference calls daily, while our optimizations directly impact thousands of AWS customers running critical AI workloads. We're focused on pushing the boundaries of large language model optimization, distributed inference architecture, and hardware-specific performance tuning. Our deep technical experts transform complex ML challenges into elegant, scalable solutions that define how AI workloads run in production.

More Seattle, WA jobs

Seattle, WA jobs · Browse all locations