Lead AI Data Engineer (Hybrid)
Genevausa · Maryland
📍 Bethesda, MD💰 $155,000 - $193,000via workdayFirst listed here 2026-09-21
Apply on company site ↗
Career Moonshot pulls this listing straight from the employer's hiring system — no recruiter middleman, no reposts. Applying takes you directly to Genevausa.
About The Position
The Lead AI Data Engineer serves as a scientific, technical, and managerial lead for data-heavy research projects. The Lead AI Data Engineer will overlap with a multidisciplinary team of government and contract researchers, academic experts and consumers. The Lead AI Data Engineer will provide technical and management support, oversee project execution, and provide key guidance on data architecture and infrastructure. The team that the Lead AI Data Engineer oversees is responsible for the curation and maintenance of large data pipelines, developing ETL pipelines, defining schemas, identifying bottlenecks, and deriving variables /features directly used in machine learning / AI model development. The Lead AI Data Engineer serves as a versatile position who designs, builds, and maintains the entire data ecosystem. As project aims evolve, collaborators join, and team processes change, an ideal candidate is flexible and can adapt quickly. The ideal candidate will also be mindful of data privacy / security and will adhere to data governance and data management best practices.
This is a full time hybrid remote position working at the Walter Reed National Military Medical Center in Bethesda, MD that will require working in office/on site at least 2 days per week. Background checks will be administered.
About The Program
Sleep Physiology Modeling Project: Sleep & Wearables Operational Readiness for Research & Defense (SWORD) Lab
Salary Range
$155,000 - $193,000. Salaries are determined based on several factors including external market data, internal equity, and the candidate’s related knowledge, skills, and abilities for the position.
Qualifications
PhD in a relevant field (e.g., Computer Science, Engineering, Data Science, Biomedical Engineering) required
8+ years experience with multimodal data analysis and data pipeline engineering required
Proven experience with multivariate signal processing (e.g., time-series biosensor data)
Hands-on experience with relational (SQL) and non-relational (NoSQL) databases
Hands-on experience with version control systems (eg, Git) and demonstrated ability to work in (and lead) a collaborative coding environment
Solid problem-solving and analytical skills to address complex technical challenges
Hands-on experience using Google Cloud Platform (GCP) cloud infrastructure (or equivalent), including setting up and managing cloud-native data warehouses (eg, BigQuery), storage, and compute resources
Ability to translate high-level scientific hypotheses into scalable engineering solutions and data products
Ability to work in a fast-paced, multidisciplinary, multi-site (sometimes asynchronous) team environment
Preferred qualifications: Strong knowledge of sleep science and hands-on experience with handling data from consumer wearable devices (eg, actigraphy, PPG, EEG); familiarity with machine learning workflows, including model development, model tuning, and deploying models at scale; Leadership and/or project management experience with the ability to oversee a team of people ingesting data
Management Responsibilities
Foster a collaborative coding and research environment, driving skill development for junior and mid-level data engineers and analysts across the data pipeline
Serve as the primary technical liaison to senior management, translating high-level research aims into actionable objectives
Communicate team progress, bottlenecks, and milestones
Responsibilities
Produce clean, well-documented, efficient code across the entire stack
Design, develop, and deploy robust, scalable applications (both front-end interfaces and back-end data pipelines) to support large-scale research
Lead the engineering workflows to acquire, ingest, and clean multimodal datasets, ensuring efficient storage, retrieval, and processing of massive datasets (+1million records)
Architect and maintain scalable infrastructure to support advanced machine learning models using physiological features and sleep microarchitectures
Optimize application performance and scalability through performance tuning, code refactoring, and database optimization techniques
Stay updated with emerging industry trends, academic literature, and technologies in data engineering, cloud architecture, and machine learning to continuously improve the lab’s technical capabilities
Ensure data integrity throughout engineering workflows. Assist in the preparation of Standard Operating Procedures (SOPs), analytical frameworks, and technical documentation
Maintain open communication with leadership, advise on technical processes, and curate progress reports
Provide technical support and oversight to team members with less experience
Assist in regulatory support, Data Sharing Agreements, and other project documentation
More Maryland jobs
Maryland jobs · Browse all locations