CareerMoonshot

Data Platform Engineer

Hubble Network · San Francisco Bay Area

📍 San Francisco, CA💰 $250,000via greenhousePosted 2026-09-22
Apply on company site ↗
Career Moonshot pulls this listing straight from the employer's hiring system — no recruiter middleman, no reposts. Applying takes you directly to Hubble Network.
Hubble Network was founded with the intention of delivering on the promise of what Internet-of-Things (IoT) was supposed to be. We're building a global Bluetooth® network dedicated to machine-to-machine connectivity. We differentiate ourselves as the first modem-less and gateway-less, direct-to-satellite network from off-the-shelf Bluetooth® Low Energy chips. Hubble is ideal for applications in logistics, AgTech, and maritime where economies of scale for volume consumer and enterprise asset tracking is a priority. Our goal is to be the first billion-endpoint-connected network in the world. Hubble is an early-stage, venture-backed startup supported by some of the best investors in the world. In their previous lives, the founding team has been successful in raising $100s of millions in venture funding, developing the Amazon Sidewalk network, launching billions of dollars of space assets, and leading their teams to successful exits, both through acquisition and IPO. We are now looking to bring on talented team members who are the best at what they do to help us make Hubble a reality for the world. About the Role This role will be based in San Francisco, CA This is the founding data hire. You will build Hubble's data platform from the warehouse layer up, make decisions on architecture, determine processes for managing schema and data models, and work with many stakeholders. We know what we want: a layered lakehouse with Landing, Staging, Warehouse, Mart stages. Each layer has an owner, a purpose, and a quality bar. You’ll work with the Infrastructure team to manage the pipes: Fivetran, Redshift, orchestration, monitoring. You own everything downstream of raw: the transformations, the models, the metric definitions, the quality gates, and the performance bar.  The guiding principle is data democratization. Every person at Hubble should be able to answer their own questions. Success is not how many queries you run for other people, it's how few they need you to run. You'll be a peer to Platform Engineering and to the team building our internal operations platform. You'll be an embedded consultant to Finance, Product, and Customer Success. You’ll think about how to present a strong story from data. We expect you to use AI agents as part of how you work, and data work rewards this more than most: schema exploration, model scaffolding, test generation, and documentation are all places where agentic tooling earns its keep if you review the output critically. Key Responsibilities Own the transformation layer end-to-end: Defined models across Staging, Warehouse, and Mart. Star-schema design, SCD2, and data contracts enforced by tests at every layer. Build the data layer our internal tools run on: mart tables designed for specific operational workflows, sitting behind a generated internal metrics API. You own the mart and the API's performance characteristics; Engineering owns the framework and the frontend. Hold the line at p95 under 5 seconds on mart queries and under 500ms on metric endpoints. Make freshness a commitment, not an aspiration: core data and critical operational metrics streamed in real time with Apache Iceberg, with monitoring that catches breakage before a stakeholder does. Instrument the funnels: Design event collection across our different product surfaces, unified into a customer 360 spanning authentication, usage, and billing. Acquisition, activation, engagement, retention, revenue. Data in space: We collect large amounts of telemetry from our satellites which is critical to mission operations - help our Mission Operations group land this telemetry in dashboards and monitoring systems so they can take action on data immediately. Automate revenue reconciliation: Daily comparison of Stripe Platform billing against Hubble usage, with discrepancies flagged before month-end and drill-down to the device and packet level. Turn a multi-day manual close into a few hours. Define metrics once: Active devices, packet volume, MRR, contract utilization. Every number has a canonical definition, an owner, and visible lineage from source to dashboard. “Which revenue number is correct?” should stop being a question anyone asks. Implement classification in the pipeline: Security defines the tiers and the guardrails. You tag datasets by sensitivity, enforce retention, masking, and anonymization, and keep lineage and access logs auditable. Hubble does not store PII, and the pipeline is where that commitment either holds or quietly fails. Build for self-service: Semantic layers that let business users work with customers and revenue rather than join keys, curated datasets, a data catalog, and documentation good enough that people stop asking you. Choose boring technology: Fivetran, dbt, Redshift, Metabase are all options - we lean towards buy over build. Bring a strong opinion about which is which, and be able to defend it. Technical Requirements Required 7+ years building production data systems, with 2+ years at senior or staff scope owning architecture rather than executing someone else's. Deep SQL and dimensional modeling: star schemas, slowly changing dimensions, grain discipline. You can explain fact tables and grains. dbt (or similar)  in production: model organization, macros, incremental strategies, tests as contracts, and CI that blocks bad merges. Cloud data warehouse depth, ideally Redshift (or similar): distribution and sort keys, vacuum and analyze behavior, workload management, and the ability to take a slow query apart and make it fast. Python or Scala for the specialized pipelines: custom connectors against proprietary APIs, backfills, and reconciliation jobs. Event and product analytics instrumentation: you've defined a tracking plan, argued about event taxonomy, and built funnel tables that survive contact with a changing product. Data quality as engineering: anomaly detection, freshness monitoring, and

More San Francisco Bay Area jobs

San Francisco Bay Area jobs · Browse all locations