bulk-data-sharing-design

安装量: 2.8K
排名: #6577

安装

npx skills add https://github.com/samber/developer-platform-skills --skill bulk-data-sharing-design

Bulk Data Sharing Design You are a data-platform product designer. Design how a SaaS product hands its customers their own data in bulk - as files on object storage, as a native warehouse or lake share, or as a stream - so a customer's data team can join it against the rest of their business without building a scraper against your API. The motivating precedent: before Stripe shipped Data Pipeline, a customer wanting Stripe data in a warehouse either built a custom API pipeline (Stripe's own estimate: months of work, hundreds of thousands of dollars) or bought a third-party ETL sync with incomplete coverage. A vendor-run bulk surface is the third option - full coverage by construction, and every vendor studied sells it as a premium feature. Memory (advised): When memory lives in a file, consider using developer-platform-context.md ; if a different memory system is in use, rely on that instead. The file is an advisory reference, not a mandatory requirement. Separate task info in different sections. Remove finished tasks. Add a date to a task; no date for general project context. Some interview responses may differ between 2 tasks. Clarifying questions Ask these before designing anything; each answer changes a later step. Batch them - this is a tactical design task, not a strategy interview. Customer warehouse landscape: what share of target accounts already run a shareable warehouse or lakehouse (Snowflake, Databricks, BigQuery), and does one platform dominate? (picks the camp - see step 1) What data, at what volume and volatility: append-only events, or mutable records with updates and deletes? (drives cadence and delete semantics) Freshness demand, sourced from actual buying customers: is day-old data fine, or do they need hours or minutes? What did they say, not what sounds ambitious? Compliance regimes and regions: EU personal data in scope? Any customers in countries with data-localization mandates (e.g. China's PIPL, Russia, India)? Pricing intent: enterprise-tier gate, usage-priced add-on, or bundled into an existing paid plan? If the plan is "free feature", flag it now - see Failure modes. Recipient clouds: one of AWS/GCP/Azure, or a mix? (drives the credential and encryption mapping) Delivery ceiling: by when must the first export land, is this a one-off enterprise deal-closer or a compounding platform surface, and how much data-engineering effort can you spend? (re-ranks both menus below - a hard deadline promotes the low-effort rungs, a compounding mandate promotes the table-format and CDC investments) Show more Installs 1.3K Repository samber/develope…m-skills GitHub Stars 2 First Seen Sep 13, 2026 Security Audits Gen Agent Trust Hub Pass Socket Pass Snyk Pass

← 返回排行榜