Arena: Market Conditions Before Snowflake
Companies' data and analytics teams · 2012 · United States
Changing technical requirementsEnabling technology shiftHigh setup and upkeep costs
In 2012 a company that wanted to analyse its data had two main choices. An analytic warehouse such as Teradata or Netezza ran on a fixed pool of machines sized in advance, needed administrators to plan capacity, tune it and apply upgrades, and handled poorly the web and application data, often in JSON, that companies were now collecting SN11 SN1 SN2. Hadoop stored anything cheaply but was a toolkit: teams wrote code and hired specialists to get answers from it SN13 SN2. Demand for analysis was also becoming less predictable, with large jobs that needed a lot of computing for a few hours and then none. Public clouds had begun to rent computing and storage on demand, but the warehouse software itself still assumed it owned its machines SN2.
How each step happened
Step 1 of 5 · 2012–16
Innovation: the warehouse rebuilt for the cloud, storage apart from compute
Snowflake's founding bet concerned how a warehouse should use the cloud. Warehouses of the day, Amazon Redshift included, used a shared-nothing design: each node kept a slice of the data on its own disks, so compute and storage grew together and every resize reshuffled data SN2 SN13. Snowflake kept one copy of the data in cloud storage and put as many separate compute clusters against it as customers needed, each able to start, grow or stop without touching the data SN2 SN13. Customers got compute sized to each job, which Redshift could not offer without being rebuilt.
Benoit Dageville and Thierry Cruanes, database architects at Oracle, left in 2012 rather than pitch it there, judging that a company with a mature database could not build it SN1 SN8. With most people expecting Hadoop to win, their 2016 paper calls a classic SQL warehouse written from scratch a contrarian and risky choice SN7 SN2.
They ran it only as a service, with no tuning, table statistics or vacuuming for users and one production version the company could fix quickly SN2. Mike Speiser of Sutter Hill, chief executive until June 2014, kept it in stealth for an 18 to 20 month technology head start, the first salesperson, Chris Degnan, recalls SN1 SN22. Under Bob Muglia the service became generally available on Amazon's cloud in June 2015 SN1 SN2. The separation made possible, first, a new way to buy compute.
Rivals Amazon Redshift launched in November 2012 on technology licensed from ParAccel, with nodes holding 2 or 16 TB, so more compute meant more nodes and redistributed data SRX-1 SN2. It separated storage from compute only with RA3 nodes in December 2019 SRX-2. Hadoop, Muglia said, was a toolkit that needed code and specialists SN13.
Novel ArchitectureFounder Domain ExpertiseManaged Service
Step 2 of 5 · 2015–21
Each workload gets its own compute, billed by use
Because compute no longer held the data, each workload could get its own compute cluster, which Snowflake calls a virtual warehouse. Data loading, analysts' reports and a data science job ran on separate machines, so one did not slow another; each shut off when idle, and customers paid only for the storage and compute they used SN2. Compute is now billed by the second SN25.
It was also easier to run. In 2021 the fintech Tide left Redshift because its cluster kept running out of space and one bad query could force a restart and hours of reloading; on Snowflake it scaled compute in seconds, and it says maintenance, not cost, drove the move SRX-5.
By 2020 most customers signed annual capacity commitments, could use more than they had bought, and rolled unused capacity forward when they bought more SN1. Revenue followed use, not a fixed subscription, so a release that made queries cheaper cut near-term revenue SN1. But every new workload was new revenue, and finding workloads became the sales force's job.
Rivals Redshift was priced by the node-hour: $0.85 an hour on demand, or under $1,000 per terabyte a year reserved SRX-1. It added Concurrency Scaling for queued queries in March 2019, and per-second billing only with Redshift Serverless in July 2022 SRX-3 SRX-7.
Usage-Based PricingEase of Use
Step 3 of 5 · 2017–22
Land one workload, then migrate the rest
Sellers landed one workload, then worked to migrate the rest of the customer's SN1. Capital One signed in June 2017 and moved its analytics workloads; use spread to other lines of business, and Capital One was 17% of revenue in the fiscal year to January 2019 SN1. A European retailer started in 2018 with one brand's analytics and kept moving workloads SN1. Degnan says the team used its own product to spot customers with new uses SN22.
Net revenue retention compares what customers pay now with what the same customers paid a year earlier. Snowflake's exceeded 150% at January 2019 and 2020 and reached 177% in the fourth quarter of fiscal 2022 SN1 SRX-8. That year 44% of migrations came from cloud platforms SRX-8.
Selling to the largest companies
Degnan had called an early enterprise focus a waste of time, and Muglia had never run a direct sales force SN21 SN22. Frank Slootman, formerly of ServiceNow, took over in April 2019; he says modernizing old analytics workloads worked as an entry strategy but was not where Snowflake could stay SN1 SN20. By the 2020 listing a sales force segmented by customer size was aiming at large enterprises, and customers included 146 of the Fortune 500 SN1 SN26.
Rivals AWS set up a dedicated sales team in 2018 to answer Snowflake's growth and Redshift's losses, and in February 2020 Slootman said Snowflake had taken business from Redshift SRX-9. Teradata and other installed warehouses held many of the large accounts Snowflake migrated SN1.
Land and ExpandTop-Down Selling
Step 4 of 5 · 2017–20
Neutral across clouds, with live data shared between accounts
This step ran alongside the sales push from 2017 and gave sellers an answer no cloud's own warehouse had. Redshift ran only on AWS, Azure's warehouse only on Azure and BigQuery only on Google SN1 SN2. Owning no cloud, Snowflake could run on all three. Muglia called Amazon a partner Snowflake competed with every day SN12. Snowflake reached Azure in September 2018 and Google Cloud in February 2020, with databases replicated across clouds SN16 SN19. A layer hiding each cloud's specifics let customers' applications move with it SN17.
Data sharing grew from the same design. Any compute could reach any data with the owner's consent, so from June 2017 one account could share live data with another without copying it; the provider paid nothing and the consumer paid for its compute, so every share was also consumption SN7 SN14. In March 2020 hundreds of customers queried Starschema's COVID-19 data from their own accounts, and in May FactSet put 24 data sets on the platform for shared clients SN1. The prospectus describes the network effect: the more customers adopt, the more data they can exchange SN1. Slootman made this access the center of his Data Cloud strategy SN20.
Rivals Redshift stayed on AWS only, and its data sharing arrived in preview in December 2020, between Redshift clusters on RA3 nodes SRX-4. Databricks now shares data with recipients outside its own platform SN6.
Counter-positioningCollaboration & exchange network effects
Step 5 of 5 · 2018–24
Switching costs
Each workload migrated in step 3 brought pipelines, permissions and reports built around it. Capital One consolidated analytics, marketing and petabytes of logs on one platform used across its lines of business SN1, and DoorDash engineers describe the ETL jobs and dataset tuning they maintain in Snowflake SN5. Leaving means moving the data and rebuilding that work. The prospectus named customers' investment in older warehouses as a brake on adopting Snowflake; once the workloads had moved, the same brake worked for Snowflake SN1.
The barrier has narrowed. Degnan says Snowflake gave Databricks room by building tools for data science teams too slowly SN22, and since June 2024 Iceberg tables let customers keep data in their own storage, readable by other engines SN3. What holds customers now is mostly the work built on the data, which is why step 3's one-workload-at-a-time migration mattered.
Rivals Redshift matched the architecture late: separated storage in December 2019, data sharing in December 2020 and per-second serverless billing in July 2022 SRX-2 SRX-4 SRX-7. Instead, by 2020 AWS was paying its own field staff on Snowflake consumption, a first, Slootman said SRX-9.
Switching costsData Gravity