The short version
- The prize compounds: Snowflake (opens in a new tab) reported 828 customers with more than $1 million of trailing product revenue and a net revenue retention rate of 126%, meaning existing customers spent 26% more than a year earlier.
- Enterprise reach is the battleground: Databricks (opens in a new tab) says over 60% of the Fortune 500 use its platform, and IBM (opens in a new tab) says Confluent serves more than 40% of the Fortune 500.
- The vendor map is being redrawn: IBM agreed to buy Confluent for $11 billion, Salesforce (opens in a new tab) agreed to buy Informatica for about $8 billion, and Fivetran and dbt Labs (opens in a new tab) merged in 2026.
- Data practitioners already work inside AI tools: in dbt Labs’ 2025 survey (opens in a new tab) of 459 data professionals, 80% used AI in their daily work, up from 30% a year earlier.
- The buyer’s pain is specific: poor data quality was the top challenge, cited by over 56% of the same respondents, and data quality tools were the second-largest area of planned new investment.
Who chooses a data platform, and how much can one customer spend?
A chief data officer or head of data platform leads, with architects, security, finance and often a systems integrator.
Data infrastructure covers the pipes and platforms that move, store, transform and serve data: ingestion and integration tools, warehouses and lakehouses, streaming platforms, transformation, quality and governance. The buyer is usually a data leader with an architecture problem, not a developer picking a tool for one app. That is the main difference from choosing a database for one application: here the decision is the shape of the whole data estate. The cloud underneath is a separate choice, covered in how cloud providers win enterprise deals through AI.
The money is in expansion. Most platforms charge for consumption, so revenue grows as the customer moves more workloads onto them. Snowflake’s second quarter of fiscal 2027 shows the pattern:
- Product revenue of $1.49 billion, up 37% year on year.
- 828 customers with trailing 12-month product revenue above $1 million, up 27%.
- 829 Forbes Global 2000 customers, and $9.00 billion of remaining performance obligations.
- 692 net new customers in the quarter.
Databricks, which is private, says more than 15,000 customers use its platform and valued itself at over $100 billion in its 2025 funding round. Confluent, the streaming company, had more than 6,500 clients when IBM agreed to buy it. Fivetran and dbt Labs say their combined business supports more than 100,000 data teams.
Budgets are moving in the right direction for sellers. The dbt Labs survey found that 30% of respondents reported budget increases for their data teams, against 9% the year before, while 40% reported headcount increases.
How far into data platform decisions has AI already reached?
In the daily work of the people who shape the decision, and in buyers’ research before vendor calls.
The people who evaluate data platforms use AI constantly. In the dbt Labs survey, 80% of data practitioners used AI in their daily workflow, and 70% used it for analytics development, mainly through general-purpose assistants such as ChatGPT, Claude and Gemini. The same tools they use to write code are a short step from the tools they ask about architecture. The same practitioners weigh in on reporting tools, covered in how analytics and BI tools get shortlisted.
Technology buyers in general use AI in research and check it. In TrustRadius’s 2026 survey (opens in a new tab) of nearly 2,500 technology buyers and vendors, 63% of buyers used AI during their purchase journey, 94% of those fact-checked its answers at least some of the time, and 83% shortlisted three or fewer products. Because TrustRadius sells review visibility, its numbers are best read as one input among several. The finding matters for data infrastructure because a shortlist of three is a small door for any platform.
Which questions do data leaders ask AI assistants?
Architecture, comparison, migration, cost and vendor-risk questions. These prompts are our own illustrations of how data leaders phrase architecture questions; we did not collect them from real buyers.
| Decision | Illustrative prompt |
|---|---|
| Architecture | “Should a 2,000-person insurer build a lakehouse on open table formats or stay with a cloud data warehouse?” |
| Streaming | “Do we need Kafka for real-time fraud scoring, or can our warehouse handle it?” |
| Integration | “Which data integration tools have reliable connectors for SAP and Salesforce with change data capture?” |
| Cost | “How do Snowflake and Databricks compare on cost for 50 TB of mostly SQL analytics?” |
| AI readiness | “What data stack do we need before we put AI agents on our customer data?” |
| Vendor risk | “What changes for Informatica customers after the Salesforce acquisition?” |
| Governance | “Which catalogs support column-level lineage across dbt, Spark and our warehouse?” |
Two features of these questions favor some vendors over others, we infer. First, they are framed around architectures, not products, so an answer usually describes a pattern and then names the vendors that fit it. A vendor that is not tied to the pattern in public writing is easy to leave out. Second, the vendor map changes fast: an assistant answering from older training data may still describe a merged or acquired company as independent, as we explain in do AI assistants answer from training data.
How does an AI answer turn into consumption revenue for a data platform?
Through the architecture decision: the AI answer frames the pattern, the shortlist follows, the proof of concept decides.
- A data leader or architect asks an assistant how to solve a problem: real-time data, AI readiness, a migration.
- The answer describes an architecture and names a few vendors that fit it.
- The team checks those vendors with peers, analysts, partners and documentation.
- Two or three run a proof of concept on the buyer’s own data, often with a systems integrator.
- The winner signs a consumption commitment and then grows as workloads move.
The AI answer matters most at step 2, where the frame is set. A vendor framed as “the streaming option” or “the cheaper warehouse” enters the evaluation with that label. Step 5 is where the value is: a net revenue retention rate of 126%, as Snowflake reported, means each year’s customers keep spending more. Losing step 2, by this logic, costs a stream of revenue rather than one deal. That is our inference from the published figures, not a measured link.
Tracing a consumption contract back to an AI answer is hard. A buyer who met you in an AI answer may arrive through a partner, a cloud marketplace or an inbound demo request months later. Our article on what lost clicks mean for pipeline looks at how to account for that delay.
Why does an assistant name one warehouse, lakehouse or pipeline vendor and not another?
No assistant publishes its vendor-selection rules; studies point to independent coverage, and data teams test claims on their own data.
Documented by the platforms. Google documents that AI Overviews and AI Mode may run several related searches (opens in a new tab) for one question, and OpenAI documents that ChatGPT search rewrites questions (opens in a new tab) into targeted queries. Neither explains how a warehouse or integration vendor earns a place in the answer.
Observed in studies. In our four-assistant study, assistants agreed on brands more for B2B software than for any other industry, with an overlap of 0.543 on a scale from 0 to 1. That still leaves plenty of disagreement: a vendor named by one assistant may be missing from another. Of the signals in our brand study, independent coverage mattered most, with each tenfold rise in independent sites naming a brand going with 4.7 times the odds of a recommendation.
What data buyers check. Data leaders verify with tests on their own data, reference customers, partner opinions and documentation. TrustRadius found analyst reports were used by only 13% of buyers to make purchase decisions, a 63% fall since 2022, while 74% of buyers used reviews.
Our inference. Data platforms are judged on specifics: connector coverage, supported formats, latency, governance features, pricing units and where data stays. Those are facts an assistant can repeat only if they are public and consistent. A reasonable expectation is that vendors with clear, current, independently confirmed specifics are described more accurately and named more often for the right architectures. That idea has not yet been tested on warehouse, streaming or integration vendors.
What slips away when a data platform is absent from architecture answers?
Lost architecture decisions, which are rarely reopened, though no study has measured the cost directly.
- Platforms are sticky. Once pipelines, models and governance are built on a platform, switching means rebuilding them. A vendor left out of an architecture decision may wait years for the next one, we infer.
- Expansion goes to the incumbent. With consumption pricing and retention above 100%, the vendor that wins the first workload usually wins the next ones too.
- Consolidation reshuffles the map. After the Confluent, Informatica and dbt deals, buyers will ask who now owns what and what changes for them. An assistant that answers with outdated information can send buyers elsewhere.
- AI projects raise the stakes. The dbt Labs survey found 45% of respondents planned to increase investment in AI tooling, the largest category. Data platforms that are described as “ready for AI” in answers stand to gain from that spending, our inference. Data science platforms chase it too; see how data science platforms win enterprise buyers.
What does GEO look like for a warehouse, streaming or integration vendor?
It connects your platform to the data architectures buyers ask about, through sources assistants already rely on. Whether an assistant then names you stays outside anyone’s control.
- One clear position per architecture. Say plainly which patterns you fit (streaming, lakehouse, integration, quality) and which you do not, the same way on your site, docs, partner pages and marketplace listings.
- Public specifics. Publish connector lists, supported formats, limits, security attestations and pricing units on readable pages. If those pages are hard for machines to load, read when AI agents can’t read your site.
- Independent validation. Reference architectures written by cloud providers and integrators, customer talks at practitioner conferences, and independent benchmarks with published methods give assistants sources other than you. For data vendors, this is the practical side of how brands build authority for AI search.
- Up-to-date corporate facts. After a merger, rename or acquisition, update every profile, partner listing and knowledge source quickly. When an answer still shows a pre-merger owner or an old product name, how to fix wrong brand information in AI answers sets out the steps.
- Architecture content, not only product pages. Write honest guides to the decisions buyers face, including when a competitor’s approach fits better. Roundups of warehouses, pipelines and catalogs published by others also count, as our review of best-of lists shows.
- Tracking by architecture question. Record how ChatGPT, Gemini, Perplexity, Claude, Copilot and Google’s AI features answer your lakehouse, streaming, migration and vendor-risk questions, asking each one repeatedly.
What is still unknown about AI’s role in data platform decisions?
No public study shows whether AI answers change which data platform wins the architecture decision.
- No data-leader survey on AI research. The dbt Labs survey measures AI use at work, not in vendor research; the TrustRadius figures cover technology buyers in general. The broader software-buying picture is in how B2B SaaS companies generate revenue from AI search.
- Company figures are self-reported. Customer counts and Fortune 500 shares come from the companies’ own releases.
- How assistants handle mergers is untested. We found no study of how quickly AI answers reflect acquisitions.
- Revenue effects are the least proven part. The cross-industry evidence is reviewed in does AI visibility drive business results.
Which architecture questions should a data platform test before its next proof of concept?
The warehouse, streaming, migration and vendor-risk questions behind your proofs of concept, and how answers frame you.
List the decisions that lead to your proofs of concept: the architectures, migrations, integrations and vendor-risk questions. Ask each one in several assistants, more than once, because a single answer is a thin sample. Record whether you are named, how you are labeled, which sources are cited, and whether your connectors, pricing and ownership are described correctly. In data infrastructure, the gaps usually trace back to thin independent proof or to ownership and product facts that predate a merger.
For a second pair of eyes, talk to us about a data platform visibility review. We will map how assistants label your platform in the architecture choices that come before proofs of concept and consumption commitments, and flag the gaps most likely to be costing you enterprise pipeline. Tying your platform to the right architectures, publishing connector and pricing specifics and keeping post-merger facts current are covered on our generative engine optimization service page.
Frequently asked questions
Do data engineers use AI to research vendors?
They use AI daily for work: 80% of data practitioners in dbt Labs’ survey. No public study measures their vendor research separately.
Does an acquisition hurt a data vendor’s AI visibility?
It can confuse answers. Assistants that answer from older training data may describe the company as it was before the deal.
Are analyst reports still worth the effort?
Less than before for buyers directly: TrustRadius found 13% of buyers used them to decide, down 63% since 2022.
Should we publish our connector list and pricing units?
Yes. Integration and cost are among the first questions data leaders ask, and assistants can only repeat facts they can read.
How long does it take for GEO to show in data platform pipeline?
Expect quarters, not weeks. Architecture decisions are slow, and proofs of concept add months before a commitment is signed.
Sources
- Snowflake, via 01net (2026-09-02), Snowflake Reports Financial Results for the Second Quarter of Fiscal 2027 (opens in a new tab)
- Databricks (2025-08-19), Databricks is raising a Series K Investment at >$100 billion valuation (opens in a new tab)
- IBM (2025-12-08), IBM to Acquire Confluent to Create Smart Data Platform for Enterprise Generative AI (opens in a new tab)
- Salesforce (2025-05-27), Salesforce Signs Definitive Agreement to Acquire Informatica (opens in a new tab)
- Fivetran (2026-06-01), Fivetran + dbt Labs Complete Merger to Create the Data Infrastructure for Trusted AI Agents (opens in a new tab)
- dbt Labs (2025), 2025 State of Analytics Engineering Report (opens in a new tab)
- TrustRadius, via Demand Gen Report (2026), TrustRadius: AI Has Changed How Buyers Research, But Not What They Trust (opens in a new tab)
- Google Search Central (2025), AI features and your website (opens in a new tab)
- OpenAI Help Center (2025), ChatGPT search (opens in a new tab)
- Underneath (2026), Do ChatGPT, Gemini, Perplexity and Claude agree on brands?
- Underneath (2026), Do Wikipedia and schema make AI assistants recommend a brand?