Concept · Technology
Data advantage
Why the data a company collects only matters if it makes a decision better than a rival could make it.

An asset only when it improves a decision.
On August 5, 2024, Judge Amit Mehta of the U.S. District Court in Washington found that Google had unlawfully maintained a monopoly in general search. In its 286 pages was a description of how the monopoly feeds itself. Google receives nine times more queries each day than all its rivals combined, and nineteen times more on mobile. Its NavBoost ranking model runs on 13 months of Google click-and-query data, which the court equated to more than 17.5 years of Bing data. Of 3.7 million unique query phrases analyzed, 93% were seen only by Google and 4.8% only by Bing.
A data advantage exists when a company's own data makes its decisions or product better than rivals can match with the information they can get.
The court describes Google's internal training as treating positive user reactions as one signal of relevance. That signal is a proxy, not ground truth: clicks can reflect placement, query and user context as well as result quality. The ruling supports the mechanism Google said it used; it does not establish that each click improved the next search. Query volume also correlates with product design, engineering, defaults and other inputs, so its standalone effect is difficult to isolate.
A data advantage exists when data a business can use improves a specific decision or product outcome in a way a rival cannot match at comparable cost. The data need not be secret, but it must be meaningfully harder to obtain or reproduce. Test the decision, the measurable improvement, the rival's alternative and the cost of collecting and using the data. A large archive that no decision depends on is a storage cost.
It is also different from a network effect. A network effect exists when participants make the product more valuable to other participants. A data advantage exists when useful information from activity improves a decision or the product; users contribute only when their actions generate data that can be used. The two can coexist, but a growing user count proves neither one by itself.
Forms of data advantage
Data advantages differ in where the data comes from and how quickly it goes stale. The source determines how easily a rival can copy it.
BehavioralClicks, queries and choices made by users of your product
01
Clicks, queries and choices made by users of your product
OperationalRecords from your own fleet, plants or processes
02
Records from your own fleet, plants or processes
MarketObservations of prices or demand you do not control
03
Observations of prices or demand you do not control
A continuum, not a switch
A data advantage grows as proprietary data improves a decision and the decision generates more data. It shrinks when rivals can buy similar information, or when the data cannot predict what the decision depends on.
“Unique data improves over time.”
Why it matters
A data advantage can strengthen with use when new activity produces useful information, a better decision changes the customer outcome, and that outcome attracts more relevant activity. If any link fails, volume alone does not compound quality. The Google court treated scale as one barrier to entry: a new search engine needs users to learn from and needs to be useful to attract them. Judge Mehta observed that Microsoft has run Bing since 2005 and bought Yahoo's search data in 2009, and Bing remains well behind in absolute scale.
The claim also has to be tested, and it can be contested. Google argued that Bing had reached diminishing returns, and its expert concluded that only 2.9% of the quality gap between the two engines came from the volume of user data. The court gave that study little weight, in part because Google had never used the finding for anything beyond the lawsuit: if less data produced the same quality, the company would have stopped storing so much of it. The remedy the court ordered in September 2025 is the strongest evidence of how it viewed the data. Google must share certain search-index and user-interaction data with qualified competitors, so that it does not keep the fruits of its exclusionary conduct. The order excludes advertising data.
The advantage is worth what the decision is worth. Data that changes a decision made millions of times a day has value even when each improvement is small. UPS expects its route-optimization system to cut 100 million miles and 10 million gallons of fuel a year, and drivers on ORION routes already average six to eight fewer miles a day. A saving of that size requires the company to act on the data: the drivers, dispatch systems and rules must change. Data without a workflow to use it is a report.
The practical test is the counterfactual. Ask what a competitor with public information and a good analyst would decide, then measure the gap. If the gap is small or cannot be measured, the company has data but not an advantage.
Real-world examples
The same concept shows up in different ways across industries.
The August 2024 findings in United States v. Google describe queries as the raw material for ranking: Google sees nine times as many queries…
UPS built ORION from data gathered by its drivers' handheld devices and vehicle telematics, and spent about a decade developing it before the 2013…
Zillow had more than 220 million average monthly unique users and the Zestimate, and about three and a half years earlier had decided to…
When it breaks
Zillow shows the failure most managers do not plan for: the data is real, the scale is real, and it answers the wrong question. Then the pandemic froze the housing market and prices began to rise at an unprecedented rate.
The Q3 2021 results show the pattern. Zillow bought 9,680 homes in the quarter and sold 3,032. It wrote down about $304 million of inventory because it had bought homes at higher prices than its estimates of future selling prices, and said it expected a further $240 million to $265 million of losses in Q4. Its own explanation is that higher conversion rates than it had seen before led it to purchase homes unintentionally at those prices. The letter lists a pandemic, a temporary freeze of the market and a supply-demand imbalance as the shocks, and the lesson is not that the data was bad. Knowing what homes had sold for, and what shoppers looked at, did not tell Zillow what the market would pay when it resold months later.
A behavioral dataset fits a problem where the system is stable and feedback is immediate: a search result gets a click within seconds. A market forecast has neither property. Buying at risk turned a data question into a balance-sheet question, and errors were paid for in cash. The company kept the audience and the Zestimate, and it said it would build asset-light ways to help people sell.
Advantages built on data also fail when the law changes what a company may keep or must share, as the Google order shows, and when rivals reach similar quality without matching the volume.
Key takeaways
- 01
Which decision should this data improve, and what outcome would show an improvement over a credible rival or baseline?
- 02
Does use generate timely, relevant data that changes the next decision, or does the dataset add collection and storage cost without changing action?
- 03
Is the observed data a reliable signal for this decision, or a proxy shaped by selection, ranking or a different time horizon? What would disconfirm it?
Sources
- United States v. Google LLC, Memorandum Opinion (liability), Case No. 20-cv-3010 (APM), Document 1033 · U.S. District Court for the District of Columbia, 2024-08-05. Findings of fact 86-89 (nine times and nineteen times query volume; 93% and 4.8% of unique phrases); Part V.A.2.b (NavBoost, 13 months equals 17.5 years of Bing data); Part V.A.2.c (Fox study, 2.9%, diminishing returns, Bing history)
- United States v. Google LLC, Memorandum Opinion (remedies), Case No. 20-cv-3010 (APM), Document 1436 · U.S. District Court for the District of Columbia, 2025-09-02. Introduction: sharing of certain search index and user-interaction data, not ads data, with Qualified Competitors
- UPS Routing Program ORION Helps Drivers Trim Miles, Reduce Costs · Transport Topics, 2016-08-09. Decade of development, 2013 rollout, expected 100 million miles, 10 million gallons and $300-400 million a year; DIAD telematics data; algorithm tailored to UPS; drivers encouraged to use judgment
- Analytics Success Story: UPS's ORION Project · InformIT. $250 million to build and deploy; six to eight fewer miles per route per day for drivers using ORION
- Zillow Group Q3 2021 Shareholder Letter (Exhibit 99.3 to Form 8-K) · U.S. Securities and Exchange Commission, 2021-11-02. Wind-down; approximately 25% workforce reduction; market maker versus market risk taker; three to six month price forecast; 200 basis point guardrail; 1,200 basis point swing; 220 million monthly users
- Zillow Group Reports Third Quarter 2021 Financial Results (Exhibit 99.1 to Form 8-K) · U.S. Securities and Exchange Commission, 2021-11-02. $304 million inventory write-down; 9,680 homes purchased and 3,032 sold in Q3; additional $240-265 million of Q4 losses; unintentional purchases at higher prices