More on the topic…
Google Search dominates European search with roughly 90% market share, giving it an enormous advantage: access to hundreds of millions of user interactions that reveal what people actually search for, how long they stay on results, and which sites they find useful. This behavioral data is gold for training search algorithms. When you search for "grand prix," bounce quickly to try "grand prix attack," click a Reddit thread and spend time there, then watch a chess video, Google learns that you wanted chess content, not Formula 1 racing—and that Reddit and specific YouTube creators are reliable sources for that query. Competitors can't match this. They operate with a fraction of the data, and as Google's own researchers noted, more data has an "unreasonable effectiveness" for machine learning tasks. Without access to similar signals, no rival search engine can compete.
The European Commission tackled this through the Digital Markets Act in 2022, specifically Article 6(11), which requires Google to share anonymized search data—queries, clicks, and viewing patterns—with competitors on fair terms. Google started offering a dataset in March 2024, but nobody took it. The data was useless because Google's anonymization approach stripped out so much information that competitors couldn't actually learn anything useful from it. The Commission opened formal proceedings in January 2026 to force Google to do better. The author, a privacy researcher, was hired to help design a solution that would let competitors access genuinely useful data while still protecting user privacy.
The core tension is this: anonymization under the DMA must prevent anyone from identifying individual users who submitted queries, but it only needs to protect end users—not the people or companies mentioned in search results. So if Alice searches for "Bob's company scandal," it's fine to share that query and result; you just can't tell that Alice made the search. The law also allows a hybrid approach combining technical measures (actually modifying the data) with non-technical ones (like access controls and contractual restrictions). The technical anonymization doesn't need to be perfect; it just needs to reduce re-identification risk to an acceptable level, which other safeguards then protect. This is unusual territory—the shared data still counts as personal data under GDPR, so privacy obligations don't disappear just because it's anonymized.
Questions about this article
No questions yet.