arXiv · 2609.24407
Auditing Source Exposure in Baidu and Google AI Search
Abstract
AI-generated overviews are becoming an increasingly prominent layer of search interfaces, yet their behavior in Chinese-language search remains underexplored. We conduct a cross-lingual audit of AI overview behavior on Baidu and Google using English queries sampled from MS MARCO and their translated Chinese counterparts. Our analysis examines when overviews are triggered across platform-language settings, which host domains receive visible exposure in Chinese-language overviews, how concentrated that exposure is, and how source overlap varies across settings. We also compare the embedding-based semantic similarity of generated answers for matched query intents. The results reveal substantial differences across platform-language settings in overview availability and visible source exposure. At the aggregate level, the settings exhibit low overlap in visible host-domain inventories, while matched-query answers yield median cosine similarities ranging from 0.701 to 0.813. These findings indicate that answer-level semantic similarity and aggregate source exposure capture distinct dimensions of AI-mediated search. Evaluations of AI search should therefore consider not only the content of generated answers but also how source visibility is distributed across platforms, languages, and information environments.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yibo Li, Enci Guan, Yuedan Cai, Geng Liu, Francesco Pierri. 2026-09-21. Auditing Source Exposure in Baidu and Google AI Search. https://arxiv.org/abs/2609.24407
Cite the original work for its findings. Save a collection to share your selection of sources.