Search arXivSearch

arXiv subjects

Benjamin Tannenbaum

Publications and source records attributed to Benjamin Tannenbaum.

6 recordsLinked to original sources

Scoring With the Engine: Retrieval Exposure, Cross-Engine Divergence, and the Limits of Engine-Agnostic GEO Scores

Recent work asks whether generative-engine visibility can be approximated with deterministic, engine-free page scores. We separate two stages such scores can conflate: exposure to a live engine and citation selection conditional on exposure. In an observational audit of ChatGPT, Microsoft Copilot, Google, and Perplexity, 15 fixed commercial prompts produced 589 citation observations on 6 June 2026, corresponding to 528 unique URLs and 356 domains. Same-prompt cross-engine URL overlap was extremely small: mean pairwise Jaccard similarity was 0.0079, the median was zero, and 84.9% of engine pairs shared no cited URL. On the ten prompts observed on all four engines, mean exact-URL Jaccard was 0.0072. A matched-size hypergeometric baseline preserving each prompt's four-engine URL universe and each engine's list length predicts 0.1272, so observed overlap was only 5.7% of that baseline; zero URL overlap occurred in 86.7% of comparisons versus 12.3% expected. Top-five exact-URL overlap was zero in all 60 pairwise comparisons. A single engine captured only 11.4%-42.6% of the four-engine URL union, and 96.4% of observed URLs appeared in only one engine. A separate 5-to-6 June same-engine comparison found 67.0% mean URL-set turnover. These results do not invalidate engine-free page scoring; they identify its estimand. A score computed without a live engine can estimate page quality or query-page fit, while end-to-end visibility additionally depends on engine-specific exposure and selection. We therefore argue for reporting page fit, observed exposure, conditional selection, and final visibility as distinct quantities.

cs.IR

From Prompt to Recommendation: A Fitted Stage Model of Brand Visibility in AI Search

We analyze 34,960 unbranded prompt-engine observations from 75 anonymized Aiso projects, covering 2,854 distinct monitored prompts and repeated GPT and Gemini runs from June-September 2026. When neither the target brand nor its own domain appears in the observable live retrieval path, target mention rates are 2.8% for GPT and 3.8% for Gemini. With an own-domain citation but no branded fan-out, they rise to 49.0% and 58.4%. When both own-domain exposure and a branded fan-out occur, mention rates reach 91.4% and 100%. The relationship persists within the same project, prompt, and engine across repeated runs: among prompt cells that vary in own-domain exposure while holding branded fan-out absent, exposure is associated with a mean mention-rate increase of 40.2 percentage points on GPT and 49.0 points on Gemini. Prior visibility is independently persistent. A previous non-mention plus no current own-domain exposure yields next-run mention rates of 1.6% and 1.9%; previous mention plus current exposure yields 80.5% and 83.7%. We fit a chronological diagnostic model using prior-run history and contemporaneous retrieval indicators: $ \operatorname{logit}P(M_t=1)=α_e+β_e\operatorname{logit}(\widetilde P_{t-1})+γ_e E_t+δ_e F_t+θ_e^\top X. $ On the latest 30% holdout, the full model achieves AUC 0.963 on GPT and 0.942 on Gemini, compared with 0.937/0.917 for prior history alone and 0.880/0.840 for live signals alone. A manually curated prompt sensitivity gives nearly identical AUCs (0.960 and 0.943). A separate 199-prompt page-corpus validation finds that prompt-page match predicts Gemini exposure (AUC 0.641) more clearly than GPT exposure (0.545), placing relevance upstream of a larger engine-mediated exposure effect. The equation is predictive and observational, not a causal description of proprietary engine internals.

cs.IR

Purchase Advice and Observable Buyer Responses in Real AI Conversations

How often does a generative assistant persuade someone to buy, or persuade them not to buy? Conversation logs contain recommendations, but they do not necessarily record subsequent decisions. We audit 317 historical interactions from Aiso's proprietary research database of licensed, consent-based, de-identified conversations with commercially available AI assistants. Single-agent AI-assisted screening identifies 68 purchase-directed records; collapsing one shared-prefix copy yields 67 retained episodes, dated April 2023 to July 2025. Assistant responses provide candidate options, acquisition channels, or conditional preferences in 52 episodes (77.6%). One episode contains conditional redirection away from a named accommodation candidate. No episode is coded as advice to abandon or defer the purchase category. Only 18 episodes (26.9%) contain a subsequent user turn within the same purchase-related mission, compared with 23 (34.3%) that contain any later user turn. Using conversation depth alone therefore overstates this follow-up availability by 27.8%. Across 47 retained user follow-up messages, no explicit post-advice purchase commitment, completed-purchase report, or purchase-category abandonment statement is observed. These zeros describe recorded statements, not conversion or persuasion rates. The paper supplies operational definitions, text-free annotations, and reproducible descriptive results. Its central finding is a measurement limitation: recommendation content is observable much more often than a buyer's subsequent decision. The selected historical sample, unvalidated AI annotations, and missing transaction outcomes do not support a population-level or causal estimate of persuasion.

cs.IR

Beyond the Final Prompt: Measuring the Effect of Within-Conversation Context on AI Answers

An isolated final user message is often treated as the query in evaluations of AI systems. In a conversation, however, the actionable request may be distributed across preceding turns. We directly test whether that omitted within-conversation context changes answers. For each of 180 English multi-turn conversations sampled from a governed commercial corpus and the public PRISM dataset, we hold the final user message and requested answer model constant while generating three answers: one from the full role-labelled conversation, one from the final message alone, and one from the final message plus a prefix-only reconstruction capped at 160 words. A separately requested judge model evaluates answers under randomized labels. The prespecified primary endpoint is a material difference that could change what the user does, rather than a difference in style or detail. After inverse-probability weighting to the eligible cohorts, the full-conversation and isolated-final answers differ materially in 44.7% of cases (95% bootstrap CI 33.8% to 56.1%). Full-conversation answers score 0.49 points higher on a 0 to 4 request-satisfaction scale (0.32 to 0.67). Adding the compressed prefix reduces the material-difference rate to 30.8% (20.2% to 42.1%), a 13.9-point reduction (4.9% to 24.1%), and reduces the mean satisfaction gap to 0.01 points (-0.12 to 0.13). Yet compression is not equivalent to the complete dialogue context: almost one third of answers remain materially different. An order-swapped repeat on 48 cases yields 91.7% agreement and kappa = 0.83 for the primary decision. The study concerns preceding turns in the same conversation and does not test persistent memory across separate conversations.

cs.IR

The Prompt Is Not the Query: How Request State Evolves Across Multi-Turn AI Conversations

AI-search evaluation commonly treats a prompt as a stable query that can be counted, classified, and replayed in isolation. A conversation makes that unit of analysis questionable: each user turn can add a constraint, revise an assumption, request evidence, or refer to alternatives established earlier. We replace latent "intent" with an observable construct, conversation-conditioned request state, and measure how that state is distributed across user turns. The analysis reuses frozen rules and the governed cohort of a preceding conversation study: 670 English commercial multi-turn conversations in a discovery-replication design and 7,463 public PRISM conversations from 1,389 participants. In the commercial corpus, the final prompt contains a median 35.6% of the session's unique user-side content vocabulary; in PRISM, the median is 36.4%. The final prompt contains at most half of that vocabulary in 68.4% and 74.3% of conversations, respectively. More importantly, transparent rules detect at least one request-state dimension in history but not in the final prompt in 50.3% of commercial conversations and 44.8% of PRISM conversations. Among dimension-bearing conversations, the final prompt reproduces the full observed dimension set in only 26.1% and 26.2%. At the same time, the final prompt adds a previously unseen dimension in 17.9% and 19.3%, showing that the endpoint is neither a summary nor merely a reference: it is often another state update. Length-matched nulls show that low lexical coverage is largely a consequence of turn length, so vocabulary results are interpreted as information availability, not semantic drift. The categorical results support session-level measurement for AI search. They do not estimate the causal effect of history on model answers.

cs.IR

Answer-Reconstruction Search Density: Measuring the Query and Source Work Compressed by Conversational Answers

Conversational systems can collapse a visible sequence of web queries, result inspections, and source comparisons into a single synthesized answer. Existing retrieval metrics evaluate ranking, effort, or factual support, but they do not quantify the minimum conventional search work represented by a completed answer. We define answer-reconstruction search density (\ARSD): the minimum number of distinct query actions required, under a fixed and dated reconstruction policy, to support a target share of atomic retrievable answer units. A parallel page-density measure separates query compression from source compression.

cs.IR