Search arXiv⌕ Search

arXiv subjects

Phanish Puranam

Publications and source records attributed to Phanish Puranam.

9 recordsLinked to original sources

Why They Disagree: Decoding Differences in Opinions about AI Risk

Identifying the reasons for disagreements between influential points of view on issues that affect the public can help produce informed policy responses, even if they do not bring disagreeing parties closer to agreement. We present a methodology for extracting reasoning chains - the sequences of premises that motivate or justify opinions - from natural discourse, and for characterizing the types of premises (facts, forecasts, definitions, causal beliefs, and evaluations) that make up these chains. We demonstrate the utility of this approach for two practical goals: diagnosing specific points of contention and aggregating arguments across speakers. We illustrate the methodology through an analysis of the debates on the nature of risks that AI poses to the public, using a corpus of interviews from the Lex Fridman podcast. We find that differences in perspectives among the podcast's guests on existential risk and employment risk from AI arise primarily from differences in causal premises and forecasts, whereas in the case of AI's effects on human social relationships, premises regarding what is valued and definitions about what counts as genuine human connection play a distinctively larger role. Our approach to analyzing reasoning chains at scale, using an ensemble of LLMs to parse textual data, can be applied to facilitate deliberation and aggregation of opinions on any topic.

cs.CY↗

Artificial intelligences and human scientists exhibit complementary strengths in theory building

We investigate the effectiveness of artificial intelligences (AI)-specifically large language models (LLMs)-relative to human scientists at high-level cognitive tasks in social science such as theory formulation, predictions of novel empirical results, and theory revision in response to new evidence. The research domain was academic discourse regarding gender and race inequality. Our findings, comparing 25 LLMs with 13 senior researchers and 60 doctoral scholars, reveal that the AIs outperformed most humans individually on most of the present tasks, while human theories were more diverse and exhibited greater gains in predictive accuracy from aggregation. AI-generated theories were more extensively elaborated, involving additional theoretical paths and latent variables, and were rated as higher quality than human theories by independent raters blinded to source. However, this theoretical complexity was in part ornamental, in that it was not associated with more accurate predictions about empirical patterns in data; in contrast, human scientists achieved greater predictive efficiency with simpler theories. The AIs were significantly more likely than human scientists to revise their theories to incorporate new evidence; human scientists updated their beliefs in a selective way that is sensitive to prior prediction errors. We speculate that the superior processing capacity of artificial intelligences makes them especially well-suited to tasks requiring grappling with complexity, but that the greater diversity of human ideas is essential to wise crowds and collective creativity.

cs.AI↗

How Generative AI Adoption Alters the Demand for Cognitive and Social Skills Within Roles: A Skill-Centric Analysis

A common view holds that generative AI (GenAI) automates cognitive tasks, reshaping roles to emphasize social skills over cognitive ones. Drawing on the framework we develop in this paper, we argue that other outcomes are theoretically possible. We analyze seven million job postings from 595 U.S. public firms that adopted GenAI in 2022-2024, estimating difference-in-differences models comparing GenAI-adopting and non-adopting roles around ChatGPT's launch. We find no evidence of greater emphasis on social skills in GenAI-adopting roles; instead, their relative demand for social skills declined by 3.4%, while demand for cognitive skills held steady. This decline reflects the addition of functional administrative skills to roles outside those functions. Concentrated in non-managerial roles, this pattern points toward GenAI-enabled self-service rather than intensified cross-functional coordination.

econ.GN↗

Predictive AI Can Support Human Learning while Preserving Error Diversity

We examined the effects of predictive AI deployment on the immediate performance and learning of medical novices. In two pre-registered field experiments, we varied whether AI input was provided during the training or practice of lung cancer diagnoses, or both. Our results show that different AI deployments have distinct implications for human professionals. AI input during training or practice independently improves individuals' diagnostic accuracy, whereas deployment across both phases yields gains that exceed either approach alone. Furthermore, AI input in both training and earlier practice can improve the accuracy of individuals' subsequent independent diagnoses. Beyond individual accuracy, AI deployment affects the diversity of errors across individuals, with consequences for the accuracy of group decisions (e.g. when getting a second or third opinion on a diagnosis).

cs.HC↗

Thinking with Many Minds: Using Large Language Models for Multi-Perspective Problem-Solving

Complex problem-solving requires cognitive flexibility--the capacity to entertain multiple perspectives while preserving their distinctiveness. This flexibility replicates the "wisdom of crowds" within a single individual, allowing them to "think with many minds." While mental simulation enables imagined deliberation, cognitive constraints limit its effectiveness. We propose synthetic deliberation, a Large Language Model (LLM)-based method that simulates discourse between agents embodying diverse perspectives, as a solution. Using a custom GPT-based model, we showcase its benefits: concurrent processing of multiple viewpoints without cognitive degradation, parallel exploration of perspectives, and precise control over viewpoint synthesis. By externalizing the deliberative process and distributing cognitive labor between parallel search and integration, synthetic deliberation transcends mental simulation's limitations. This approach shows promise for strategic planning, policymaking, and conflict resolution.

cs.CL↗

Can LLMs Help Improve Analogical Reasoning For Strategic Decisions? Experimental Evidence from Humans and GPT-4

This study investigates whether large language models, specifically GPT4, can match human capabilities in analogical reasoning within strategic decision making contexts. Using a novel experimental design involving source to target matching, we find that GPT4 achieves high recall by retrieving all plausible analogies but suffers from low precision, frequently applying incorrect analogies based on superficial similarities. In contrast, human participants exhibit high precision but low recall, selecting fewer analogies yet with stronger causal alignment. These findings advance theory by identifying matching, the evaluative phase of analogical reasoning, as a distinct step that requires accurate causal mapping beyond simple retrieval. While current LLMs are proficient in generating candidate analogies, humans maintain a comparative advantage in recognizing deep structural similarities across domains. Error analysis reveals that AI errors arise from surface level matching, whereas human errors stem from misinterpretations of causal structure. Taken together, the results suggest a productive division of labor in AI assisted organizational decision making where LLMs may serve as broad analogy generators, while humans act as critical evaluators, applying the most contextually appropriate analogies to strategic problems.

cs.AI↗

Why Trust in AI May Be Inevitable

In human-AI interactions, explanation is widely seen as necessary for enabling trust in AI systems. We argue that trust, however, may be a pre-requisite because explanation is sometimes impossible. We derive this result from a formalization of explanation as a search process through knowledge networks, where explainers must find paths between shared concepts and the concept to be explained, within finite time. Our model reveals that explanation can fail even under theoretically ideal conditions - when actors are rational, honest, motivated, can communicate perfectly, and possess overlapping knowledge. This is because successful explanation requires not just the existence of shared knowledge but also finding the connection path within time constraints, and it can therefore be rational to cease attempts at explanation before the shared knowledge is discovered. This result has important implications for human-AI interaction: as AI systems, particularly Large Language Models, become more sophisticated and able to generate superficially compelling but spurious explanations, humans may default to trust rather than demand genuine explanations. This creates risks of both misplaced trust and imperfect knowledge integration.

cs.AI↗

LLMs as mediators: Can they diagnose conflicts accurately?

Prior research indicates that to be able to mediate conflict, observers of disagreements between parties must be able to reliably distinguish the sources of their disagreement as stemming from differences in beliefs about what is true (causality) vs. differences in what they value (morality). In this paper, we test if OpenAI's Large Language Models GPT 3.5 and GPT 4 can perform this task and whether one or other type of disagreement proves particularly challenging for LLM's to diagnose. We replicate study 1 in Koçak et al. (2003), which employes a vignette design, with OpenAI's GPT 3.5 and GPT 4. We find that both LLMs have similar semantic understanding of the distinction between causal and moral codes as humans and can reliably distinguish between them. When asked to diagnose the source of disagreement in a conversation, both LLMs, compared to humans, exhibit a tendency to overestimate the extent of causal disagreement and underestimate the extent of moral disagreement in the moral misalignment condition. This tendency is especially pronounced for GPT 4 when using a proximate scale that relies on concrete language specific to an issue. GPT 3.5 does not perform as well as GPT4 or humans when using either the proximate or the distal scale. The study provides a first test of the potential for using LLMs to mediate conflict by diagnosing the root of disagreements in causal and evaluative codes.

cs.CL↗

Learning what they think vs. learning what they do: The micro-foundations of vicarious learning

Vicarious learning is a vital component of organizational learning. We theorize and model two fundamental processes underlying vicarious learning: observation of actions (learning what they do) vs. belief sharing (learning what they think). The analysis of our model points to three key insights. First, vicarious learning through either process is beneficial even when no agent in a system of vicarious learners begins with a knowledge advantage. Second, vicarious learning through belief sharing is not universally better than mutual observation of actions and outcomes. Specifically, enabling mutual observability of actions and outcomes is superior to sharing of beliefs when the task environment features few alternatives with large differences in their value and there are no time pressures. Third, symmetry in vicarious learning in fact adversely affects belief sharing but improves observational learning. All three results are shown to be the consequence of how vicarious learning affects self-confirming biased beliefs.

econ.TH↗