Search arXiv⌕ Search

arXiv · 2610.02386

Social bot detection in the age of ChatGPT: Challenges and opportunities

Abstract

We present a comprehensive overview of the challenges and opportunities in social bot detection in the context of the rise of sophisticated AI-based chatbots. By examining the state of the art in social bot detection techniques and the more salient real-world application to date, we identify gaps and emerging trends in the field, with a focus on addressing the unique challenges posed by AI-generated conversations and behaviors. We suggest potentially promising opportunities and research directions in social bot detection, including (i) the use of generative agents for synthetic data generation, testing and evaluation; (ii) the need for multimodal and cross-platform detection based on network and behavioral signatures of coordination and influence; (iii) the opportunity to extend bot detection to non-English and low-resource language settings; and, (iv) the room for development of collaborative, federated learning detection models that can help facilitate cooperation between different organizations and platforms while preserving user privacy.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Emilio Ferrara. 2026-10-01. Social bot detection in the age of ChatGPT: Challenges and opportunities. https://doi.org/10.5210/fm.v28i6.13185

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Validity, Reliability, and Transparency in Artificial Intelligence Regulation

AI systems increasingly produce claims about people that shape access to employment, healthcare, and other consequential services. Yet lawful data processing and predictive accuracy do not establish that these claims justify the treatment that follows. We argue that regulation must therefore examine the legitimacy of the inferential step connecting data to claims and claims to decisions. This requires construct, internal, and external validity: evidence must support the meaning attributed to an output, the relationship asserted, and its application to the intended people and setting. We show why predictive performance alone cannot provide this warrant. Non-identifiability limits what observations can explain, while omissions and confounding require domain-specific judgment about what the evidence supports. Building on validity provisions in the EU AI Act and NIST AI RMF Playbook, we develop an explicit evidentiary burden for consequential reliance. Its normative basis follows from the dignity, autonomy, and informational-privacy principles in \emph{Puttaswamy} \citep{Puttaswamy2017Privacy}, which we extend from scrutiny of information acquisition to the justification of derived claims and their use. On this account, benefits that depend on an inference can carry weight in proportionality assessment only to the extent that the inference is substantiated. We make this requirement operational through a claim-specific \emph{Validity Case} linking evidence and assumptions to permitted uses, independent review, monitoring, and remedies. Hiring and LLM applications in healthcare and legal assistance illustrate how the framework can guide oversight of consequential public and private services through the relevant legal instruments.

cs.CY↗

Curriculum Brain: Constructing Curriculum Knowledge Graphs as a Substrate for Cognitive Diagnosis

Cognitive Diagnostic Models (CDMs) identify which specific skills a student has and has not mastered, the signal a personalized learning path needs and a single aggregate score cannot give. Yet they are rarely deployed. The obstacle is their precondition: the Q-matrix, a mapping from every assessment item to the skills it requires, historically authored by hand. We separate the task into two stages: first construct the curriculum's own knowledge graph, the full space of concepts and skills it contains, independent of any item; then map items against that graph on demand. This paper addresses the first stage only. The item-mapping stage is designed but not implemented here, so the claim that this shifts judgment cost from once per item to once per curriculum is a design rationale rather than a finding. We present Curriculum Brain, a two-repository system pairing a version-controlled knowledge base with an agentic pipeline of eleven single-responsibility agents under a thin deterministic orchestrator. It generates candidate concept-skill mappings from official curriculum documents, checks them against accumulated rules, and compares them with a concept-skill map extracted independently from the textbook, repairing its own failures and escalating to a human only when it cannot resolve a case itself. Across 241 chapter runs (168 distinct chapters), 41.5% produced a Generator output passing both checks without a patch, and 67.6% resolved without escalation. Both are measured against criteria the system itself produced, so both describe internal consistency rather than agreement with an external standard, and both pool two pipeline configurations separated by a single change at run 77; after it the figures are 57.0% and 91.5%. Observed spend was $1.19 per chapter, API spend only, excluding human review. We release both the framework and the resulting curriculum dataset.

cs.CY↗

AI-Generated Disinformation in the UK: Risk of Harm, Context, and Classification

More people across the UK are today exposed to more AI-generated and AI-altered disinformation and misinformation than ever before - exposing individuals and society to a growing range of harmful consequences and potential consequences. We analyse evidence from a dataset of 112 pieces of AI-generated or AI-altered disinformation or misinformation seen tens of millions of times across the UK between 1 January 2025 and 31 March 2026. The dataset of examples studied does not, of course, provide an exhaustive sample of all AI-generated false information in circulation in the period, but rather a snapshot of examples showing some, but not all, of the potential effects. Our analysis of this sample found a substantive risk of causing or contributing to harm to individuals and society in eight distinct fields, from contributing to incidents of serious social unrest and vigilante violence to causing direct harms to health and causing the sort of serious financial loss that can be caused by online scams and fraud. Other fields of risk included: abuse serious enough to affect individuals' health and behaviour; public engagement with the police and justice systems; susceptibility to false conspiracy theories with the potential to cause direct harms; and broader changes to social and political attitudes with potential to affect political, social events over the longer term. More than four in five pieces of content we assessed added to reasons for the public to distrust information as not merely inaccurate but substantively false or misleading: a broad disinformation effect with potential for significant effects for society.

cs.CY↗