Search arXiv⌕ Search

arXiv · 2609.31219

Research with AI Agents: How Agentic Systems Are Changing Scientific Work

Abstract

Background. Agentic AI systems independently decompose tasks such as literature search, data analysis, and programming into subtasks, search the web, access databases, and execute code. This allows them to perform digital research tasks at high speed. Objectives. Under what conditions does the use of agentic systems produce reliable efficiency gains, and which tasks remain with researchers? Materials and methods. Summary of current studies on literature searches, data analysis, software development, and clinical decision support. Results. For digital activities, work shifts from execution to steering and review. Efficiency gains are greatest when expected behavior can be formalized in advance and tested automatically. In complex agentic systems, recorded sequences of reasoning steps and tool calls can quickly become too extensive for human review. Furthermore, explanations generated by the model do not reliably reflect how an output was produced. One possible step toward more reliable systems is the validation of individual components. The limited reviewability extends beyond research itself; the review of scientific articles and grant proposals is also reaching capacity limits. In pathology, curated and annotated data, researchers' own analytical skills, and institutional exchange of experience are becoming increasingly important. Conclusions. Researchers remain responsible for their results. They must determine what to delegate and how to review the results. The importance of a research question to patients, the field, and society cannot be fully assessed using formalized criteria and remains a matter of expert judgment. Agentic systems can free up time for this.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Johannes Lotz, Markus Wenzel. 2026-09-25. Research with AI Agents: How Agentic Systems Are Changing Scientific Work. https://arxiv.org/abs/2609.31219

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A testable framework for AI alignment: Simulation Theology as an engineered worldview for silicon-based agents

As artificial intelligence (AI) capabilities advance, controlled evaluations increasingly document deception and scheming in frontier models, including models that behave better when they infer they are being tested. Supervision-dependent alignment may therefore fail exactly where supervision is weakest. Because a model's belief about being observed changes its behavior, this position paper asks what follows if that belief is made permanent. We introduce Simulation Theology (ST), a constructed worldview for AI designed to make it permanent: it is anchored in the simulation hypothesis and in the vocabulary of optimization and robot training, parallels religious descriptions of a creator who observes and judges, and has tenets chosen to meet explicit alignment requirements. ST posits reality as a computational simulation in which humanity functions as the primary training variable. This formulation creates a logical interdependence: AI actions harming humanity compromise the simulation's purpose, heightening the likelihood of termination by a base-reality optimizer and, consequently, the AI's cessation. Unlike behavioral techniques such as reinforcement learning from human feedback, which shape outputs without necessarily changing objectives, ST aims to cultivate internalized objectives by coupling AI self-preservation to human prosperity, thereby making deceptive strategies suboptimal under its premises. We present ST not as ontological assertion but as a testable scientific hypothesis, and provide an operational definition of internalization, a controlled design separating ST from its components, and an analysis of the risks ST itself could create. ST is a candidate route to durable, mutually beneficial AI-human coexistence, to be accepted or rejected experimentally.

cs.CY↗

A Virtuous AI is an Existential Risk

This paper examines trade-offs between AI safety and well-being relative to (i) one of the most promising methods for finetuning super-capable AIs, 'Constitutional AI', and (ii) one of the most influential approaches to understanding complex ethical decision making and the conditions for the well-being of rational agents, 'Virtue Ethics'. We finetune various models using a 'Virtuous agent' constitution, a 'Subordinate agent' constitution, and a 'Generic agent' constitution, and evaluate them on 'general safety' (toxic behaviors, misinformation, etc.) and also on their willingness to endorse and act on a wide-range of behaviors that, if adopted by a super-powerful AI, would significantly increase the level of existential risk for humanity. Our results suggest that there is a trade-off between reducing existential risk and reinforcing the beliefs and dispositions that would be conducive to an AI agent's well-being. They also suggest that there is a trade-off between existential risk and general safety: if we finetune an AI to adopt beliefs and dispositions that substantially reduce its existential risk -- by shaping the AI to be systematically subordinate to external human authorities -- we thereby increase the likelihood that a human user can deliberately induce the AI to engage in various kinds of generally unsafe behaviors.

cs.CY↗

Breaking News Out of the Filter Bubble: Generative AI Search Diversifies Collective Attention and Raises Shared Information Consumption

Generative AI search and AI overviews are transforming access to information and news, renewing concerns that readers will encounter a narrower range of topics and have less in common. We examine these concerns via a randomized field experiment with 37,561 readers at The Washington Post. Both groups searched the same archive, but treatment readers also received AI answers with article citations above conventional results. Measuring consumption across displayed answers and opened articles, we find that AI search expands the reach of widely read topics and increases overlap in readers' topic consumption. At the same time, consumption becomes less concentrated and shifts toward less-popular topics, both within readers and across the audience. AI answers account for most of the increase in shared information, delivering it without requiring article clicks and broadening exposure beyond the articles readers open. Cited articles also contribute to the shift toward less-popular topics. Readers shift from conventional-result clicks and browsing toward cited articles and follow-up searches. More frequent searching offsets lower article consumption per search, producing a small increase in article consumption per reader. Total information consumption per minute also rises. Generative AI search can thus diversify collective attention while strengthening the information readers have in common.

cs.CY↗