Search arXivSearch

arXiv subjects

Anna Neumann

Publications and source records attributed to Anna Neumann.

6 recordsLinked to original sources

It is not enough to give your moderation rules to ChatGPT: Policy-as-Prompt Moderation and Its Potential Impacts on Community Governance

Content moderation practices and governance paradigms are changing rapidly, as fewer human moderators are deployed as `experts' by social media companies in a centralized manner. Instead, the companies are focusing more on community approaches, relying on volunteers to provide accurate information and make correct decisions. In decentralized moderation, communities have always relied on volunteers, updated community guidelines, and internal discussions thereof. For both content moderation paradigms, Artificial Intelligence (AI) seems like it could help ease moderation burdens of time, mental health, and accuracy. One possible way to operationalize AI in content moderation is a `policy-as-prompt'' approach, where the policy is formulated as a natural-language prompt and then passed to a large language model (LLM). This model then aids in moderation tasks. In this paper, we briefly lay out the technical and governance properties of this approach, and argue that its limitations lead to specific risks and harms that have to be addressed. Towards alleviating them, we lay out multiple considerations towards more effective prompt governance, but ultimately find that writing prompts alone is not appropriate for ensuring meaningful community governance.

cs.CY

Traces of Abuse: How Generative AI Impacts Image-Based Sexual Abuse (IBSA) Investigations

The introduction of generative AI (GAI) into the workflow of image-based sexual abuse (IBSA) only worsened the ease of creation and distribution, victimizing more people than ever. We outline how the introduction of generative AI (GAI-IBSA) impacts the creation of traces and the type of reasoning they allow. We illustrate the impact by comparing the forensic traces available in four different IBSA scenarios. We discuss the impacts on the (possibility of) investigation, arguing that the advent of generative AI overall benefits abusers by making perpetration easier, and the perpetrator harder to trace.

cs.CY

Prompt Governance? On Governing Technologies Governed by Natural Language

Generative artificial intelligence (GenAI) is increasingly operated by natural language instructions (prompts). Across the pipeline, stakeholders designate various forms, e.g. end-user guidelines, developer specifications, or system prompts, as prompt governance instruments. These textual artifacts are intended to shape model behaviour by specifying constraints, priorities, and compliance rules. Policymakers and regulators have begun to treat system-level instructions as accessible prompt-based GenAI intervention points, assuming they function (directly or indirectly) as behavioural control. Yet whether these instructions operate reliably and predictably enough across contexts to support such governance frameworks remains underexplored. Towards this, we systematically evaluate (i) how researchers discuss and treat system-level instructions in the literature, focusing on large language models (LLMs) as they isolate language effects; (ii) how policymakers position system-level instructions as governance objects, incorporating analysis of two policy frameworks (US Exec. Order on Preventing Woke AI, and EU General-Purpose AI Code of Practice); and (iii) whether misalignments between these perspectives warrant closer inspection of the viability of governing AI through natural language. We identify a fragmented literature advancing varying and contradictory claims about what goals system-level instructions can achieve, which we distil into a typology of claims. Further, we show how divergent claims complicate policy approaches that treat system-level instructions as stable, interpretable control mechanisms. We argue that given such misalignments, careful consideration must be given to prompt governance approaches. Our findings have broad implications, extending from a LLM policy context to the use of natural language as control mechanism in technical systems more generally.

cs.CY

AI Safety Evaluations Need To Consider Cascading Effects

AI systems comprise a range of interactions across the technical and organisational components of a range of actors. These components work together to provide the systems' functionality. This socio-technical assemblage is increasingly described as an algorithmic supply chain. Given their role in supporting a wide range of systems, foundation models (FMs) are increasingly a key part of many algorithmic supply chains. In practice, various technical and non-technical components work to mediate, adapt, and augment the behaviour of models, such as FMs, both in general, and for their use in specific application contexts. However, many AI safety evaluations tend to focus on capabilities of FMs themselves and/or assess these components independently, at certain levels of abstraction, with less consideration on how these components could interact, influence, reinforce or counteract each other. In this paper, we introduce the cascade as a concept for supporting more holistic AI evaluations. The cascade captures how the interactions between socio-technical components along the AI supply chain can compound and produce cumulative effects with downstream consequences. Specifically, we (i) identify gaps in current AI auditing approaches, using LLMs as our case study; (ii) demonstrate how cascade problems manifest in deployed AI systems through a characterisation of cascades through different viewpoints (component and stakeholder), and (iii) propose research directions for assessing the cascade specifically as an object of analysis. As these cascades can significantly impact transparency, accountability, security, and safety, we advocate for a paradigm shift in AI system auditing towards systems-oriented audits that incorporate cascading effects, to complement model centric evaluations.

cs.CY

Who Controls the Conversation? User Perspectives On Generative AI (LLM) System Prompts

System prompts - instructions that shape the behaviour of generative AI systems - strongly influence system outputs and users' experiences. They define the model's guidelines, `personality', and guardrails, taking precedence over user inputs. Despite their influence, transparency is limited: system prompts are generally not made public and most platforms instruct models to conceal them, leaving users disconnected from and unaware of a key mechanism guiding and governing their AI interactions. This paper argues that system prompts warrant explicit, user-centred design attention and, focusing on large language models (LLMs), asks: what do system prompts contain, how do end-users perceive them, and what do these perceptions offer for design and governance practice? Our results reveal user perspectives on: the benefits and risks of system prompts; the values they prefer to be associated with prompt-design; their levels of comfort with different types of prompts; and degrees of transparency and user control regarding prompt content. From these findings emerge considerations for how designers can better align system prompt mechanisms with user expectations and preferences over these mechanisms that directly shape how generative AI systems behave.

cs.CY

Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs)

System prompts in Large Language Models (LLMs) are predefined directives that guide model behaviour, taking precedence over user inputs in text processing and generation. LLM deployers increasingly use them to ensure consistent responses across contexts. While model providers set a foundation of system prompts, deployers and third-party developers can append additional prompts without visibility into others' additions, while this layered implementation remains entirely hidden from end-users. As system prompts become more complex, they can directly or indirectly introduce unaccounted for side effects. This lack of transparency raises fundamental questions about how the position of information in different directives shapes model outputs. As such, this work examines how the placement of information affects model behaviour. To this end, we compare how models process demographic information in system versus user prompts across six commercially available LLMs and 50 demographic groups. Our analysis reveals significant biases, manifesting in differences in user representation and decision-making scenarios. Since these variations stem from inaccessible and opaque system-level configurations, they risk representational, allocative and potential other biases and downstream harms beyond the user's ability to detect or correct. Our findings draw attention to these critical issues, which have the potential to perpetuate harms if left unexamined. Further, we argue that system prompt analysis must be incorporated into AI auditing processes, particularly as customisable system prompts become increasingly prevalent in commercial AI deployments.

cs.CY