Search arXivSearch

arXiv · 2607.02201

The Eticas AI Risk Taxonomy: Open Infrastructure for Operationalizing AI Audits

Abstract

The rapid deployment of AI systems across high-stakes domains has created urgent demand for standardized evaluation, yet the field remains fragmented across competing risk taxonomies that catalog risks without showing how an audit is executed. At least 74 AI risk taxonomies exist, and almost all stop at the catalog. The hard part of auditing is not naming a risk but operationalizing it: turning it into a test run against a real system, a measured value, a calibrated severity, and a defensible grade. This paper leads with that bridge. We present the operationalization layer Eticas has built and run, shown end to end on a single risk (PII leakage) against a public benchmark, and then the open taxonomy that makes the method scale. On GPT-4-0314, a disclosure risk that seven external frameworks require be controlled is measured at 0%, 51%, and 84% disclosure as adversarial conditioning increases, mapping through calibrated severity bands to a subcategory grade of E with a SYSTEMIC pattern. Around this example, the Eticas AI Risk Taxonomy v3.0.0 organizes 70 active subcategories across 10 categories and 21 sub-groups, with mappings to 18 external frameworks across compliance, reference, and academic tiers. Its established layer - categories, sub-groups, and the 32 established subcategories - is published under CC BY 4.0 as open semantic infrastructure with stable URIs and SKOS/JSON-LD distributions, and a worked subcategory example shows the operational layer down to its severity thresholds. The contribution is the demonstrated bridge from concept to graded finding, anchored by a clean separation of risks from the mechanisms by which they surface, and framed by an open-core model in which the conceptual scaffold is open and the methodology calibration is the practitioner layer. This is the infrastructure the AI auditing field needs: shared, open, and demonstrably operable.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Gemma Galdon Clavell, Pablo Accuosto, Usman Gohar. 2026-07-27. The Eticas AI Risk Taxonomy: Open Infrastructure for Operationalizing AI Audits. https://arxiv.org/abs/2607.02201

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A Brief AI Literacy Intervention Does Not Significantly Reduce Over-Reliance and Increases Under-Reliance on ChatGPT: A Randomized Study

In this study, we examined whether a brief AI literacy intervention influences high school students' reliance on recommendations from large language models (LLMs). In a randomized experiment, students were assigned to either a control group receiving a brief introduction to LLMs or an intervention group receiving additional information about how LLMs work, their limitations, and effective usage strategies. Participants then solved eight math puzzles with ChatGPT's advice, which was incorrect in half of the trials. Results indicated widespread over-reliance, with incorrect recommendations adopted in 52.1% of the trials. The intervention did not significantly reduce over-reliance. Instead, it led to an increase in under-reliance, as students were more likely to reject correct recommendations. These findings provide preliminary evidence that brief text-based interventions may be ineffective in fostering appropriate reliance. More comprehensive and interactive approaches may be required to meaningfully influence students' real-world reliance on LLMs.

cs.CY

Examining Community-Requested Fact-Checking: Request Alerts Are Associated with Greater Diversity and Visibility of Community Notes

Crowdsourced fact-checking systems such as Community Notes are increasingly used on social media platforms, yet concerns remain about which content receives scrutiny and how visible that scrutiny is. X allows users to request notes for specific posts. When sufficient requests accumulate, an alert is displayed, creating an interface cue that may guide contributor behavior. We present a quantitative, non-causal analysis comparing the diversity and visibility of community notes written for X posts with and without request alerts. We infer alert presence at note submission and analyze 10,432 alerted and 44,442 non-alerted English notes from 318 top writers. We find that, alerted notes are associated with greater individual-level topical diversity, but also with stronger collective concentration in the Politics category. Mixed-effects models estimate that alerted notes are 8.4-20.2 percentage points more likely to be modeled helpful and visible, though this visibility gain diminishes as topics diverge from writers' prior interests.

cs.CY

Scarcity and Predictive Uncertainty: Implications for Societal Resource Allocation

An emerging literature examines the critical question of when and how prediction can be useful in allocating scarce societal resources. We examine a novel variant of this question: What happens when predictive uncertainty differs systematically across the population? This can occur in several situations; for example, when machine learning models have significantly different accuracies across different demographics. We show that this uncertainty has serious implications for resource allocation when coupled with commonly used binary measures of societal benefit from allocation. We formulate a novel mathematical model of scarce resource allocation that accounts for heterogeneous predictive uncertainties and analyze implications for both the allocation mechanism and the realized population-level benefits. We find that when resources are very scarce, maximum marginal benefit (MMB) prioritization favors individuals with lower predictive uncertainty even at the identical underlying initial state. However, we observe a flip in prioritization when resources are abundant, targeting higher-uncertainty individuals. We illustrate the implications of our results on the PISA educational testing dataset. Our findings have meaningful ramifications for the distributional outcomes of prioritization policies in many domains touched by the theory of local justice, including the allocation of public education resources, medical triage, and homelessness services. They also reveal a new moral dilemma in the ethics of scarce resource allocation - is it just to allocate a resource to one person over another solely based on predictive uncertainty about their futures? We also assess efficiency losses under both MMB and the vulnerability-first (VF) prioritization. Our model predicts efficiency losses across all resource levels, but particularly in low-resource settings.

cs.CY