Search arXiv⌕ Search

arXiv · 2610.04123

Agent Reliability Profiles in Financial Services

Abstract

AI agents can take actions. At times, those actions can go beyond what is intended. Agent reliability can be defined as assurance that an agent will stay within intended bounds and operate within limits. Today, there is no shared framework or language for describing, validating, and benchmarking the reliability of agentic deployments in financial services. This makes it difficult for financial institutions, vendors, and regulators to assess and trust agents at scale, thus limiting the pace of development and adoption. A standardized, shared representation of agent reliability would fill the gap. This paper introduces the Agent Reliability Profile, a per-agent unit of assurance evidence for agent deployments in financial services. Each Profile records a bounded, falsifiable claim, this agentic system reliably functions within its operating boundary. We define "operating boundary" as an agent having; (1) a defined autonomy tier, (2) a defined operational design domain, (3) defined classes of action, and (4) a defined control envelope. Production assurance progresses through three levels while the Profile schema remains constant: a Profile Builder compiles a Level 1 Asserted Profile from institutional evidence, a Profile Validator tests the deployment in its own environment to produce a Level 2 Validated Profile, and operation of the same tests by a qualified independent assessor produces a Level 3 Verified Profile. Separately a Benchmarked Profile reports results comparable across institutions under reference conditions. We describe the architecture, the artifact, the assurance ladder, the comparability flag, associated tools, an evaluation methodology, applications for financial institutions and supervisors, limitations, and a staged implementation program.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mike Hsu, Medha Bankhwal, Béatrice Moissinac, Kevin Werbach, Lukasz Szpruch, Bennett Hillenbrand. 2026-10-02. Agent Reliability Profiles in Financial Services. https://arxiv.org/abs/2610.04123

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Large Language Models and Augmented Democracy

Artificial intelligence enables computational agents to represent political preferences and take part in collective decision-making. In this thesis, I investigate the opportunities and challenges of digital twins (DTs) based on Large Language Models (LLMs) as intermediaries in augmented democracy, focusing on individual preference representation, collective representation of political organizations, and the vulnerability of those representations to attackers. First, using data from an online experiment in Brazil, I examine whether personalized DTs can predict citizens' preferences for unseen policy proposals. Second, I extend the DT framework from individuals to political organizations. Using Swiss parliamentary data, I build topic-specific knowledge graphs from lawmakers' legislative records and connect them to LLM-based lawmaker agents, which are organized into party-level DTs representing collective positions. Agentic deliberation among these agents tests whether aggregated party representations capture a broader range of intra-party perspectives than official party communications. Finally, I study the vulnerability and robustness of LLM-mediated deliberation against prompt-injection attacks that amplify viewpoints, suppress opinions, or redirect consensus. Using data from a 2023 deliberative experiment in the United Kingdom, I analyze how attack effectiveness varies with the distribution of opinions and rhetorical strategies, and evaluate a pipeline combining injection detection, structured opinion representations, and reinforcement learning to improve resistance. These findings characterize the opportunities and challenges of LLM-based digital twins in augmented democracy, stressing accurate preference representation, faithful aggregation, and robustness to strategic interaction.

cs.CY↗

Enhancing Knowledge Tracing through Leakage-Free and Recency-Aware Embeddings

Knowledge Tracing (KT) aims to predict a student's future performance based on their sequence of interactions with learning content. Many KT models rely on knowledge concepts (KCs), which represent the skills required for each item. However, some of these models are vulnerable to label leakage, a phenomenon in which the input data inadvertently reveal the correct answer, particularly in datasets with multiple KCs per question. We propose a straightforward yet effective solution to prevent label leakage by masking ground-truth labels during input embedding construction whenever such leakage could occur. To accomplish this, we introduce a dedicated MASK label, inspired by masked language modeling (e.g., BERT), to replace ground-truth labels. In addition, we introduce Recency Encoding, which encodes the step-wise distance between the current item and its most recent previous occurrence. This distance is important for modeling learning dynamics such as forgetting, which is a fundamental aspect of human learning, yet it is often overlooked in existing models. Recency Encoding demonstrates improved performance over traditional positional encodings on multiple KT benchmarks. We show that incorporating our embeddings into KT models such as DKT, DKT+, AKT, and SAKT consistently improves prediction accuracy across multiple benchmarks. The approach is both efficient and widely applicable.

cs.CY↗

"Lighting The Way For Those Not Here": How Can Technology Researchers Help Resist the Missing and Murdered Indigenous Relatives (MMIR) Crisis?

Indigenous peoples across Turtle Island face disproportionate rates of disappearance and murder, a genocide rooted in settler-colonial violence and systemic erasure. Technology plays a crucial role in the Missing and Murdered Indigenous Relatives (MMIR) crisis: it perpetuates systemic violence and impedes investigations, yet also enables sites of advocacy, healing, and resistance. For example, Native communities utilize AMBER alerts, digital news, sovereign crowdsourced databases, social media groups, and resistance movements to mobilize searches, amplify awareness, and honor missing relatives. Yet little research in HCI has critically examined the role of technology in shaping the MMIR crisis. Thus, we qualitatively analyze 140 webpages to identify sociotechnical barriers that hinder communities' efforts, while highlighting actions that foster healing, safety, and resilience. We grounded our analysis in stories that resist epistemic erasure through relational accountability, critical humility, cultural sensitivity, and refusal. Finally, we provide recommendations for HCI to recognize self-determination and sovereignty of Indigenous technologies, direct action to support families, and honor Indigenous onto-epistemologies that cease epistemic violence.

cs.CY↗