Search arXiv⌕ Search

arXiv · 2610.04418

CORE-RL: Confidence-Oriented Reliability Evaluation of Black-Box Reinforcement Learning Policies

Abstract

The deployment of Reinforcement Learning (RL) agents in critical domains must be preceded with a pipeline to evaluate the alignment of the RL agent with complex multi-objective specifications and robustness under real-world environmental drift. However, to protect intellectual property, the RL agent may be delivered for evaluation as opaque executable or remote API, which makes traditional evaluation techniques based on the internals of the policies infeasible. To address this gap, CORE-RL: Confidence-Oriented Reliability Evaluation of black box RL policy is proposed in this paper. The CORE-RL pipeline introduces a Unified Reliability Metric that formally integrates early task termination and safety constraint violations, preventing unsafe policies from masking failures through premature episode halts. By subjecting the policy to a noise certification envelope of perceptual noise, actuation noise and change in environment dynamics, the pipeline computes the finite-sample Clopper-Pearson bounds on unified reliability metric and Hoeffdings' lower bound on reward and safety cost. The pipeline then defines safe operational design domain to report high-confidence certificates for safety and expected performance. Experiments on continuous control tasks demonstrate the CORE-RL pipeline's ability to automatically reject non-compliant policies and map the safe Operational Design Domain (ODD) of safety-aware policies. Thus CORE-RL provides an evaluation framework towards a quantitative, transparent and reproducible, statistical rationale necessary to safely evaluate, compare, and deploy black box RL solutions.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Santhosh GS, Ananya Ravi, Devika Jay, Abhishek Sarkar, Perepu Satheesh Kumar, Saurav Prakash, Kaushik Dey, Balaraman Ravindran. 2026-10-03. CORE-RL: Confidence-Oriented Reliability Evaluation of Black-Box Reinforcement Learning Policies. https://arxiv.org/abs/2610.04418

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

PaTAS: A Framework for Trust Propagation in Neural Networks Using Subjective Logic

Trustworthiness has become a key requirement for deploying artificial intelligence in safety-critical applications, yet conventional metrics such as accuracy fail to capture uncertainty or the reliability of predictions, particularly under adversarial or degraded conditions. This paper introduces the Parallel Trust Assessment System (PaTAS), a framework for modeling and propagating trust in neural networks using Subjective Logic (SL). PaTAS operates in parallel with standard neural computation through Trust Nodes and Trust Functions that propagate input, parameter, and activation trust across the network, refining parameter trust during training (Parameter Trust Update) and returning a per-inference trust opinion, quantifying belief, disbelief, and uncertainty, on each prediction (Inference-Path Trust Assessment). Experiments on real-world and adversarial datasets show that these estimates are interpretable and stable and complement accuracy: PaTAS flags predictions that remain accurate yet rely on parameters learned from corrupted data, detects adversarially patched inputs without knowledge of the trigger, and audits training-label corruption without any test data. Compared with output-only uncertainty baselines and dedicated input- and feature-space detectors, out-of-distribution and corruption detection is carried by the conformity-derived input opinions PaTAS consumes, within a regime that a per-dataset diagnostic identifies in advance; trust propagation adds model-side signals these detectors lack, and the composed trust degrades gracefully where any single signal fails. The mechanism extends to convolutional networks. PaTAS thereby provides a foundation for transparent, quantifiable trust reasoning across the AI lifecycle.

cs.AI↗

High-Precision Estimation of the State-Space Complexity of Shogi via the Monte Carlo Method

Determining the state-space complexity of Shogi (Japanese Chess) has been a challenging problem, with previous combinatorial estimates leaving a gap of five orders of magnitude ($10^{64}$ to $10^{69}$). This gap arises from the difficulty of distinguishing positions reachable from the initial position among the vast number of board configurations. In this paper, we present a high-precision statistical estimate of the number of reachable Shogi positions. Here, a position consists of the side to move, the board configuration, and the pieces in hand, without move history, and positions related by exchanging the two players or by horizontal reflection of the board are counted once. Our method combines Monte Carlo sampling with a reachability test that performs reverse search toward the set of ``King-King only'' (KK) positions rather than toward the single initial position. Since every KK position and the initial position are mutually reachable, reaching any KK position proves reachability, and removing pieces from the board is a much simpler objective than reconstructing the initial position. Based on a sample of five billion positions, we estimate the number of such positions in Shogi to be $6.55 \times 10^{68}$; both ends of the approximate $3σ$ confidence interval round to this value at three significant digits. We also applied this method to Mini Shogi, obtaining approximately $2.38 \times 10^{18}$.

cs.AI↗

DERM-3R: A Resource-Efficient Multimodal Agents Framework for Dermatologic Diagnosis and Treatment in Real-World Clinical Settings

Skin diseases impose a substantial and growing global health burden. Modern dermatologic therapies control acute manifestations rapidly, but single-target treatment paradigms, recurrent disease courses, and neglected systemic comorbidities limit their long-term effectiveness. Traditional Chinese medicine (TCM) offers a complementary, holistic approach through syndrome differentiation and individualized treatment, yet its clinical practice is hindered by non-standardized knowledge, incomplete multimodal records, and the difficulty of scaling expert-driven reasoning. We propose DERM-3R, a resource-efficient multimodal agents framework that models TCM dermatologic diagnosis and treatment under limited data and computational resources. We decompose and redefine real-world dermatologic decision-making into three essential issues: fine-grained lesion recognition, multi-view lesion representation with specialist-level pathogenesis modeling, and holistic clinical reasoning for syndrome differentiation and treatment planning. DERM-3R comprises three collaborative agents, DERM-Rec, DERM-Rep, and DERM-Reason, each addressing one of these issues. Built on a lightweight multimodal large language model and fine-tuned by partial-parameter finetuning on 103 real-world TCM psoriasis cases, DERM-3R performs strongly across multiple dermatologic reasoning tasks. Automatic metrics, LLM-as-a-Judge assessments, and human doctors show that, despite extremely limited training data and parameter updates, DERM-3R matches or even surpasses hundred-billion-parameter general-purpose multimodal models such as GPT-5.1 and Gemini-3-Flash. Our results indicate that structured, domain-aware multi-agent modeling is an effective alternative to brute-force scaling for complex clinical tasks, offering a practical and scalable paradigm for multimodal AI in dermatology and integrative medicine.

cs.AI↗