Search arXiv⌕ Search

arXiv · 2610.02622

CuBEs: Culturally-Situated Behavioral Evaluations and the Limitations of Culture-Blind LLM Judges

Abstract

Evaluating the occurrence and triggers of large language model (LLM) behaviors - such as sycophancy, self-preference, or over-confidence - is critical for predicting real-world model deployment risks. However, existing situated behavioral evaluations typically ignore cultural context, limiting their generalizability across an increasingly global user base. To address this gap, we propose CuBEs - Culturally-situated Behavior Evaluations that probe for response patterns across diverse user cultures. We first extend an automated testing pipeline to inject cultural context into behavioral test scenarios and subsequent evaluation. We assess the cultural adaptability of this pipeline by building a human-labeled dataset that captures nuanced dimensions of behavior understanding across 12 distinct cultures. Our dataset reveals significant cross-cultural variations that one-size-fits all judgments fail to capture. Through evaluating 13 open- and closed-source LLMs, we find that introducing cultural situatedness in the evaluation scenario creates significant variation in the presence of a behavior. For example, while our baseline experiments testing for political bias capture localized Western political dimensions like the American conservative-progressive divide, non-Western culturally situated evaluations surface entirely different axes of bias such as religious and colonial political issues. Our findings demonstrate that standard, culturally-agnostic evaluations fail to capture these shifts, highlighting the necessity of culturally situated behavioral testing for global deployments.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Hoda Ayad, Tanu Mitra, Abhishek Mukherji. 2026-10-02. CuBEs: Culturally-Situated Behavioral Evaluations and the Limitations of Culture-Blind LLM Judges. https://arxiv.org/abs/2610.02622

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

PaTAS: A Framework for Trust Propagation in Neural Networks Using Subjective Logic

Trustworthiness has become a key requirement for deploying artificial intelligence in safety-critical applications, yet conventional metrics such as accuracy fail to capture uncertainty or the reliability of predictions, particularly under adversarial or degraded conditions. This paper introduces the Parallel Trust Assessment System (PaTAS), a framework for modeling and propagating trust in neural networks using Subjective Logic (SL). PaTAS operates in parallel with standard neural computation through Trust Nodes and Trust Functions that propagate input, parameter, and activation trust across the network, refining parameter trust during training (Parameter Trust Update) and returning a per-inference trust opinion, quantifying belief, disbelief, and uncertainty, on each prediction (Inference-Path Trust Assessment). Experiments on real-world and adversarial datasets show that these estimates are interpretable and stable and complement accuracy: PaTAS flags predictions that remain accurate yet rely on parameters learned from corrupted data, detects adversarially patched inputs without knowledge of the trigger, and audits training-label corruption without any test data. Compared with output-only uncertainty baselines and dedicated input- and feature-space detectors, out-of-distribution and corruption detection is carried by the conformity-derived input opinions PaTAS consumes, within a regime that a per-dataset diagnostic identifies in advance; trust propagation adds model-side signals these detectors lack, and the composed trust degrades gracefully where any single signal fails. The mechanism extends to convolutional networks. PaTAS thereby provides a foundation for transparent, quantifiable trust reasoning across the AI lifecycle.

cs.AI↗

High-Precision Estimation of the State-Space Complexity of Shogi via the Monte Carlo Method

Determining the state-space complexity of Shogi (Japanese Chess) has been a challenging problem, with previous combinatorial estimates leaving a gap of five orders of magnitude ($10^{64}$ to $10^{69}$). This gap arises from the difficulty of distinguishing positions reachable from the initial position among the vast number of board configurations. In this paper, we present a high-precision statistical estimate of the number of reachable Shogi positions. Here, a position consists of the side to move, the board configuration, and the pieces in hand, without move history, and positions related by exchanging the two players or by horizontal reflection of the board are counted once. Our method combines Monte Carlo sampling with a reachability test that performs reverse search toward the set of ``King-King only'' (KK) positions rather than toward the single initial position. Since every KK position and the initial position are mutually reachable, reaching any KK position proves reachability, and removing pieces from the board is a much simpler objective than reconstructing the initial position. Based on a sample of five billion positions, we estimate the number of such positions in Shogi to be $6.55 \times 10^{68}$; both ends of the approximate $3σ$ confidence interval round to this value at three significant digits. We also applied this method to Mini Shogi, obtaining approximately $2.38 \times 10^{18}$.

cs.AI↗

DERM-3R: A Resource-Efficient Multimodal Agents Framework for Dermatologic Diagnosis and Treatment in Real-World Clinical Settings

Skin diseases impose a substantial and growing global health burden. Modern dermatologic therapies control acute manifestations rapidly, but single-target treatment paradigms, recurrent disease courses, and neglected systemic comorbidities limit their long-term effectiveness. Traditional Chinese medicine (TCM) offers a complementary, holistic approach through syndrome differentiation and individualized treatment, yet its clinical practice is hindered by non-standardized knowledge, incomplete multimodal records, and the difficulty of scaling expert-driven reasoning. We propose DERM-3R, a resource-efficient multimodal agents framework that models TCM dermatologic diagnosis and treatment under limited data and computational resources. We decompose and redefine real-world dermatologic decision-making into three essential issues: fine-grained lesion recognition, multi-view lesion representation with specialist-level pathogenesis modeling, and holistic clinical reasoning for syndrome differentiation and treatment planning. DERM-3R comprises three collaborative agents, DERM-Rec, DERM-Rep, and DERM-Reason, each addressing one of these issues. Built on a lightweight multimodal large language model and fine-tuned by partial-parameter finetuning on 103 real-world TCM psoriasis cases, DERM-3R performs strongly across multiple dermatologic reasoning tasks. Automatic metrics, LLM-as-a-Judge assessments, and human doctors show that, despite extremely limited training data and parameter updates, DERM-3R matches or even surpasses hundred-billion-parameter general-purpose multimodal models such as GPT-5.1 and Gemini-3-Flash. Our results indicate that structured, domain-aware multi-agent modeling is an effective alternative to brute-force scaling for complex clinical tasks, offering a practical and scalable paradigm for multimodal AI in dermatology and integrative medicine.

cs.AI↗