Search arXiv⌕ Search

arXiv · 2610.06744

ufakzeka-karar: An Open Turkish Typed-Decision Model with Order-Invariant Option Scoring

Abstract

ufakzeka-karar is an open Turkish decision model with 182,494,466 parameters. Given a Turkish text and questions of a fixed answer type (a choice, a level on an ordered scale, or yes or no), it returns a temperature-scaled probability for every option and an expected error that serves as a "not sure" signal, without generating text and in one CPU forward pass for up to ten options. Built on the lab's ufakzeka-1-base, its head scores each option blind to the others at shared positions, so the answer does not depend on option order. A sequential head trained with shuffled options was about as accurate but changed 2.3 to 2.8 percent of its answers when only the option order changed; REINFORCE lost 10.2 points (0.102) of macro F1 to cross-entropy. On the open set of HakemBench v1.0 (4,275 questions, 7 tracks) the released model ranks 7th of 16 rows with a composite of 0.660 (95% interval 0.642 to 0.677). Temperature scaling lowers calibration error (smooth ECE) on the development set but raises it on held-out support questions, from 0.027 to 0.045 for the first scored run, which never trained on them; the released model later trained on them, so its 0.036 to 0.064 is not an unseen-question test. The released model is the last of three runs scored on HakemBench, and its numbers are not blind. The second run's new training data was aimed at the first run's errors on the full test set in guardrails, moderation and customer support, and the released run was trained after the second run's guardrail results on the full test set were read, under a protocol fixed in writing before any of its data, code or runs. All its numbers come after these readings; its guardrail, moderation and customer support numbers carry the flag "shaped by reading the test results". With every model scored on the other four tracks only, its composite is 0.678, 6th of 16. Weights and code are under Apache-2.0.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Sait Furkan Teke. 2026-10-05. ufakzeka-karar: An Open Turkish Typed-Decision Model with Order-Invariant Option Scoring. https://arxiv.org/abs/2610.06744

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Boosting Large Language Models with Mask Fine-Tuning

The large language model (LLM) is typically integrated into the mainstream optimization protocol. However, it remains underexplored whether maintaining the model integrity is \textit{indispensable} for promising performance. In this work, we introduce Mask Fine-Tuning (MFT), a novel LLM fine-tuning paradigm demonstrating that carefully breaking the model's structural integrity can surprisingly improve performance without updating model weights. MFT learns and applies binary masks to well-optimized models, using the standard LLM fine-tuning objective as supervision. Based on fully fine-tuned models, MFT uses the same fine-tuning datasets to achieve consistent performance gains across domains and backbones (e.g., an average gain of 2.70/4.15 on IFEval with LLaMA2-7B/3.1-8B). Detailed ablation studies and analyses examine the proposed MFT from different perspectives, including the sparse ratio and the loss surface. Additionally, when deployed on well-trained models, MFT is compatible with other LLM optimization procedures to improve overall model performance. Furthermore, this study extends the masking operation beyond its conventional use in network pruning for model compression to encompass a broader range of model capabilities.

cs.CL↗

Too Categorical to be Human: Emotion Concepts in LLMs and Humans

Understanding human emotions is central to user-facing AI applications, safety alignment, and the simulation of human behavior. As emotional stimuli shape high-stakes behavior in Large Language Models (LLMs), there is increasing interest in how models represent emotion concepts internally. Mechanistic accounts of these representations, however, cannot be compared directly against humans: emotion processing in humans is highly distributed and yields no equivalent neural representation. To understand whether LLMs internalize emotion concepts in a way similar to humans, we propose characterizing the abstract concept of an emotion using external behavioral signatures, which we term behavioral representations. Using the theory of cognitive appraisals, which enables representing emotional situations along interpretable evaluative dimensions, we create a benchmark dataset of emotional scenarios spanning 15 emotion categories. We elicit behavioral representations of emotion concepts from LLMs and humans using our benchmark, and study their structural similarity. We find that LLMs represent emotion concepts more categorically, homogeneously, and determinately than humans, representing a single emotion concept with less internal diversity, and place different emotions further apart. The categorical structure of representations in LLMs is further robust to contextual variation, including with different task framing and demographic personas. Analyzing model checkpoints across different training stages, we also find that the discretized nature of representations appears after the mid-training stage itself and is unaffected by different post-training strategies. Through our results, we highlight a key difference in how LLMs behaviorally represent emotion concepts, curbing the subjectivity inherent to the human experience of emotions.

cs.CL↗

FedCoT: Communication-Efficient Federated Reasoning Enhancement for Large Language Models

Enhancing LLM reasoning in federated settings is nontrivial due to stringent computational, communication, and privacy constraints, especially in healthcare, where clinically consequential decisions require not only accuracy but also interpretable, auditable rationales to meet safety, accountability, and regulatory requirements. Conventional federated fine-tuning largely imitates final answers rather than cultivating step-by-step reasoning, often relying on privacy-sensitive centralized distillation and still incurring substantial communication overhead. We address this gap with \textbf{\ours{}}, a federated reasoning framework that combines lightweight chain-of-thought resampling with a compact discriminator for selection, and client-aware LoRA stacking with weighted classifier aggregation to accommodate heterogeneity while reducing aggregation noise and communication; clients generate candidate chains and supervision locally, and only lightweight modules are aggregated on the server. Experiments on medical reasoning benchmarks show consistent gains under tight resource budgets while keeping data local and respecting privacy, offering an interpretable and resource-efficient solution. Our code is made publicly available at https://github.com/DIaacKr/FedCoT

cs.CL↗