Search arXiv⌕ Search

arXiv · 2610.04174

DimSteer: Steering LLM Authoring with Automatically Discovered Stylistic Controls

Abstract

Large language model writing interfaces often make users steer outputs by repeatedly articulating desired changes in natural language. Yet writers may recognize useful stylistic directions only after seeing alternatives, making revision recall-heavy. We present DimSteer, an authoring interface that samples prompt-local completions, discovers high-variance activation-space axes of variation, labels them, and exposes them as sliders with pole previews, diff comparison, and reset controls. Users can manipulate discovered dimensions, reducing the need to reformulate prompts for each stylistic adjustment. In a within-subjects study with 16 participants against a matched prompt-only baseline, DimSteer reduced mental demand, effort, and frustration while preserving comparable perceived success. Participants valued the surfaced dimensions, yet 15 of 16 disagreed that they would have thought to request the same changes in a prompt. Results suggest prompt-local controls can shift LLM authoring from recall-based prompting toward recognition-based exploration and direct manipulation, while preserving prompting for open-ended edits.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ajit Mallavarapu, Ziwei Gu. 2026-10-03. DimSteer: Steering LLM Authoring with Automatically Discovered Stylistic Controls. https://arxiv.org/abs/2610.04174

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

VIRENA: Virtual Arena for Research, Education, and Democratic Innovation

Digital platforms shape how people communicate, deliberate, and form opinions. Studying these dynamics has become harder because of restricted data access, ethical limits on real-world experiments, and the technical demands of existing research tools. VIRENA (Virtual Arena) is a platform for controlled experiments in realistic social media environments. Several participants can interact at the same time in replicas of feed-based platforms (Instagram, Facebook, Reddit, X) and messaging apps (WhatsApp, Messenger). AI agents powered by large language models (LLMs) join the participants with configurable personas and human-like timing. Researchers set up experiments in a visual interface without programming: they define conditions, schedule stimulus content, add moderation rules, assign participants randomly to conditions, and export the data. VIRENA supports designs that were hard to run before, such as studying human--AI interaction in realistic social settings, comparing moderation interventions, and observing group deliberation as it unfolds. Participants and data stay within institutional control, and the platform links to survey and recruitment tools. This paper describes how VIRENA works and how to use it.

cs.HC↗

Evaluating Just Noticeable Differences in Layered Opacity Visualizations

Opacity is a widely used channel in data visualization, but it remains less well understood compared to channels such as color, length, size, etc. Recent work from Meng et al. investigated the impact of opacity across competing color schemes, finding that certain color schemes were associated with better participant accuracy. We examine these effects further in a controlled two-alternative forced-choice setup to determine whether opacity differences are truly equal across possible opacity comparison ranges. In a within-subjects study with 96 trials, including two competing color schemes (best and worst from Meng et al.) and 48 opacity pairs, we find little differences between color schemes but larger individual differences in accuracy. Further, results show stable performance in middle opacity ranges, with more errors occurring when comparing extreme values. We discuss potential implications for design guidelines and further study and make our study materials, analysis scripts, and data available at https://osf.io/zv9dx/overview?view_only=38d03cea1b3d42788593e3e6b1016cfd.

cs.HC↗

Benchmarking Psychological Dynamics in Generative Agents

Large language models (LLMs) are increasingly deployed to simulate human behavior, acting as computational replicas of human subjects. Yet the lived psychological experience of humans is difficult to benchmark, particularly as it unfolds over time. We introduce a psychometric benchmark for computational replicas: personas that carry a fixed identity through an evolving sequence of events. Built entirely from published norms and meta-analytic effects, the benchmark scores two dimensions of psychological realism. The first, internal validity, quantifies whether generated trajectories reproduce the internal structure of repeated human measurement: distributions, the between- versus within-person variance partition, temporal dependence, and range. The second, external validity, quantifies whether replicas recover established trait, state, and indicator relations. Across 36 open-weight and proprietary LLMs from nine developers (1B-671B parameters), most recover the direction of established relations (84.3% mean agreement) and the variance partition (26 of 34), yet the within-person correlation and distributional structure elude recovery. Internal validity is independent of scale and capability: a mid-size open LLM (Gemma-3-27B) strikes the best trade-off between the two dimensions. The benchmark is a precondition for using computational replicas in causal inference across domains (e.g., marketing, healthcare), and identifies within-person grounding as the central challenge ahead.

cs.HC↗