Search arXiv⌕ Search

arXiv · 2610.11539

Design Creativity Bench: Measuring creativity in LLM-Generated UI

Abstract

As leading LLMs improve on capability evaluations, their limitations in producing creative outputs on design tasks remain insufficiently characterised. Our work introduces Design Creativity Bench, a benchmark that evaluates diversity and appropriateness in UI designs. It measures distinctiveness among models on the same prompt (originality), how much a model's designs change between two prompts for the same UI goal in different product domains (creative range), and the share of a brief's acceptance criteria each design meets (appropriateness). Originality is 0.592 for same-prompt design pairs from different models (95% CI [0.582, 0.602]), far below the 0.764 for same-prompt human-model pairs (95% CI [0.751, 0.778]). Creative range is 0.581 across models (95% CI [0.567, 0.597]), against 0.902 for human designs (95% CI [0.884, 0.919]). Appropriateness is above 90% for every model, and the best model reaches 99.2%, slightly above the 98.0% for human designs. Our work shows that the default output of LLMs, though generally appropriate, is substantially more repetitive than the human baseline. This calls for strong measures to address the issue.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Aman Rusia, Abhijit Bhole, Prashank Gupta, Dipanjan Dey. 2026-10-08. Design Creativity Bench: Measuring creativity in LLM-Generated UI. https://arxiv.org/abs/2610.11539

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Trends in EHR Satisfaction and Interoperability Among Family Physicians: A Five-Year Analysis of the ABFM Continuous Certification Questionnaire, 2022-2026

Objective: To assess trends in family physicians' (FPs) electronic health record (EHR) satisfaction and their experience of interoperability across organizations. Materials and Methods: Serial cross-sectional analysis of five waves of the American Board of Family Medicine's Continuous Certification Questionnaire (2022-26). Adjusted logistic regression compared odds of being very satisfied across 11 EHRs; Cochran-Armitage tests assessed vendor trends. Ease of using outside information was analyzed separately for same- and different-vendor sources. Results: Across 36,785 FPs, adjusted odds of being very satisfied relative to Epic ranged from 0.91 (95% CI, 0.74-1.11) for Elation Health to 0.17 (95% CI, 0.15-0.20) for Oracle Health/Cerner. Only Epic's users became significantly more satisfied (Benjamini-Hochberg adjusted P < .001), widening the vendor satisfaction gap. Within comparable periods, ease of using different-vendor information did not improve; fewer than 13% rated it very easy. Epic's interoperability advantage reversed by exchange type: non-Epic users had less than half the odds of Epic users of rating same-vendor interoperability very easy (adjusted OR, 0.41; 95% CI, 0.38-0.44) but 1.54 times the odds for different-vendor interoperability (95% CI, 1.35-1.75). Discussion: Rising national exchange volume has not yielded easier use of outside information. Epic's advantage is confined to its own network, so market concentration may shift, rather than solve, the different-vendor problem. Conclusion: The ABFM data are already informing EHR policy but also have utility for the EHR marketplace. While EHR satisfaction has known implications for burden and burnout, poor interoperability is a quality and safety threat that may be compounded by artificial intelligence.

cs.HC↗

Before Bringing It Up: When and How AI Companions Should Use Memory

Memory can sustain AI companionship, yet even accurate recollection can be inappropriate to use. Two rounds of formative interviews with 14 users (n = 6 exploratory, n = 8 memory-focused) motivate asking what a companion should consider before using past information. Eight themes inform Reconsider, a single-call procedure with five checks and four handling modes, evaluated on 80 scenarios across five models over 400 blinded within-model pairs. Two LLM judges favored Reconsider by net margins of +15 and +23 percentage points, with bootstrap intervals excluding zero for three of five models but not for GPT or Claude. Evaluator analysis linked judge scoring differences to model family, and a preliminary matched-guidance control isolating memory-specific content gave positive margins. We contribute an interview-grounded design framework for memory use and an evaluation that scrutinizes its own evaluators.

cs.HC↗

Intent Graph: Navigating the Analytical Reasoning Space for Exploratory Data Analysis

Exploratory data analysis (EDA) is rarely open-ended in practice: analysts work from high-level domain questions toward the concrete analyses that can answer them, prioritizing directions with domain knowledge and prior hypotheses. Large language models (LLMs) can supply such knowledge, but their responses are unstructured, leaving analysts no way to see what has been explored, what is missing, or why one direction was chosen over another. We present DAG-EDA, a system that lets analysts and an LLM co-navigate the space of possible analyses through two linked structures. An intent graph, governed by a grammar of analytical intent, decomposes an ambiguous natural-language question into progressively concrete analysis tasks, keeping alternative framings open and letting analysts branch, backtrack, and compare paths. A multi-layered knowledge graph externalizes the LLM's domain knowledge, linking domain concepts to the dataset variables that can measure them, so analysts can inspect and contest how their question is grounded in the data. Both graphs are constructed from only the dataset and the analyst's question, and the analyses the analyst reaches are rendered as interactive dashboards. We illustrate the system through a usage scenario and describe a user study design for examining whether the system scaffold analysts' reasoning and navigation.

cs.HC↗