Search arXivSearch

arXiv subjects

Michael Correll

Publications and source records attributed to Michael Correll.

At least 19 recordsLinked to original sources

Visualizing Uncertainty in Non-linear Projections with Ensembles

Widely used non-linear dimensionality reduction (NLDR) methods such as UMAP and t-SNE are stochastic--repeated runs on the same data can produce different low-dimensional projections. In this paper, we explore two problems related to projection variability: on some datasets clusters, structure, and outliers may change run-to-run, and on others projections can be extremely stable when overfitting noise. To address the first problem, we propose visualizing the median of multiple NLDR outputs rather than relying on individual projections. To address the second, we perturb input data before creating consensus embeddings. We find that taking the median of multiple projections performs comparably to individual runs on multiple quality metrics, while increasing perturbation emphasizes global over local structure. We show through a set of exploratory visualizations that even relatively simple ensemble presentations can be used to better communicate the reliability of projection patterns.

cs.HC

Seeing Through the Forecast Clutter: Communicating Climate Forecast Distributions with Weighted Multiple Forecast Visualizations

Forecasts often diverge because different models make varying assumptions to account for underlying uncertainty. Readers who consume forecasts may wish to survey the shape and spread of these multiple forecasts to get a full account of the different predictions. One approach to visualizing multiple forecasts is through Confidence Interval (CI) plots. However, while the summative CI plots can communicate uncertainty of an ensemble, they obscure attributes of individual forecasts that can lead to inaccurate perceptions of the distribution of these forecasts (e.g., implying a normal distribution when non-existent). To address this challenge, we investigate the use of multiple forecast visualization (MFV) in communicating nuanced forecast distributions through two preregistered experiments using climate forecast data. In Experiment 1 (480 participants), we compared how well MFV and CI plots can represent the distribution of multiple forecasts. We found that, compared to CI plots, MFV improved participants' ability to identify the underlying distribution of forecasts and reduced the likelihood of assuming normality. Building on Experiment 1, we examined in Experiment 2 (900 participants) whether a downsampled MFV showing 9 forecasts might be able to communicate additional forecast properties using linewidth and opacity without negatively impacting distribution perception. We found that visually weighting forecasts by linewidth or opacity preserves readers' perception of the underlying distribution. We discuss how these findings suggest the use of downsampled and weighted MFV to cut through forecast clutter by aligning perceived distribution with the underlying forecast distribution, while opening up design opportunities to use weighting to communicate additional forecast attributes.

cs.HC

Charting the Moral Universe: Capturing Virtues and Values of Data Visualization Practice

What do we value in our visualizations, and in the people who design them? Despite a growing body of work on critical data visualization, the conception of what it is to do ethical data visualization work can often be narrow (for instance, holding that our ethical duties are discharged merely by avoiding overtly lying or manipulating data), or entangled with potentially problematic implicit value structures (such as the assumption of the objectivity and neutrality of data, and so the designer's role being merely the passive conveying of numbers as efficiently as possible). Yet, what it means to act ethically in data visualization is broad and multifaceted, and the virtues to which we should aspire as data visualization researchers and designers are worth explicating. We conducted an interview study with a broad spectrum of 20 experienced data visualization researchers, practitioners, and data artists to solicit their values and ethical considerations around doing visualization work. We report on a list of 68 values, organized into nine virtue clusters, that we encountered in our interviews. These virtues and values together describe a diverse space of matters of care and concern in data visualization: from the unease around the best use of visualization as a tool for persuasion, to the tightrope that visualization practitioners often walk between their professional responsibilities and their personal moral commitments. The virtues themselves, as well as our interviewees' reflections on ethical practice, offer practitioners, researchers, and educators in data visualization a richer vocabulary for ethical reflection and provide a broader foundation for considering and applying visualization ethics.

cs.HC

"Trust Junk" Leads to Unjustified Support for Highly Discriminatory Predictive Models

The persuasive power of data visualizations can go awry: for instance, in an explainable AI (XAI) context, visualizations can produce over-trust of predictive models. In this paper, we use a crowdsourced study to show that providing accurate (but superfluous or irrelevant) data in a model explanation can, in fact, result in unjustified trust and other positive beliefs about a model, even when the model is patently discriminatory and unfair. Our results suggest that XAI designers and developers need to consider the implicit or explicit rhetorics of their work, and beware of the potential of visualizations to imbue models with unearned trust.

cs.HC

Should We Dangle a Carrot? The Effect of Performance-based Incentives in Visualization Experiments

A perennial research question in visualization involves identifying which visual encodings for a particular dataset are most effective for users in performing a specific task. The relative effectiveness of the different encodings are commonly identified through controlled experiments. However, designing an experiment involves making many, often ad hoc, decisions about the experimental setup such as whether to include a training module, whether to provide performance-based incentives to participants, etc. Yet, there is limited guidance on how these decisions should be made, and we do not fully understand the impact of these subjective decisions on empirical results. In this paper, we investigate the impact of one such key design decision: monetary rewards. Specifically, we ask: does providing or not providing participants with performance-based financial incentives affect the results and the conclusions that we draw from visualization studies? We conducted two crowdsourced studies investigating the impact of incentives on (i) a low-level, perceptual task (perception of correlations in scatterplots or parallel coordinate plots), and (ii) a task involving reasoning (decision-making based on a weather forecast represented as intervals or density plots). In each of these studies, we manipulate both the visual representation and the presence of incentives as between-subject conditions. We expected to find no effect of incentives on the perceptual task, but to see an effect for the decision-making task. However, we found no effect on task performance in either study. While these are results of only two studies and should be replicated, they suggest that performance-based financial incentives may not always have the intended effect on participants that we presumed, and calls for a reflection of how incentivized studies should be designed.

cs.HC

"The Worst Weather In America": Augmenting the Information Design of Extreme Cold Weather Forecasts

Mount Washington is home to extreme, and extremely volatile, weather conditions. Consulting a weather forecast of conditions at the summit is vital for making one's visit as safe as possible. Using the discussion and suggestions arising from a participatory workshop as input, we test a design intervention employing color-coded hazard icons to function as visual summaries of Mount Washington Observatory's current text-heavy forecast through a crowd-sourced study. We find that the use of icons increases the perceived risk of activities involving visiting the mountain. However, we highlight remaining questions around visualization design and design ethics that warrant further study in the domain of how best to communicate cold weather hazards in ways that are mindful of the diversity of literacies and experiences of visitors.

cs.HC

Visual Stenography: Feature Recreation and Preservation in Sketches of Noisy Line Charts

Line charts surface many features in time series data, from trends to periodicity to peaks and valleys. However, not every potentially important feature in the data may correspond to a visual feature which readers can detect or prioritize. In this study, we conducted a visual stenography task, where participants re-drew line charts to solicit information about the visual features they believed to be important. We systematically varied noise levels (SNR ~5-30 dB) across line charts to observe how visual clutter influences which features people prioritize in their sketches. We identified three key strategies that correlated with the noise present in the stimuli: the Replicator attempted to retain all major features of the line chart including noise; the Trend Keeper prioritized trends disregarding periodicity and peaks; and the De-noiser filtered out noise while preserving other features. Further, we found that participants tended to faithfully retain trends and peaks and valleys when these features were present, while periodicity and noise were represented in more qualitative or gestural ways: semantically rather than accurately. These results suggest a need to consider more flexible and human-centric ways of presenting, summarizing, pre-processing, or clustering time series data.

cs.HC

A Priest, a Rabbi, and an Atheist Walk Into an Error Bar: Religious Meditations on Uncertainty Visualization

In this provocation, we suggest that much (although not all) current uncertainty visualization simplifies the myriad forms of uncertainty into error bars around an estimate. This apparent simplification into error bars comes only as a result of a vast metaphysics around uncertainty and probability underlying modern statistics. We use examples from religion to present alternative views of uncertainty (metaphysical or otherwise) with the goal of enriching our conception of what kind of uncertainties we ought to visualize, and what kinds of people we might be visualizing those uncertainties for.

cs.HC

The Many Tendrils of the Octopus Map

Conspiratorial thinking can connect many distinct or distant ills to a central cause. This belief has visual form in the octopus map: a map where a central force (for instance a nation, an ideology, or an ethnicity) is depicted as a literal or figurative octopus, with extending tendrils. In this paper, we explore how octopus maps function as visual arguments through an analysis of historical examples as well as a through a crowd-sourced study on how the underlying data and the use of visual metaphors contribute to specific negative or conspiratorial interpretations. We find that many features of the data or visual style can lead to "octopus-like" thinking in visualizations, even without the use of an explicit octopus motif. We conclude with a call for a deeper analysis of visual rhetoric, and an acknowledgment of the potential for the design of data visualizations to contribute to harmful or conspiratorial thinking.

cs.HC

Complexity as Design Material

Complexity is often seen as a inherent negative in information design, with the job of the designer being to reduce or eliminate complexity, and with principles like Tufte's "data-ink ratio" or "chartjunk" to operationalize minimalism and simplicity in visualizations. However, in this position paper, we call for a more expansive view of complexity as a design material, like color or texture or shape: an element of information design that can be used in many ways, many of which are beneficial to the goals of using data to understand the world around us. We describe complexity as a phenomenon that occurs not just in visual design but in every aspect of the sensemaking process, from data collection to interpretation. For each of these stages, we present examples of ways that these various forms of complexity can be used (or abused) in visualization design. We ultimately call on the visualization community to build a more nuanced view of complexity, to look for places to usefully integrate complexity in multiple stages of the design process, and, even when the goal is to reduce complexity, to look for the non-visual forms of complexity that may have otherwise been overlooked.

cs.HC

The Data-Wink Ratio: Emoji Encoder for Generating Semantically-Resonant Unit Charts

Communicating data insights in an accessible and engaging manner to a broader audience remains a significant challenge. To address this problem, we introduce the Emoji Encoder, a tool that generates a set of emoji recommendations for the field and category names appearing in a tabular dataset. The selected set of emoji encodings can be used to generate configurable unit charts that combine plain text and emojis as word-scale graphics. These charts can serve to contrast values across multiple quantitative fields for each row in the data or to communicate trends over time. Any resulting chart is simply a block of text characters, meaning that it can be directly copied into a text message or posted on a communication platform such as Slack or Teams. This work represents a step toward our larger goal of developing novel, fun, and succinct data storytelling experiences that engage those who do not identify as data analysts. Emoji-based unit charts can offer contextual cues related to the data at the center of a conversation on platforms where emoji-rich communication is typical.

cs.HC

Data Guards: Challenges and Solutions for Fostering Trust in Data

From dirty data to intentional deception, there are many threats to the validity of data-driven decisions. Making use of data, especially new or unfamiliar data, therefore requires a degree of trust or verification. How is this trust established? In this paper, we present the results of a series of interviews with both producers and consumers of data artifacts (outputs of data ecosystems like spreadsheets, charts, and dashboards) aimed at understanding strategies and obstacles to building trust in data. We find a recurring need, but lack of existing standards, for data validation and verification, especially among data consumers. We therefore propose a set of data guards: methods and tools for fostering trust in data artifacts.

cs.HC

When the Body Became Data: Historical Data Cultures and Anatomical Illustration

With changing attitudes around knowledge, medicine, art, and technology, the human body has become a source of information and, ultimately, shareable and analyzable data. Centuries of illustrations and visualizations of the body occur within particular historical, social, and political contexts. These contexts are enmeshed in different so-called data cultures: ways that data, knowledge, and information are conceptualized and collected, structured and shared. In this work, we explore how information about the body was collected as well as the circulation, impact, and persuasive force of the resulting images. We show how mindfulness of data cultural influences remain crucial for today's designers, researchers, and consumers of visualizations. We conclude with a call for the field to reflect on how visualizations are not timeless and contextless mirrors on objective data, but as much a product of our time and place as the visualizations of the past.

cs.HC

Heuristics for Supporting Cooperative Dashboard Design

Dashboards are no longer mere static displays of metrics; through functionality such as interaction and storytelling, they have evolved to support analytic and communicative goals like monitoring and reporting. Existing dashboard design guidelines, however, are often unable to account for this expanded scope as they largely focus on best practices for visual design. In contrast, we frame dashboard design as facilitating an analytical conversation: a cooperative, interactive experience where a user may interact with, reason about, or freely query the underlying data. By drawing on established principles of conversational flow and communication, we define the concept of a cooperative dashboard as one that enables a fruitful and productive analytical conversation, and derive a set of 39 dashboard design heuristics to support effective analytical conversations. To assess the utility of this framing, we asked 52 computer science and engineering graduate students to apply our heuristics to critique and design dashboards as part of an ungraded, opt-in homework assignment. Feedback from participants demonstrates that our heuristics surface new reasons dashboards may fail, and encourage a more fluid, supportive, and responsive style of dashboard design. Our approach suggests several compelling directions for future work, including dashboard authoring tools that better anticipate conversational turn-taking, repair, and refinement and extending cooperative principles to other analytical workflows.

cs.HC

Toward a Scalable Census of Dashboard Designs in the Wild: A Case Study with Tableau Public

Dashboards remain ubiquitous artifacts for presenting or reasoning with data across different domains. Yet, there has been little work that provides a quantifiable, systematic, and descriptive overview of dashboard designs at scale. We propose a schematic representation of dashboard designs as node-link graphs to better understand their spatial and interactive structures. We apply our approach to a dataset of 25,620 dashboards curated from Tableau Public to provide a descriptive overview of the core building blocks of dashboards in the wild and derive common dashboard design patterns. To guide future research, we make our dashboard corpus publicly available and discuss its application toward the development of dashboard design tools.

cs.HC

Teru Teru B\=ozu: Defensive Raincloud Plots

Univariate visualizations like histograms, rug plots, or box plots provide concise visual summaries of distributions. However, each individual visualization may fail to robustly distinguish important features of a distribution, or provide sufficient information for all of the relevant tasks involved in summarizing univariate data. One solution is to juxtapose or superimpose multiple univariate visualizations in the same chart, as in Allen et al.'s "raincloud plots." In this paper I examine the design space of raincloud plots, and, through a series of simulation studies, explore designs where the component visualizations mutually "defend" against situations where important distribution features are missed or trivial features are given undue prominence. I suggest a class of "defensive" raincloud plot designs that provide good mutual coverage for surfacing distributional features of interest.

cs.HC

Fitting Bell Curves to Data Distributions using Visualization

Idealized probability distributions, such as normal or other curves, lie at the root of confirmatory statistical tests. But how well do people understand these idealized curves? In practical terms, does the human visual system allow us to match sample data distributions with hypothesized population distributions from which those samples might have been drawn? And how do different visualization techniques impact this capability? This paper shares the results of a crowdsourced experiment that tested the ability of respondents to fit normal curves to four different data distribution visualizations: bar histograms, dotplot histograms, strip plots, and boxplots. We find that the crowd can estimate the center (mean) of a distribution with some success and little bias. We also find that people generally overestimate the standard deviation, which we dub the "umbrella effect" because people tend to want to cover the whole distribution using the curve, as if sheltering it from the heavens above, and that strip plots yield the best accuracy.

cs.HC

Are We Making Progress In Visualization Research?

In this work, I use a survey of senior visualization researchers and thinkers to ideate about the notion of progress in visualization research: how are we growing as a field, what are we building towards, and are our existing methods sufficient to get us there? My respondents discussed several potential challenges for visualization research in terms of knowledge formation: a lack of rigor in the methods used, a lack of applicability to actual communities of practice, and a lack of theoretical structures that incorporate everything that happens to people and to data both before and after the few seconds when a viewer looks at a value in a chart. Orienting the field around progress (if such a thing is even desirable, which is another point of contention) I believe will require drastic re-conceptions of what the field is, what it values, and how it is taught.

cs.HC