Search arXivSearch

arXiv · 1511.04134

Whom Should We Sense in "Social Sensing" -- Analyzing Which Users Work Best for Social Media Now-Casting

Abstract

Given the ever increasing amount of publicly available social media data, there is growing interest in using online data to study and quantify phenomena in the offline "real" world. As social media data can be obtained in near real-time and at low cost, it is often used for "now-casting" indices such as levels of flu activity or unemployment. The term "social sensing" is often used in this context to describe the idea that users act as "sensors", publicly reporting their health status or job losses. Sensor activity during a time period is then typically aggregated in a "one tweet, one vote" fashion by simply counting. At the same time, researchers readily admit that social media users are not a perfect representation of the actual population. Additionally, users differ in the amount of details of their personal lives that they reveal. Intuitively, it should be possible to improve now-casting by assigning different weights to different user groups. In this paper, we ask "How does social sensing actually work?" or, more precisely, "Whom should we sense--and whom not--for optimal results?". We investigate how different sampling strategies affect the performance of now-casting of two common offline indices: flu activity and unemployment rate. We show that now-casting can be improved by 1) applying user filtering techniques and 2) selecting users with complete profiles. We also find that, using the right type of user groups, now-casting performance does not degrade, even when drastically reducing the size of the dataset. More fundamentally, we describe which type of users contribute most to the accuracy by asking if "babblers are better". We conclude the paper by providing guidance on how to select better user groups for more accurate now-casting.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jisun An, Ingmar Weber. 2015-11-13. Whom Should We Sense in "Social Sensing" -- Analyzing Which Users Work Best for Social Media Now-Casting. https://arxiv.org/abs/1511.04134

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

User Influence Analysis Based on Blogs

Rumor and word of mouth spread at the same speed as the highway of information diffusion in the age of the internet. Social networks play quite an important role in the huge internet. Nowadays, social networks have become indispensable in our lives, especially for the government and enterprises. A social network becomes a complex information diffusion network with users working as nodes and the relationships between users working as the vehicle. In this paper, we propose three kinds of algorithms for computing user influence based on the behavior of a user's forwarding microblogs and the symbol of @ in microblogs. We evaluate the effectiveness of the algorithms by comparing the results of our work with the training data in the dataset, and in the end, it proves that our algorithms work well.

cs.SI

Location transparency reduces activity by accounts misrepresenting their location on X

Concerns about inauthentic accounts, including foreign actors posing as domestic voices, are central to debates about online discourse. Yet, little is known about accounts with inaccurate location claims and how they behave when discrepancies between their claimed and actual locations become publicly visible. In November 2025, X introduced an "About this account" feature that discloses each account's platform-inferred location of operation. We leverage this intervention in a large-scale quasi-experimental study of 8,200 politically engaged accounts claiming a U.S. location, comparing accounts whose disclosed locations matched versus contradicted their claims across 1.3 million posts and 3.6 million replies over 21 weeks. Before disclosure, location-mismatched accounts posted more misleading, scam-related, and cryptocurrency-related content, but showed no distinctive partisan leaning. Difference-in-differences estimates show that disclosure reduced the posting activity of location-mismatched accounts by 13.1% with the largest declines among accounts revealed to be in Africa (29.2%) and Asia (24.4%), and among accounts with VPN flags, username changes, or scam- and crypto-heavy content. Additionally, the decline in their replies was concentrated in interactions with U.S.-based recipients (10.3%), whereas replies to non-U.S.-based recipients showed no statistically significant change. Conversely, there was no significant change in average audience engagement with their posts. Location transparency thus works primarily by inducing restraint among the disclosed accounts rather than by shifting audience behaviour, and the accounts it constrains look at least as much like cross-border fraud as foreign political influence.

cs.SI

The Benefit of Collective Intelligence in Community-Based Content Moderation is Limited by Overt Political Signalling

Social media platforms face increasing scrutiny over the rapid spread of misinformation. In response, many have adopted community-based content moderation systems, including Community Notes (formerly Birdwatch) on X (formerly Twitter), Community Notes on Meta, and Footnotes on TikTok. However, research shows that the current design of these systems can allow political biases to influence both the development of notes and the rating processes, reducing their overall effectiveness. We hypothesise that enabling users to collaborate on writing notes, rather than relying solely on individually authored notes, can enhance the overall quality of their notes. To test this idea, we conducted an online experiment in which participants jointly authored notes on politically misleading posts. We find that collaboration improves the helpfulness of notes, although the average effect depends on the interactional context. In particular, the benefits of collaboration decline when participants are made aware of one another's political affiliations. We also find that politically diverse teams improve note quality when evaluating Republican posts, while team composition does not meaningfully affect note quality for Democrat posts. These findings underscore the complexity of community-based content moderation and highlight the importance of understanding group dynamics and political diversity when designing more effective moderation systems.

cs.SI