Linking Scalar-Intensity Language to Structural Polarization with Validated Signed-Network Measures
Polarization in online communities is often studied through either language or interaction structure, but the two views are rarely connected within a unified framework. Prior work has linked them by constructing interaction graphs from human judgements of agreement and disagreement, leaving a gap between language as observed text and structure as an engineered representation of that text. We address this gap with a language-grounded signed-network pipeline that derives signed relations directly from conversational exchanges and links window-level language patterns to structural polarization over time. Before examining this relationship, we compare spectral and frustration-based polarization measures on synthetic benchmarks and real interaction networks. We find that frustration-based measures normalized by the graph's cycle-space capacity provide a more suitable basis for comparing polarization across networks of different sizes and densities. We therefore carry forward two complementary frustration-based measures: a weighted form that incorporates stance-model confidence and a count form based on edge signs. The weighted measure provides the better-behaved structural estimate and aligns more closely with polarization measured from human-labelled interactions, while the count measure reveals a stronger relationship with language. Across monthly Reddit Brexit discussions, greater prevalence of scalar-intensity language is associated with greater structural polarization, with a similar rank-level pattern in the human-labelled network. We find little evidence that language in one month predicts polarization in the next, whereas contemporaneous scalar-intensity prevalence provides useful information about polarization within the same month.