A Vision-Language Framework for Measuring Social Life on Sidewalks
While a number of methods exist for counting pedestrians in street-view imagery, these mostly ignore the social dimensions of pedestrian activity. A street traversed by a high volume of pedestrians has the same headcount as a street where people linger, sit, and socialize. This paper presents a vision-language framework for extracting social indicators from street-level imagery. Panoramic street-level imagery is reprojected to sidewalk-facing sideviews with preserved timestamps. A vision-language model (VLM)-based activity detection system codes each person across ten independent observable dimensions, resolving a systematic failure mode in which models prompted with high-level social categories conflate observable states with contextual inferences. The resulting social indicator system produces a Social Dwelling Index (SDI) that jointly considers pedestrian grouping and dwelling, provides activity labels documenting behavioral diversity, and issues binary flags for the presence of accessibility-sensitive populations. We apply the framework to 102,514 sideviews in New York City, revealing that pedestrian volume and SDI are only weakly associated (r = 0.168): streets with the highest foot traffic are not where social activity is most intense. The framework provides a scalable method for measuring not only how many people are on city sidewalks, but also their grouping, posture, and activity type, summarizing the non-transient activities that occur on city sidewalks.