Search arXivSearch

arXiv · 1909.09453

Application of Clustering Analysis for Investigation of Food Accessibility

Abstract

Access to food assistance programs such as food pantries and food banks needs focus in order to mitigate food insecurity. Accessibility to the food assistance programs is impacted by demographics of the population and geography of the location. It hence becomes imperative to define and identify food assistance deserts (Under-served areas) within a given region to find out the ways to improve the accessibility of food. Food banks, the supplier of food to the food agencies serving the people, can manage its resources more efficiently by targeting the food assistance deserts and increase the food supply in those regions. This paper will examine the characteristics and structure of the food assistance network in the region of Ohio by presenting the possible reasons of food insecurity in this region and identify areas wherein food agencies are needed or may not be needed. Gaussian Mixture Model (GMM) clustering technique is employed to identify the possible reasons and address this problem of food accessibility.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Rahul Srinivas Sucharitha, Seokcheon Lee. 2019-09-18. Application of Clustering Analysis for Investigation of Food Accessibility. https://arxiv.org/abs/1909.09453

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Calibration-guided data fusion: A framework for understanding data contribution in multi-source inference through efficient Bayesian optimization with application to virus dynamics parameter estimation

We introduce a framework that uses simulation-based calibration to assess data contribution in multi-source inference, systematically evaluating what each data source contributes to parameter estimation. When inference is fast enough, calibration becomes a practical diagnostic tool. We demonstrate the framework using Bayesian optimization likelihood-free inference (BOLFI) applied to a target cell limited model of influenza A virus dynamics, combining RNA measurements and endpoint dilution assay data from single-cycle and multiple-cycle experiments. BOLFI achieves a 20-fold speed-up over traditional MCMC and requires no likelihood knowledge. A Gaussian process classifier handles simulation failures, and a discrepancy design integrates heterogeneous data types. The results reveal a counter-intuitive finding: combining all data sources does not always improve inference. The single-cycle experiment alone is most reliable for the virion entry rate, while ED-only data is best for the total RNA production rate. ED data improves calibration of the infectious virion production rate and the infectious phase duration when combined with RNA data. For the eclipse phase duration, the multiple-cycle experiment alone is more reliable, indicating that adding single-cycle data introduces inconsistency. This reveals a trade-off between precision and calibration. Based on these findings, we explore a composite diagnostic that selects the best-calibrated data type for each parameter, serving as a sensitivity check. The inference also reveals a 35-fold disparity between RNA and infectious virus production. This work demonstrates that computational efficiency enables a new level of methodological rigor: systematic calibration-guided data fusion. The framework is general and applicable to other multi-source inference problems where simulation-based models are used.

stat.AP

Causal Inference with Video Features as Treatments

We develop the first statistical methodology for causal inference with video features as treatments. Video is the most engaging content modality on the internet. A central causal question is how audience reactions change in response to treatment features that unfold over the course of a video. Unfortunately, standard causal inference methods are not applicable because confounding features are latent, high-dimensional, and dynamically related to both the treatment sequence and the outcome trajectory. To address these challenges, we first reproduce each video using a deep generative model and leverage the model's internal representations as learned, low-dimensional summaries of video content for causal estimation. We then establish that the average potential-outcome trajectory under dynamic stochastic interventions is nonparametrically identified. Lastly, we propose a consistent and asymptotically normal estimator based on a longitudinal neural network architecture. We empirically validate our approach by constructing a new causal inference benchmark consisting of $10{,}000$ Super Mario Bros. levels played by fixed Mario AI agents, where ground-truth causal effects are known by construction. Finally, we apply our method to television advertisements from the 2020 U.S. presidential campaign and find that increasing the probability of a candidate appearing over time leads to higher average viewer evaluations. With the proposed methodology, researchers can ask which visual features, appearing at which points in a video, influence audience responses, while benchmarking new methods against datasets with known ground-truth causal effects.

stat.AP

Optimization of ReaxFF parameters for the $\mathrm{Mo-S}$ system using random optimization and coordinate search

ReaxFF is a molecular dynamics method that can be considered a good approximation to quantum methods for investigating reactive molecular systems consisting of ten thousand to one hundred thousand atoms. While ReaxFF is usually a much faster alternative to quantum methods, the force field consists of nearly 100 parameters per element, which makes the force field development a high dimensional optimization problem. In addition to the high-dimensionality, non-convexity and non-continuity make it a hard problem to optimize. We use random optimization along with coordinate search strategies to optimize efficiently and sample new parameter points that yield good molecular properties close to predefined `reference values' obtained from quantum mechanical methods for the $\mathrm{Mo-S}$ system. We also provide empirical error guaranties starting from any random sample of inputs. We discover new points for the $\mathrm{Mo-S}$ system at adjusted error levels of $13{,}000$ as compared to Sengul et al. (2022) at $70{,}000$ levels under the same loss function, registering over $80\%$ improvement. We also extend our algorithm to an out-of-sample system, $\mathrm{W-S}$, with no training data to record over $70\%$ improvement over Sengul et al. (2021).

stat.AP