Search arXivSearch

arXiv · 1604.01455

Picking Winners in Daily Fantasy Sports Using Integer Programming

Abstract

We consider the problem of selecting a portfolio of entries of fixed cardinality for contests with top-heavy payoff structures, i.e. most of the winnings go to the top-ranked entries. This framework is general and can be used to model a variety of problems, such as movie studios selecting movies to produce, venture capital firms picking start-up companies to invest in, or individuals selecting lineups for daily fantasy sports contests, which is the example we focus on here. We model the portfolio selection task as a combinatorial optimization problem with a submodular objective function, which is given by the probability of at least one entry winning. We then show that this probability can be approximated using only pairwise marginal probabilities of the entries winning when there is a certain structure on their joint distribution. We consider a model where the entries are jointly Gaussian random variables and present a closed form approximation to the objective function. Building on this, we then consider a scenario where the entries are given by sums of constrained resources and present an integer programming formulation to construct the entries. Our formulation uses principles based on our theoretical analysis to construct entries: we maximize the expected score of an entry subject to a lower bound on its variance and an upper bound on its correlation with previously constructed entries. To demonstrate the effectiveness of our integer programming approach, we apply it to daily fantasy sports contests that have top-heavy payoff structures. We find that our approach performs well in practice. Using our integer programming approach, we are able to rank in the top-ten multiple times in hockey and baseball contests with thousands of competing entries. Our approach can easily be extended to other problems with constrained resources and a top-heavy payoff structure.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

David Scott Hunter, Juan Pablo Vielma, Tauhid Zaman. 2019-01-23. Picking Winners in Daily Fantasy Sports Using Integer Programming. https://arxiv.org/abs/1604.01455

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Why is Regularization Underused? An Empirical Study on Trust and Adoption of Statistical Methods

Statistical practice does not automatically follow methodological innovation. Regularization methods, widely advocated to reduce overfitting and stabilize inference, are readily available in modern software, but are not consistently used by data analysts. We investigate this implementation gap in a large-scale empirical study of trust in, and acceptance of, regularization techniques, based on $N = 606$ data analysts. Drawing on measurement frameworks from technology acceptance research, we survey practitioners and embed a randomized experiment to test whether written recommendation of regularization methods increases trust or intended use. We find no evidence of such an effect. Instead, adoption intentions are strongly associated with analysts' perceptions of ease of implementation and practical benefit, such as improved bias control or interpretability. Perceived social norms also emerge as a central driver. These results indicate that uptake of statistical methodology depends less on formal recommendations than on usability, perceived utility, and community practice.

stat.OT

Exact analysis of a split--merge queue with latent Erlang-factor dependent subtask times

This paper studies a two-server split--merge queue with positively dependent subtask service times modeled through a latent-factor bivariate Erlang construction. An exact characterization of the split--merge completion time is obtained, including explicit formulas for its first two moments and the resulting mean waiting time. Under fixed marginal service-time distributions, independence is shown to stochastically increase the completion time and hence overestimate mean waiting time. Numerical illustrations show that this benchmark gap can be substantial.

stat.OT

Statistical Compatibility, Refutational Information, and Acceptability

This paper develops an interpretive framework for divergence P-values and S-values within a descriptive frequentist perspective. Statistical analysis is framed as operating within idealized worlds defined by a set of assumptions and a target hypothesis, where probabilities describe the behavior of data under the model but do not assign truth values to hypotheses. Within this view, P-values are interpreted as graded indices of compatibility between the observed result and the predictions generated by the assumed model; accordingly, small P-values should not be read as indicating logical impossibility or strict inconsistency of the model itself. Building on this distinction, the paper argues that practical inference requires moving beyond the internal logic of the model toward judgments of overall acceptability, which depend not only on data-model compatibility but also on multiple contextual considerations such as subject-matter knowledge, plausibility of assumptions, data quality, usefulness, and loss - all interpreted through the competence, intentions, perceptions, and moral values of the specific analyst. S-values are therefore interpreted not as evidence against the epistemic status of the model, but as a specific form of refutational information that contributes to the broader body of information used by the analyst to judge whether a model remains acceptable for an intended practical purpose. The paper also examines the linguistic and conceptual risks associated with the language of incompatibility, distinguishes probability from rarity, and clarifies different notions of surprise - including a possible definition of Shannon-type surprise, to be distinguished from Bayesian belief revision. Overall, the article proposes a more cautious and explicit interpretation of frequentist measures, centered on model-based description, analyst responsibility, and decision acceptability.

stat.OT