Search arXivSearch

arXiv · 1711.06261

A study of variability induced by events dependency in microelectronic production

Abstract

-Complex manufacturing systems are subject to high levels of variability that decrease productivity, increase cycle times and severely impact the systems tractability. As accurate modelling of the sources of variability is a cornerstone to intelligent decision making, we investigate the consequences of the assumption of independent and identically distributed variables that is often made when modelling sources of variability such as down-times, arrivals, or process-times. We first explain the experiment setting that allows, through simulations and statistical tests, to measure the variability potential stored in a specific sequence of data. We show from industrial data that dependent behaviors might actually be the rule with potentially considerable consequences in terms of cycle time. As complex industries require strong levers to allow their tractability, this work underlines the need for a richer and more accurate modelling of real systems. Keywords-variability; cycle time; dependent events; simulation; complex manufacturing; industry 4.0 I. Accurate modelling of variability and the independence assumption Industry 4.0 is said to be the next industrial revolution. The proper use of real-time information in complex manufacturing systems is expected to allow more customization of products in highly flexible production factories. Semiconductor High Mix Low Volume (HMLV) manufacturing facilities (called fabs) are one example of candidates for this transition towards "smart industries". However, because of the high levels of variability, the environment of a HMLV fab is highly stochastic and difficult to manage. The uncontrolled variability limits the predictability of the system and thus the ability to meet delivery requirements in terms of volumes, cycle times and due dates. Typically, the HMLV STMicroelectronics Crolles 300 fab regularly experiences significant mix changes that result in unanticipated bottlenecks, leading to firefighting to meet commitment to customers. The overarching goal of our strategy is to improve the forecasting of future occurrences of bottlenecks and cycle time issues in order to anticipate them through allocation of the correct attention and resources. Our current finite capacity projection engine can effectively forecast bottlenecks, but it does not include reliable cycle time estimates. In order to enhance our projections, better forecast cycle time losses (queuing times), improve the tractability of our system and reduce our cycle times, we now need accurate dynamic cycle time predictions. As increased cycle-time is the main reason workflow variability is studied (both by the scientific community and practitioners, see e.g. [1] and [2]), what follows concentrates on cycle times. Moreover, the "variability" we account for should be understood as the potential to create higher cycle times, even though "variability" may be understood in a broader meaning. This choice is made for the sake of clarity, but the methodology we propose and the discussion we lead can be applied to any other measurable indicator. Sources of variability have been intensely investigated in both the literature and the industry, and tool down-times, arrivals variability as well as process-time variability are recognized as the major sources of variability in that sense that they create higher cycle times (see [3] for a review and discussion). As a consequence, these factors are widely integrated into queuing formulas and simulation models with the objective to better model the complex reality of manufacturing facilities. One commonly accepted assumption in the development of these models is that the variables (MTBF, MTTR, processing times, time between arrivals, etc.) are independent and identically distributed (i.i.d.) random variables. However, these assumptions might be the reason for models inaccuracies as [4] points out in a literature review on queuing theory. Several authors have studied the potential effects of dependencies, such as [5] who studied the potential effects of dependencies between arrivals and process-times or [6] who investigated dependent process times, [4] also gives further references for studies on dependencies effects. In a previous work [3], we pinpointed a few elements from industrial data that questioned the viability of this assumption in complex manufacturing systems. Figure 1: Number of arrivals per week from real data (A) and generated by removing dependencies (B)

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Kean Dequeant, Pierre Lemaire, Marie-Laure Espinouse, Philippe Vialletelle. 2017-11-16. A study of variability induced by events dependency in microelectronic production. https://arxiv.org/abs/1711.06261

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Online activity prediction via generalized Indian buffet process models

Online A/B tests are the standard tool for data-driven decision-making at scale. Among the design choices with the largest impact on statistical power is the triggering mechanism: how many users to expose and for how long. This often requires forecasting user engagement, i.e., whether enough users will trigger, and when a target participation level will be reached, from limited pilot data. We introduce a Bayesian nonparametric model for predicting both new-user counts and total triggers, accommodating the heavy-tailed engagement patterns typical of web experiments. All predictive quantities can be computed without intensive numerical procedures such as Markov chain Monte Carlo (MCMC) or variational inference. We evaluate on three public datasets (over 450 public benchmark evaluations) and a proprietary benchmark drawn from 759 production A/B tests comprising 1,774 arms. Across the benchmark analyses, our models are competitive and frequently improve accuracy in forecasting new users, total triggers, and time to reach a target sample size compared with state-of-the-art competitors, especially when only a few pilot days are observed.

stat.AP

Spending Scarce Confirmatory PET Measurements: Target-Aligned Validation in A4/LEARN

Anti-amyloid therapies and blood-based biomarkers are changing Alzheimer disease workups into a two-stage measurement workflow: screen broadly with cheaper information, then spend scarce confirmatory amyloid measurements where they support the decision that will be reported. Amyloid positron-emission tomography (PET) remains one such protocol measurement for amyloid burden, but PET slots, trial budgets, and payer-facing evidence packages are finite. This paper asks a deliberately operational question: when is simple transparent PET validation enough, and when is a fitted residual-uncertainty score worth the added complexity? For a weighted protocol target, the first-order value of validating subject i is the product of target influence and residual protocol uncertainty. Generic uncertainty sampling uses only the second factor and can spend PET measurements on subjects that are hard to predict but weak for the scientific, clinical, or commercial claim. We apply this rule to the A4/LEARN PET archive, treating observed PET as a design laboratory for scarce-confirmation studies. For the primary APOE4 carrier versus non-carrier contrast in Centiloid 24-or-higher PET positivity, simple APOE4-balanced validation recovers nearly all of the target-specific gain: at PET budget 200, the confidence-interval width ratio relative to random validation is 0.923 for APOE4 balancing and 0.914 for target-specific scoring, while generic uncertainty sampling is 0.980. Other targets behave differently: target-specific scoring gives larger gains for an age-slope analysis and for cutoff-indexed PET positivity. The practical message is simple: spend scarce protocol measurements according to the claim being validated, not only according to prediction uncertainty.

stat.AP

Identifying Damage Pathways Linking Sequence Composition to Storage Failure in DNA Data Storage via High-Dimensional Mediation Analysis

DNA data storage offers extraordinary information density and long-term durability, but its reliability is limited by sequence-dependent errors introduced during synthesis and accumulated during storage. It remains unclear how sequence composition is associated with storage failure through specific molecular damage components. We develop a high-dimensional semiparametric mediation framework for survival outcomes. GC content is treated as the exposure, a high-dimensional baseline damage spectrum (a vector of per-read damage counts stratified by trinucleotide context and error type) as the mediator, and storage-quality failure as the outcome. Nonlinear covariate effects in both the mediator and survival models are approximated using deep neural networks. A three-step procedure combining product-of-coefficients screening, Smoothly Clipped Absolute Deviation (SCAD) penalized estimation, and joint significance testing is developed for mediator selection and inference. Applied to an aging experiment on electrochemically synthesized DNA, the method identifies 14 significant mediators, all corresponding to single-base deletions, with estimated mediated effects concentrated in trinucleotide contexts ending in C. These results reveal deletion-type damage as a major pathway linking sequence composition to reduced archival reliability and suggest candidate sequence features for future optimization and error-control strategies. The proposed framework thus offers a mechanism-oriented statistical approach for understanding and improving the reliability of DNA data storage.

stat.AP