arXiv · cs/0306074
Understanding and Coping with Hardware and Software Failures in a Very Large Trigger Farm
Abstract
When thousands of processors are involved in performing event filtering on a trigger farm, there is likely to be a large number of failures within the software and hardware systems. BTeV, a proton/antiproton collider experiment at Fermi National Accelerator Laboratory, has designed a trigger, which includes several thousand processors. If fault conditions are not given proper treatment, it is conceivable that this trigger system will experience failures at a high enough rate to have a negative impact on its effectiveness. The RTES (Real Time Embedded Systems) collaboration is a group of physicists, engineers, and computer scientists working to address the problem of reliability in large-scale clusters with real-time constraints such as this. Resulting infrastructure must be highly scalable, verifiable, extensible by users, and dynamically changeable.
Explore related subjects
Keep this discovery
Jim Kowalkowski. 2003-06-13. Understanding and Coping with Hardware and Software Failures in a Very Large Trigger Farm. https://arxiv.org/abs/cs/0306074
Cite the original work for its findings. Save a collection to share your selection of sources.