arXiv2026
Complex, high-throughput data acquisition and processing systems, such as those used in high-energy physics experiments, are increasingly moving sophisticated pattern recognition and data compression algorithms closer to the sensors themselves. To meet these needs, programmable device manufacturers offer multi-silicon die packages that commonly include dedicated co-processors within the same package. We present a technical study of a new family of such co-processors from AMD Xilinx, the Adaptive Intelligence (AI) Engine, or AIE, as part of the Versal architecture. Specifically, we focus on the deployment capabilities of AIEs in fixed latency environments with latency constraints of the order of microseconds, such as those typically encountered in colliding beam experiments like at CERN's Large Hadron Collider. We evaluate the performance of a vectorised implementation of both a Boosted Decision Tree and a Convolutional Neural Network, thereby demonstrating the feasibility of deploying AIEs for Machine Learning applications in such environments and their use as possible alternatives to traditional programmable logic-based implementations. In addition we are including an example of a CNN based architecture prototyped for the ATLAS Phase-II trigger, evaluated on open-data. We conclude that AIEs can augment the compute capabilities of field programmable gate array-based processing systems with microsecond-latency requirements.