


















Fermilab operates several clusters for lattice gauge computing. Minimal manpower is available to manage these clusters. We have written a number of tools and developed techniques to cope with this task. We describe our tools which use the IPMI facilities of our systems for hardware management tasks such as remote power control, remote system resets, and health monitoring. We discuss our techniques involving network booting for installation and upgrades of the operating system on these computers, and for reloading BIOS and other firmware. Finally, we discuss our tools for parallel command processing and their use in monitoring and administrating the PBS batch queue system used on our clusters.
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。