Abstract
MOLAR is a multi-institutional research effort that concentrates on adaptive, reliable, and efficient operating and runtime system (OS/R) solutions for ultra-scale high-end scientific computing on the next generation of supercomputers. This research addresses the challenges outlined in FAST-OS (forum to address scalable technology for runtime and operating systems) and HECRTF (high-end computing revitalization task force) activities by exploring the use of advanced monitoring and adaptation to improve application performance and predictability of system interruptions, and by advancing computer reliability, availability and service-ability (RAS) management systems to work cooperatively with the OS/R to identify and preemptively resolve system issues. This paper describes recent research of the MOLAR team in advancing RAS for high-end computing OS/Rs.
Original language | English |
---|---|
Pages (from-to) | 63-72 |
Number of pages | 10 |
Journal | Operating Systems Review (ACM) |
Volume | 40 |
Issue number | 2 |
DOIs | |
State | Published - Apr 2006 |
Keywords
- Availability
- Fault Tolerance
- Group Membership
- High-End Computing
- Monitoring
- RAS
- Reliability