High Performance Computing Data Center Operator positions focus on delivering results in their domain. This page aggregates open High Performance Computing Data Center Operator roles and what employers typically expect.
## Company Description **Join us and make YOUR mark on the World!** Lawrence Livermore National Laboratory (LLNL) has turned bold ideas into world-changing impact advancing science and technology to strengthen U.S. security and promote global stability. Our mission spans four critical national security areas nuclear deterrence, threat preparedness, energy security, and multi-domain defense empowering teams to take on the toughest challenges of today and tomorrow. With a culture built on innovation and operational excellence, LLNL is a place where your expertise can make a real impact. ## Job Description We have an opening for a **High-Performance Computing (HPC) Data Center Operator** to monitor, diagnose, troubleshoot, and repair system faults on a large number of high-performance computer (HPC) systems, storage systems and networks, working under minimal supervision. You will interact with other Livermore Computing (LC) staff to remediate problems and provide advanced technical support in a complicated HPC computing and networking environment, working either Swing (4:00pm – 12:00am) or Owl Shift (12:00am – 8:00am). This position is in the Livermore Computing Operations Group in the LC Division within the Computing Directorate. *This position requires full-time on-site presence due to the nature of the work.* This position will be filled at either level based on knowledge and related experience as assessed by the hiring team. Additional job responsibilities (outlined below) will be assigned if hired at the higher level. **You will** - Provide intermediate technical support and operational monitoring for HPC systems including large Linux clusters, file systems, storage systems, and associated infrastructure. - Apply working knowledge of Linux/Unix systems and use in-house and vendor-supplied tools to monitor systems, diagnose issues, and perform routine repairs and recovery actions. - Utilize the Laboratory’s trouble ticketing system, ServiceNow, for problem ticket tracking. - Receive, document, triage, and respond to customer issues during business and off-hours, resolving routine problems or escalating to appropriate technical staff. - Perform data center facilities monitoring, problem remediation, and emergency event response during normal daily operation and off-hours. - Participate in system installation, hardware swaps, system relocation, and decommissioning activities in support of ongoing data center operations. - Promote the use of inter-departmental resources for tools, metrics, and common solutions to team members via email and presentations. - Perform other duties as assigned. **In Addition at the 525.3 Level** - Provide advanced technical support and monitoring capabilities for the HPC systems clusters, file systems, and storage systems under minimal supervision. - Independently troubleshoot and resolve moderately complex hardware, software, operating system, network, and infrastructure issues; analyze symptoms, determine root cause, implement corrective actions, and coordinate escalation when required. - Perform advanced technical tasks including installation, diagnosis, repair and maintenance of clustered computer systems and related file systems and networks. - Analyze system events, alarms, logs, and monitoring data to identify patterns, isolate faults, and recommend improvements to procedures, tools, or operational practices. - Serve as a knowledgeable resource to team members during off-hours operations and contribute to the development, refinement, and documentation of operating procedures and response practices. ## Qualifications - Ability to obtain and maintain a U.S. DOE Q-level security clearance which requires U.S. Citizenship. - Associate’s degree in a computer-related field or equivalent combination of technical training and experience. - General working knowledge of networking concepts, protocols, and connectivity troubleshooting. - Demonstrated experience using Linux command-line utilities for sys…