
Site Reliability Engineer II
Job Description
Site Reliability Engineer II
Work Location: Berkeley, CA
Rate: Not Defined / need Candidate’s requested rate
Hours: 40 (This schedule consists of an onsite 5-day weekly schedule working Owl (Midnight–8 am) shifts to monitor the facility)
Position Details:
Seeking a Site Reliability Engineer to support a scientific computing center. As a member of a 24x7 team in the Operations Technology Group, the Site Reliability Engineer helps ensure systems remain accessible, reliable, and secure for the scientific user community. This role leverages advanced data collection and monitoring systems to proactively manage the health of the computing environment, keeping computational power an uninterrupted resource for fundamental scientific research.
Responsibilities:
Monitoring & Incident Response
Review and respond to alerts from computer systems, storage, network, and other data center/facility-related systems, triaging or engaging appropriate on-call staff.
Respond to alerts from multiple systems to ensure data collection continues 24/7, providing real-time information for diagnoses.
Perform regular physical and logical walkthroughs of the data center floor to monitor environmental health, power distribution units, and cooling infrastructure.
Provide accurate information in the trouble ticketing system for outages, maintenance updates, and other incidents.
Automation & Tooling
Create solutions to improve processes, prevent issue recurrence, and automate responses to routine service conditions.
Identify issues and propose solutions to improve monitoring capabilities or provide better automation for triage.
Develop and maintain tools within the monitoring pipeline in collaboration with the Operations Team; create new software to bring alerts and notifications from HPC system APIs into the monitoring pipeline.
Apply expertise in ServiceNow to develop and implement customized service management solutions.
Build and maintain application/tool configurations to ensure software runs reliably as data and user demands grow.
Cross-Team Coordination
Collaborate with other groups to ensure communication and workflows are clearly understood.
Work closely with other groups to coordinate center-wide maintenance activities and manage diagnostic and notification software during maintenance periods.
Problem Solving
Work on and resolve problems of diverse scope where data analysis requires evaluation of identifiable factors.
Demonstrate good judgment in selecting methods and techniques for obtaining solutions.
Work on and resolve complex issues where analysis of situations or data requires in-depth evaluation of variable factors.
Required Qualifications:
Experience in or willingness to work within a 24/7 onsite team environment supporting large-scale data centers or critical installations.
Strong hands-on experience with the Linux shell and working in a command-line (e.g., SSH) environment.
Experience developing tools using programming languages such as C, C++, Perl, Java, or Python, or a scripting language, with knowledge of standard software development practices.
Motivated self-starter able to learn technologies that improve data center management, such as Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, building management software, evaporative cooling, and power utilization.
Experience with network security, including configuring/maintaining ACLs and knowledge of firewalls.
Knowledge of large data communications networks, network protocols, and IT infrastructure supporting highly available systems and applications.
Experience collaborating across technical teams to resolve operational bottlenecks and ensure system reliability and alignment with service-level objectives.
Strong communication skills and ability to work effectively across multiple technical teams.
Preferred Qualifications:
Practical experience developing and deploying Agentic AI or autonomous automation tools — for example, autonomous agents that automate decision-making, optimize complex workflows, or enhance proactive system monitoring.
Experience with ServiceNow implementation.
Familiarity with ITSM best practices and understanding of how to align service lifecycles with business goals.
Physical Requirements: Must be capable of climbing, walking, kneeling, bending, and lifting as part of daily job responsibilities. Ability to lift a minimum of 35 – 60 lbs.
JOB NUMBER
#5274
LOCATION
Berkeley, CA
ESTIMATED DURATION OF WORK
11 months (extension likely)
