SonicJobs Logo
Left arrow iconBack to search

HPC Operations Engineer

Evergreen Statistical Trading
Posted 4 days ago, valid for 13 days
Location

Bellevue, WA, US

Salary

$200,000 per year

Contract type

Full Time

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.

Sonic Summary

info
  • Evergreen, a proprietary trading firm in Bellevue, Washington, is seeking an HPC Operations Engineer responsible for managing research clusters across multiple datacenters.
  • The ideal candidate should have strong hands-on Linux systems administration experience, particularly with Red Hat-family distributions, and proficiency in Python for automation.
  • Experience with tools such as Git, Prometheus, Grafana, Slurm, and various cluster provisioning tools is required, along with the ability to troubleshoot complex production issues under pressure.
  • The position offers an annual base salary of $200,000, along with additional signing and performance bonuses and company-paid medical benefits.
  • Candidates should possess a growth-oriented mindset and be able to work collaboratively within a team, with a passion for maintaining high-performance infrastructure.

Overview 

Evergreen is a proprietary trading firm based in Bellevue, Washington. We combine statistical research and technological excellence to trade in global financial markets.

As an HPC Operations Engineer at Evergreen, you will own the day-to-day operation of our research clusters across multiple datacenters, along with the scheduled workloads that our research and trading operations depend on. You will also have the unique opportunity to work on a technology stack that is unencumbered by legacy infrastructure, and to take on substantial projects across both software and hardware as our HPC footprint continues to expand.

Qualifications  

  • Strong hands-on Linux systems administration experience, particularly with Red Hat-family distributions such as RHEL, Rocky Linux, AlmaLinux, etc.
  • Strong proficiency with Python for automation and tooling
  • Experience with most of the following:
    • Version control tools, ideally Git, as well as versioned configuration management
    • Monitoring and alerting systems, ideally Prometheus and Grafana
    • Batch schedulers, ideally Slurm, including queue configuration, resource management and troubleshooting
    • Cluster provisioning tools, such as Warewulf, xCAT, etc.
    • ZFS and parallel filesystems, such as GPFS, Lustre, etc.
    • InfiniBand networking
  • Ability to own critical processes, design them not to break, and resolve critical issues outside of business hours when they do
  • Ability to direct hardware work through datacenter remote hands, and be hands-on when the situation calls for it
  • Ability to diagnose complex problems in production systems under time pressure
  • Growth-oriented and collaborative mindset; enjoys working within a team
  • Passion for high-performance infrastructure that runs correctly and with minimal drama

Annual base salary of $200,000. Additional signing and performance bonuses will be provided along with company-paid medical and/or other benefits. 

 

Equal Opportunity 

As an equal opportunity employer, Evergreen employs and values an equitable work environment. We hope to bring people together from many different backgrounds to create a workplace that is as diverse as the markets we trade. 




Learn more about this Employer on their Career Site

Apply now in a few quick clicks

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.