SonicJobs Logo
Left arrow iconBack to search

Staff HPC Software Engineer

San Diego Stealth Startup
Posted 21 hours ago, valid for a month
Location

San Diego, CA, US

Salary

$174,000 - $185,000 per year

Contract type

Full Time

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.

Sonic Summary

info
  • The position is for an experienced systems engineer in San Diego, CA, with a salary range of $174,000 - $185,000.
  • The role requires a PhD with 3+ years, a master's with 6+ years, or a bachelor's with 8+ years of relevant experience.
  • The engineer will focus on building and improving a high-throughput compute architecture for data ingestion, processing, storage, and delivery.
  • Key responsibilities include establishing hardware benchmarks, productizing machine learning models, and ensuring reliable data and compute paths under various constraints.
  • Preferred qualifications include experience with GPU computing, high-throughput storage, and resilient data pipelines.

Location: San Diego, CA

Job Type: Full-Time

Salary: $174,000 - $185,000

Position Overview

We are looking for an experienced systems engineer to build and improve a measured, reliable compute architecture. The work spans high-throughput data ingestion, processing, storage, and delivery. 

Role mission 

Build and improve a high-throughput compute stack, so it is fast, observable, recoverable, and practical to operate in production environments. 

This person will work with algorithms, platforms, and infrastructure engineers. They will not be expected to own every algorithm or infrastructure service. Their core responsibility is making the data and compute path reliable under real throughput, storage, network, and latency constraints. 

Early work 

In the first three to six months, this person should help: 

  • Establish reproducible hardware benchmarks for accelerated compute, CPU workloads, memory transfers, storage throughput, and network streaming. 
  • Build or harden stateful, multi-threaded pipelines that move data from ingestion through compute and output. 
  • Productize machine learning models, including neural networks, tree-based models, and unsupervised models, so they meet production requirements for performance, reliability, observability, and quality. 
  • Define backpressure, checkpointing, retry, and recovery behavior for disk pressure, slow consumers, and network outages. 
  • Compare alternative processing designs using wall-clock time, memory, storage, and quality measurements. 
  • Make the production interfaces and performance tests durable enough that later algorithm changes do not quietly break throughput or recovery behavior. 

Required experience 

  • This role requires a PhD in Computer Science, Life Sciences, or a related discipline with 3+ years of relevant experience; a master's degree with 6+ years of relevant experience; or a bachelor's degree with 8+ years of relevant experience.  
  • Has contributed to a complex production software system with state machines, concurrency, and real compute or I/O bottlenecks. They do not need to have been the technical lead but must understand how these systems fail and how to debug them. 
  • Has shipped production-quality software in at least one of C++, Rust, CUDA, C, or C#. Comfortable with the normal engineering tools: profiling, tracing, debugging, testing, code review, builds, and CI. 
  • Can reason concretely about throughput, latency, buffering, memory, storage, network behavior, scheduling, contention, and failure recovery. 
  • Uses measurements to guide performance work: can identify a bottleneck, make a targeted change, quantify the gain, and add a regression guard. 
  • Has experience productizing machine learning models, including neural networks, tree-based models, or unsupervised models. Can make these models reliable, measurable, and efficient in a production system. This is not a model-research role. 
  • Works well in a flat, highly technical team: can state tradeoffs clearly, contribute outside a narrow specialty, learn from others, and strengthen areas where the team is currently thin. 

Strongly preferred 

  • GPU computing, CUDA profiling, or heterogeneous CPU/GPU pipelines. 
  • High-throughput storage, networking, streaming, or low-latency systems. 
  • Performance-sensitive scientific computing or another data-intensive system where delayed or failed processing has operational consequences. 
  • Experience with resilient data pipelines: bounded queues, backpressure, checkpoint/restart, idempotent outputs, and operational telemetry. 

Explicit boundaries 

  • This is not primarily an MLOps, cloud-platform, data-science, or model-training position. 
  • This role complements algorithm and scientific development; it does not unilaterally set scientific requirements, quality criteria, or cloud/platform ownership. 
  • The near-term focus is the high-throughput compute path and its interfaces, not a general rewrite of company infrastructure. 

We are an equal opportunity employer. We thrive on diversity and collaboration.

 




Learn more about this Employer on their Career Site

Apply now in a few quick clicks

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.