Responsibilities
- Lead integration of scale up (e.g. NVlink, XGMI, RoCE) and scale out (e.g. NICs) interfaces for AI Platforms
- Develop understanding of Collective Communication patterns/ AI workloads and incorporate as part of new product introduction (NPI)
- Proactively create experiments and tooling to detect, reproduce and diagnose hardware/firmware/software issues
- Contribute to enabling hacks for future technology explorations in AI space
- Troubleshoot, diagnose and root-cause system failures and isolate the components/failure scenarios while working with internal & external partners
- Develop visibility through data visualization and implement systemic solutions to hardware health issues
- Leverage production experience to drive external and internal teams to continuously improve product quality
Minimum Qualifications
- Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
- 8+ years of work experience in one or more domains such as: Network ASIC/Platform Development (Silicon/Switch Platform design or bring-up or characterization), Network Product Deployment and Customer Support (Switches, NICs), Interconnect Technologies (e.g. Optics, DAC)
- Knowledge of TCP/IP and experience using tools like iperf/uperf
- Knowledge of server architecture and components
- Experience working with Linux
- Hands-on troubleshooting and debug experience
Preferred Qualifications
- Experience working with RDMA/RoCE, including scale-out networks
- Experience with Python scripting
- Experience working with Network Interface Cards (NICs)
- Experience working with AI server systems
- Experience working with large scale deployments
$173,000/year to $245,000/year + bonus + equity + benefits
Learn more about this Employer on their Career Site
