About Skyhook Bio:
Skyhook accelerates the discovery and development of differentiated therapeutics by accessing vast biological design spaces previously intractable with current state of the art technologies. Our differentiated platform employs efficient AI-guided design, synthesis, and functional screening of diverse programmable biomolecule libraries at unprecedented speed and scale.
\n
-
Develop and apply machine learning methods to model sequence-function relationships from high-throughput assay data
-
Train and evaluate models on proprietary datasets and public protein sequence, structure, and function data
-
Design and deploy active learning strategies to accelerate protein engineering cycles
-
Extract and visualize insights from large-scale experimental datasets to inform protein engineering strategy
-
Model and account for noise, uncertainty, and variability in biological assay data
-
Collaborate closely with experimental and computational scientists to integrate laboratory and in silico workflows, including NGS-based analyses
-
Communicate computational strategies, results, and recommendations to multidisciplinary teams and leadership
-
PhD in Computer Science, Machine Learning, Computational Biology, Biophysics, Bioinformatics, Statistics, Genomics, or related field with a minimum of 2+ years industry experience in pharma or biotech R&D required.
-
Titling will be commensurate with experience
-
Strong foundation in machine learning, including model development, training, evaluation, and application to large-scale biological datasets
-
Experience with exploratory data analysis, statistical modeling, and data visualization to interpret experimental results and guide modeling decisions
-
Proficiency in modern ML frameworks such as PyTorch or TensorFlow, and writing efficient, well-structured, and maintainable code (Python required)
-
Experience working with cloud-based ML workflows and data pipelines (AWS preferred)
-
Strong analytical and problem-solving skills; comfortable working in a dynamic, collaborative startup environment
-
Clear written and verbal communication skills, with the ability to explain technical concepts to diverse audiences
-
Experience with protein-focused ML models and tools (e.g., ESM, AlphaFold, or related frameworks)
-
Experience designing or improving ML infrastructure and data pipelines, including cloud-based training, inference, data versioning, automation, and deployment
-
Experience with NGS data analysis pipelines and integrating experimental data with computational workflows
We kindly ask that recruiting agencies and third-party recruiters do not contact us regarding this position. We are not seeking assistance or accepting unsolicited resumes from agencies at this time.
Learn more about this Employer on their Career Site
