SonicJobs Logo
Left arrow iconBack to search

Senior Data Engineer

Cohort AI Inc.
Posted a day ago, valid for a month
Location

Naperville, IL, US

Salary

Competitive

Contract type

Full Time

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.

Sonic Summary

info
  • Cohort AI is seeking a Senior Data Engineer with over 5 years of professional experience to enhance their data infrastructure for healthcare datasets.
  • The role involves designing, building, and maintaining scalable data pipelines, as well as developing high-performance ETL/ELT workflows using SQL, Python, and PySpark.
  • Candidates should have strong skills in data modeling, data warehousing, and distributed data systems, along with solid Linux command-line experience.
  • The position requires collaboration with various teams to create effective data solutions and support customer onboarding initiatives.
  • Salary details are not specified, but the role emphasizes a strong problem-solving capability and effective communication skills.

At Cohort AI, We’re looking for a Senior Data Engineer to help build and scale the data infrastructure behind our platform. You’ll work hands-on with large and complex healthcare datasets, develop reliable data pipelines, and collaborate closely with Data Science, Clinical Informatics, Product, and Infrastructure teams.

This is an opportunity for an experienced data engineer who enjoys solving challenging data problems, building production-quality systems, and working in an environment where healthcare data and AI come together.

What You’ll Do

  • Design, build, and maintain scalable, reliable data pipelines.

  • Develop high-performance ETL/ELT workflows using SQL, Python, PySpark, and Apache Spark.

  • Work with complex healthcare datasets, including clinical and claims data.

  • Build ingestion and transformation workflows that are reliable, maintainable, and scalable.

  • Develop and improve data quality, monitoring, alerting, and observability solutions.

  • Troubleshoot and resolve production data pipeline issues using Linux and shell-based tools.

  • Optimize data processing performance and cloud infrastructure costs.

  • Apply strong engineering practices around code quality, testing, version control, and CI/CD.

  • Contribute to data modeling, data warehousing, and distributed data architecture decisions.

  • Work closely with Data Science, Clinical Informatics, Product, Infrastructure, and Commercial teams to translate requirements into effective data solutions.

  • Support customer onboarding and complex data integration initiatives.

  • Participate in design and code reviews and contribute to improving engineering practices.

  • Help identify opportunities to improve the scalability, reliability, and efficiency of our data platform.

What You Bring

  • 5+ years of professional Data Engineering experience.

  • Strong SQL skills and experience working with large datasets.

  • Strong hands-on experience with Python and PySpark.

  • Experience working with Apache Spark and distributed data processing.

  • Hands-on experience with modern data platforms such as Databricks, Snowflake, BigQuery, Redshift, or similar technologies.

  • Experience designing and operating production-grade ETL/ELT pipelines.

  • Strong understanding of data modeling, data warehousing, and distributed data systems.

  • Solid Linux command-line experience, including bash and shell scripting.

  • Experience implementing data quality, monitoring, and alerting frameworks.

  • Experience with Git and CI/CD.

  • Strong problem-solving skills and the ability to independently investigate and resolve complex technical issues.

  • Strong communication skills and the ability to collaborate effectively with cross-functional teams.

Nice to Have

  • Experience working with healthcare data and standards such as OMOP, FHIR, HL7, ICD, CPT, claims, EHR/EMR, or related datasets.

  • Experience with AWS, GCP, or Azure.

  • Experience working with large-scale distributed systems.

  • Familiarity with Airflow, Dagster, Prefect, or similar workflow orchestration tools.

  • Exposure to Generative AI, LLMs, or AI-enabled data applications.

  • Experience working in healthcare, life sciences, health technology, or a data-intensive environment




Learn more about this Employer on their Career Site

Apply now in a few quick clicks

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.