SonicJobs Logo
Left arrow iconBack to search

Platform Engineer - Observability

Finsense Africa
Posted 7 days ago, valid for 13 days
Location

Nairobi, Nairobi County

Salary

Competitive

Contract type

Contract

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.

Sonic Summary

info
  • We are seeking a Platform Engineer with 3 to 5+ years of experience to design and implement a centralized observability platform in a large enterprise environment.
  • This is a 6-month renewable contract position focused on enhancing visibility across critical applications, APIs, and infrastructure using tools like Grafana and Prometheus.
  • The role requires strong hands-on experience with Grafana, Prometheus, OpenTelemetry, and Kubernetes or OpenShift, as well as familiarity with Microsoft Azure.
  • Key responsibilities include setting up centralized logging, metrics, and alerting systems, along with troubleshooting complex issues using logs and metrics.
  • The salary for this position is competitive and commensurate with experience.

Job Description

We are looking for an experienced Platform Engineer to design, build, and implement a centralized observability platform across a large enterprise technology environment.


This is a 6-month renewable contract.

The role will focus on improving visibility across critical applications, APIs, integration services, and infrastructure using Grafana, Prometheus, Grafana Loki, and OpenTelemetry.

The successful candidate will work across Microsoft Azure and OpenShift/Kubernetes environments to establish centralized logging, metrics, tracing, dashboards, and automated alerting.

Key responsibilities include:

  • Design, deploy, and maintain a centralized observability platform using Grafana, Prometheus, and Grafana Loki.
  • Implement and manage OpenTelemetry collectors for logs, metrics, and distributed tracing.
  • Build telemetry pipelines for applications, APIs, infrastructure, integration services, and security-related logs.
  • Configure operational and management dashboards to monitor system health, performance, availability, and incidents.
  • Set up automated alerts for application errors, API latency, service timeouts, infrastructure issues, and abnormal system behaviour.
  • Define and maintain standards for structured logging, error codes, trace IDs, correlation IDs, and application telemetry.
  • Ensure internally developed and third-party applications comply with agreed observability and logging standards.
  • Develop troubleshooting guides and operational runbooks for support teams.
  • Support L1 and L2 teams in using dashboards, logs, alerts, and traces to diagnose incidents.
  • Collaborate with software engineering, infrastructure, security, QA, and operations teams.
  • Integrate synthetic monitoring and application health checks into the wider monitoring platform.
  • Help improve Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR) across critical technology services.


Requirements

  • 3 - 5+ years of experience in Platform Engineering, DevOps, Site Reliability Engineering, Cloud Engineering, or a similar role.
  • Strong hands-on production experience with Grafana, Prometheus, and Grafana Loki.
  • Strong experience with OpenTelemetry, including collectors, instrumentation, distributed tracing, and trace propagation.
  • Experience designing and operating centralized logging, monitoring, metrics, and alerting platforms.
  • Strong experience with Kubernetes and/or OpenShift.
  • Hands-on experience with Microsoft Azure infrastructure and services.
  • Experience working in hybrid cloud and on-premise environments.
  • Experience with Infrastructure as Code and automation tools such as Terraform, Ansible, GitHub Actions, Azure DevOps, or similar.
  • Good understanding of application and infrastructure logging, log parsing, and structured logging.
  • Experience monitoring APIs, microservices, and enterprise applications.
  • Ability to troubleshoot complex application and infrastructure issues using logs, metrics, and traces.
  • Experience defining technical standards and ensuring engineering teams follow them.
  • Strong communication skills and the ability to work with engineering, infrastructure, security, and support teams.

Experience with Java/Spring Boot, Node.js, enterprise integration platforms, WAF/security logging, or regulated enterprise environments would be an added advantage.






Learn more about this Employer on their Career Site

Apply now in a few quick clicks

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.