SonicJobs Logo
Left arrow iconBack to search

Grafana & Observability Engineer - Dallas, Tampa & Jersey City

StradIT
Posted 16 hours ago, valid for 11 days
Location

Jersey City, NJ, US

Salary

Competitive

Contract type

Full Time

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.

Sonic Summary

info
  • The role of Grafana & Observability Engineer requires 5 to 10 years of experience in observability, monitoring, operations, or platform engineering.
  • This position involves administering Grafana deployments, designing observability solutions, and establishing governance processes in locations such as Jersey City NJ, Tampa FL, and Dallas TX.
  • Key responsibilities include developing Terraform modules, automating onboarding processes, and implementing monitoring and alerting strategies to enhance operational visibility.
  • Candidates should possess a Bachelor's degree in Computer Science or a related field, along with hands-on experience with Grafana and strong skills in Linux and cloud platform administration.
  • The salary for this W2 employment role is competitive and commensurate with experience.

Role: Grafana & Observability Engineer

Experience: 5 to 10 years

Employment: W2

Location: Jersey City NJ, Tampa FL & Dallas TX

Key Responsibilities

Observability Platform Engineering

  • Administer and support Grafana Cloud and on-premises Grafana deployments.
  • Design and implement enterprise observability solutions for metrics, logs, traces, synthetic monitoring, and alerting.
  • Establish and maintain observability standards, best practices, and governance processes.
  • Configure and manage Grafana data sources, alerting, RBAC, folders, teams, and integrations.
  • Ensure platform scalability, reliability, resiliency, and operational excellence.

Automation & Infrastructure as Code

  • Develop and maintain Terraform modules for Grafana infrastructure and configuration management.
  • Automate onboarding of applications, infrastructure, dashboards, alerts, and data sources.
  • Build self-service capabilities that reduce manual operational effort and improve adoption.
  • Integrate observability capabilities into CI/CD and infrastructure provisioning workflows.

Monitoring, Alerting & Incident Management

  • Design meaningful monitoring and alerting strategies based on service health and business-critical workflows.
  • Implement and optimize alerting standards to reduce noise and improve signal quality.
  • Support incident response, troubleshooting, root cause analysis, and post-incident reviews.
  • Drive continuous improvement of operational visibility and platform health.

OpenTelemetry & Telemetry Engineering

  • Implement and support OpenTelemetry instrumentation across applications and infrastructure.
  • Establish standards for logs, metrics, traces, and telemetry collection.
  • Support telemetry pipelines, agent deployments, and data collection strategies.
  • Assist teams with instrumentation design and observability adoption.

Migration & Modernization

  • Support migration initiatives from legacy observability platforms to Grafana.
  • Analyze existing monitoring, alerting, logging, and tracing implementations and recommend modernization approaches.
  • Develop reusable migration patterns, automation, and engineering standards.
  • Partner with application teams to accelerate adoption of enterprise observability capabilities.

Collaboration & Leadership

  • Work closely with application development, infrastructure, cloud, and SRE teams.
  • Provide technical leadership and mentoring to engineers across the organization.
  • Contribute to observability architecture, strategy, and roadmap development.
  • Promote observability as a core engineering practice across the enterprise.

Required Qualifications

  • Bachelor's degree in Computer Science, Engineering, Information Systems, or related field.
  • 5+ years of experience in observability, monitoring, operations, or platform engineering.
  • Hands-on experience administering Grafana in large-scale enterprise environments.
  • Strong experience with Terraform and Infrastructure as Code practices.
  • Experience implementing monitoring, alerting, logging, and distributed tracing solutions.
  • Experience with OpenTelemetry concepts, instrumentation, and telemetry pipelines.
  • Strong Linux and cloud platform administration skills.
  • Experience with scripting and automation using Python, PowerShell, Bash, or similar languages.
  • Knowledge of operational excellence, reliability engineering, and incident management practices.

Preferred Qualifications

  • Experience migrating from tools such as Splunk, Dynatrace, AppDynamics, New Relic, OpenText OBM, or similar platforms.
  • Experience with Grafana Alloy, Tempo, Loki, Mimir, or Prometheus.
  • Experience operating observability platforms in AWS environments.
  • Knowledge of Kubernetes, containers, and cloud-native observability.
  • Experience designing enterprise observability strategies and governance models.
  • Familiarity with CI/CD platforms and DevOps practices.

Desired Skills

  • Grafana Administration
  • Terraform
  • OpenTelemetry (OTEL)
  • Monitoring & Alerting
  • Observability Engineering
  • Platform Engineering
  • Linux Administration
  • AWS Cloud Services
  • Automation & Scripting
  • Incident Management
  • Infrastructure as Code
  • Telemetry Pipelines
  • Reliability Engineering
  • Root Cause Analysis
  • Enterprise Monitoring Architecture



Learn more about this Employer on their Career Site

Apply now in a few quick clicks

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.