SonicJobs Logo
Left arrow iconBack to search

AIOps Support Lead

BCE GLOBAL TECHNOLOGY CENTRE PRIVATE LIMITED
Posted 18 days ago, valid for 12 days
Salary

Competitive

Contract type

Full Time

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.

Sonic Summary

info
  • BCE Global Tech is seeking a Tier 1 Team Manager in Bengaluru to oversee a team of 14 AIOps Support Engineers, focusing on improving incident triage quality and speed.
  • The ideal candidate should have 7+ years of experience in application/production support and at least 2 years of experience managing a technical support team.
  • Key responsibilities include managing triage processes, driving data stewardship, and coordinating onboarding of new applications as the portfolio expands from 50 to 320 applications.
  • Candidates should be familiar with observability tools like Dynatrace and New Relic, and possess strong ITIL knowledge along with a technical background in Linux, networking, and SQL.
  • The position offers a competitive salary and comprehensive health benefits, along with flexible work hours and opportunities for professional development.

At BCE Global Tech we are on a mission to modernize globalconnectivity, one connection at a time. We aim to build the highway to thefuture of communications, media and entertainment, determined to emerge as apowerhouse within the technology landscape in India team in Bengaluru.

We bring ambitions to life through design thinking thatbridges the gaps between people, devices and beyond, fostering unprecedentedcustomer satisfaction through technology.

Our core values support a customer-centric approach and theharnessing of cutting-edge technology to provide business outcomes withpositive societal impact. Guided by innovation and a commitment to progress,we’re shaping a brighter future for the generations of today and tomorrow.

If you would like to be a part of a team of thought-leaderspioneering advancements in 5G, MEC, IoT and cloud-native architecture, we’dlove to hear from you.



Requirements

What You'll Do

•     Manage the Tier 1 team: directly manage ateam of 14 AIOps Support Engineers performing manual triage of alarms andalerts across a diverse, 50-to-320-application portfolio including hiring,coaching, scheduling, and performance management.

•     Own triage quality and speed: set and monitorstandards for how quickly and accurately the team detects, classifies, androutes incidents, and drive continuous improvement in mean-time-to-triage.

•     Drive data stewardship: partner withapplication teams to standardize alarm and alert data across heterogeneous logaggregation tools (Dynatrace, New Relic, ManageEngine, Glass box, and others)into a clean, consistent telemetry backbone built on Open Telemetry.

•     Manage the reactive-to-proactive shift: reduce reliance onreactive, manual triage over time by improving alert quality, correlation, andearly-warning signals laying the groundwork for future automated and Agentictriage.

•     Navigate a diverse, moving application landscape: support applicationsspanning different technology stacks and different architecture dispositions(Invest, Tolerate, Retire, Migrate), reprioritizing team focus as the portfolioshifts.

•     Coordinate onboarding of new apps: run a repeatableprocess for bringing new applications into Tier 1 coverage as the programscales from 50 to 320 applications, including support group and applicationowner mapping.

•     Manage stakeholders: act as the primary point of contactfor support groups, application owners, and AIOps program leadership on Tier 1status, incidents, and data-quality issues.

•     Manage shift/roster coverage: ensure the team of 14provides consistent triage coverage across required hours as the applicationcount grows.

•     Report on outcomes: track and report team KPIs triagetime, alert-to-incident accuracy, false-positive rates, coverage growth toprogram leadership.


 
What We're Looking For

•     Experience: 7+ years in application/production support(L1/L1.5/L2) or site reliability, with 2+ years directly managing a technicalsupport team.

•     Hybrid environment expertise: proven experiencesupporting applications across both on-premises and cloud environments, withexposure to modern microservices architectures.

•     Observability tooling: hands-on experience with monitoringand observability platforms such as Dynatrace, New Relic, AWS CloudWatch,ManageEngine, or Glass box; working knowledge of Open Telemetry and distributedtracing concepts.

•     ITIL discipline: strong grounding in incident, problem, andchange management practices, with ServiceNow or Jira ticket managementexperience.

•     Technical range: comfortable with Linux and Windowstroubleshooting, basic networking (TCP/IP, DNS, HTTP/HTTPS, SSL, loadbalancers), SQL/database query analysis, and API/integration troubleshooting.

•     People management: demonstrated ability to hire, coach,and retain a team of 10+ technical support staff through a period ofsignificant scale-up (5x application coverage growth).

•     Analytical mindset: able to turn noisy, inconsistent alertdata into clear, actionable insight, and to build repeatable frameworks ratherthan one-off fixes.

•     Comfort with ambiguity: willing to support amoving target  a portfolio spanningInvest, Tolerate, Retire, and Migrate applications and to adapt priorities as the programevolves.


Must-Have Skills

•     IncidentManagement Lifecycle, and working knowledge of Problem, Change Request, andService Request concepts (ITIL)

•     CMDBconcepts and their use in incident and asset traceability

•     Hands-onexperience with log aggregation technologies (Dynatrace, New Relic,ManageEngine, Glass box, or similar)

•     Workingknowledge of JSON and XML, and basic file/task automation

•     Understandingof IT infrastructure and basic networking: VMs, firewalls, load balancers,containers, OpenShift (OCP), Kubernetes

•     UnixShell scripting; Windows batch file creation

•     Basiccloud concepts: compute, storage, and security fundamentals

•     Securityfundamentals: TLS, SSL, tokens, and secret management

•     Familiaritywith API gateways and API testing toolkits (Postman, SOAP UI, or similar)

•     Outageresponse management and experience leading cross-functional coordination duringmajor incidents

•     Abilityto drive Root Cause Analyses (RCAs) and build reusable knowledgearticles/runbooks

•     SLA/SLOmanagement and reporting, including availability calculation

•     Workingknowledge of data concepts: data latency, data fragmentation, data lineage, anddata marts


Nice-to-Have Skills

•     Familiaritywith AI concepts such as prompt engineering, knowledge graphs, andRetrieval-Augmented Generation (RAG)

•     Experiencewith cloud-native observability on AWS, Azure, or GCP

•     Exposureto Agentic AI or automation-driven triage tooling


 Success Looks Like

•     A14-person Tier 1 team that reliably triages alerts across 50+ applications withclear, standardized data.

•     Ameasurable, ongoing reduction in reactive manual triage as proactive detectionimproves.

•     Aclean, well-governed Open Telemetry-based data backbone that the future Agenticautomation layer can build on.

Arepeatable onboarding process ready to scale coverage from 50 to 320 applications.



Benefits

What We Offer

Competitive salaries and comprehensive health benefits

Flexible work hours and remote work options.

Professional development and training opportunities.

A supportive and inclusive work environment






Learn more about this Employer on their Career Site

Apply now in a few quick clicks

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.