TechOps-CloudOps-Monitoring Lead-Senior-GDSN02
About this role
The Opportunity
We are seeking a Senior Service Delivery & Operations Lead to lead enterprise-scale Cloud Operations, Application Production Support, Monitoring, Service Management, and Operational Excellence across Azure-based Data & Analytics platforms. This role is accountable for ensuring operational stability, service reliability, customer satisfaction, and governance through effective management of global 24x7 operations.
The ideal candidate will bring strong experience in Cloud Operations, Application Support, Incident Management, Service Delivery, and Project Execution while partnering with engineering, business, and delivery teams to drive operational excellence, service modernization, automation, and continuous improvement.
Your Key Responsibilities
- Lead end-to-end service delivery, operational governance, and support operations across cloud, enterprise application, and Data & Analytics platforms, ensuring high availability, reliability, and customer satisfaction.
- Manage and mature global 24x7 support organizations, including workforce planning, shift governance, service quality management, operational reporting, resource allocation, performance management, and team development.
- Own service performance through effective governance of SLAs, KPIs, compliance requirements, operational reviews, executive reporting, and stakeholder communication forums.
- Lead Incident, Problem, Change, Major Incident, and Service Request Management processes aligned with ITIL best practices, serving as the primary escalation point for critical service disruptions and ensuring timely service restoration, root cause analysis, and implementation of corrective actions.
- Oversee L1.5 production support, monitoring, and operational management of Azure-based Data & Analytics applications, batch data pipelines, cloud services, and enterprise platforms, ensuring proactive issue detection, operational stability, and data availability.
- Drive observability and operational excellence initiatives through continual enhancement of monitoring, alerting, logging, dashboards, automation, and self-healing capabilities using platforms such as Azure Monitor, Log Analytics, Power BI, Grafana, and related technologies.
- Partner closely with L2/L3 support teams, Data Engineering, Architecture, DevOps, and delivery organizations to improve service reliability, platform performance, operational resilience, and continuous service improvement.
- Lead service transition, operational readiness, and support model design for new applications, cloud migrations, platform modernizations, releases, upgrades, and transformation initiatives.
- Maintain and enhance runbooks, SOPs, knowledge repositories, governance frameworks, and operational processes to ensure standardized, scalable, and efficient support operations.
- Drive continuous improvement initiatives focused on automation, reduction of alert noise, process optimization, operational efficiency, and overall service maturity.
- Provide strong leadership through mentoring, coaching, and capability development of support teams while fostering a culture of accountability, collaboration, innovation, and operational excellence.
- Support project delivery through resource forecasting, budgeting, risk management, capacity planning, and coordination with cross-functional stakeholders to ensure successful execution of operational and transformation program.
- Responsible for decision-making, optimizing processes, resource management, and overseeing team management as needed for task execution.
- Accountable for allocating personnel, supervising team members, assigning tasks, ensuring that the team has the necessary tools and support to succeed in their roles and optimizing and evaluating their performance to meet organizational goals.
- Responsible for decision-making, optimizing processes, resource management, and overseeing team management as needed for task execution.
- Accountable for allocating personnel, supervising team members, assigning tasks, ensuring that the team has the necessary tools and support to succeed in their roles and optimizing and evaluating their performance to meet organizational goals.
Skills and attributes for success
- Strong experience leading Managed Services, Service Delivery, Cloud Operations, and Application Support organizations, with proven success managing large-scale 24x7 operations in global delivery environments.
- Demonstrated leadership in operational governance, incident management, service resilience, and continuous service improvement, ensuring SLA compliance, operational excellence, and business continuity.
- Proven ability to lead distributed and shift-based teams, including workforce planning, capacity management, stakeholder engagement, and executive-level communication.
- Strong understanding of monitoring, observability, automation, AIOps, and operational support frameworks, with the ability to drive proactive issue detection, rapid incident resolution, and service optimization.
- Strong project management, delivery governance, and problem-solving capabilities, with a focus on prioritization, risk management, process adherence, and driving successful business outcomes.
To qualify for the role, you must have
- 8-12+ years of experience in Service Delivery, IT Operations, Cloud Operations, Application Support, Managed Services, or Production Support.
- Minimum 3-5 years of experience leading global 24x7 operations and support teams.
- Experience managing enterprise-wide Incident, Problem, Change, and Major Incident Management processes.
- Experience leading service transition and operational readiness initiatives.
- Experience managing global stakeholders and executive communications.
- Experience with operational governance, KPI management, SLA reviews, and service reporting.
- Experience working in complex enterprise environments supporting business-critical platforms.
Technologies and Tools
Must have
- Strong experience in Cloud Operations, Application Production Support, and Managed Services, with hands-on exposure to Azure cloud environments; AWS/GCP exposure is an added advantage.
- Proven experience managing 24x7 global support operations, including workforce planning, capacity management, shift governance, and service continuity.
- Strong expertise in Incident, Problem, Change, Release, and Major Incident Management, including leading L1.5 triage calls and coordinating resolution across L2/L3 and engineering teams.
- Experience working with ITSM platforms and operational governance frameworks, including SLA/KPI management, service reporting, escalation management, and continuous service improvement.
- Strong understanding of Monitoring, Observability, Logging, Application Performance Management, API and Integration Support, with the ability to drive proactive issue detection and operational stability.
- Proven leadership experience in managing support teams, driving operational excellence, stakeholder communications, resource planning, and service delivery outcomes in enterprise environments.
- Certifications such as
- ITIL v4 Foundation / Managing Professional
- Microsoft Azure Administrator Associate
Good to have
- PMP (Project Management Professional)
- PRINCE2 Practitioner
- Agile / Scrum Certification
- Site Reliability Engineering (SRE) Certification
- Cloud Operations Certifications</