Site Reliability Engineering Lead_Truist
Required skills for this role
KubernetesAnsibleSite Reliability Engineering(SRE)Python - OpenSystemPythonDevOpsLinuxUnix
About this role
The Site Reliability Engineering Lead is a senior, hands-on technical leader within the Corporate Technology and Operations organization. This teammate is accountable for elevating the reliability, resiliency, and operational excellence of critical enterprise platforms across hybrid cloud and onprem environments.
Roles & Responsibilities
- Reliability Engineering & Automation
- Architect and deliver automation solutions that eliminate toil, reduce MTTR, and increase service resilience. Experience in Ansible, Puppet or Chef is a plus.
- Implement intelligent alerting, anomaly detection, and event correlation leveraging AI and AIOps tools.
- Guide and enforce SLO/SLI adoption across product teams, ensuring metrics inform decision-making and prioritization.
- Utilize Infrastructure-as-Code (IaC) tools for automating deployment of assets within cloud tenants.
- Observability & Operational Excellence
- Ensure operational readiness of applications and platforms through resiliency testing, chaos engineering, and failure-mode validation.
- Cross-Functional Leadership & Influence
- Partner with Delivery, Architecture, Security, and Risk teams to embed reliability and resilience into design and execution.
- Standardization & Documentation
- Develop, maintain, and enforce runbooks, response playbooks, and automated recovery patterns.
- Follow best practices and internal processes for Non-Functional requirements to improve resiliency and reliability.
- Mentorship & Technical Development
- Coach and mentor Associate, Professional, and Senior SREs to build technical depth and operational discipline.
- Provide thought leadership in SRE methodologies, cloud-native operational patterns, and automated reliability engineering.
- Incident Leadership & Production Operations
- Lead P1/P0 incident bridges and direct technical investigation efforts.
- Perform hands-on triage using logs, traces, metrics, and application telemetry.
- Drive mitigation, recovery, RCA development, and follow-through remediation.
- Provide executive communications during major incidents.
- Build operational automation based on recurring production issues.
- Establish credibility through technical leadership during live service disruptions.
Technical Requirements
- 7+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Infrastructure Operations.
- Deep hands‑on experience with distributed systems, container orchestration (Kubernetes), and cloud-native operational tooling.
- Proficiency with automation and scripting languages (Python, Go, PowerShell, Ansible).
- Strong understanding of observability platforms (Splunk, Dynatrace) and event-driven monitoring.
- Proven leadership in major incident management and cross-team technical coordination.
- Strong grasp of networking, Linux/Unix internals, and modern infrastructure patterns.
- Excellent communication skills, including executive-level situational awareness during critical incidents.
- Demonstrated ability to influence technical roadmaps and drive adoption of reliability best practices.
Additional
- Experience enabling large-scale SRE transformations or modernization initiatives.
- Demonstrated proficiency with GitLab Duo, or similar AI technologies.
- Familiarity with chaos engineering, resilience assessments, and service failure modeling.
- Exposure to hybrid-cloud and multi-cloud operational frameworks.
- Experience contributing to or leading Center for Enablement functions or Communities of Practice.
- Expertise with highly regulated industries preferred.