Site Reliability Engineering Lead
Required skills for this role
Site Reliability Engineering(SRE)PythonAWSAzureGCPDockerKubernetesTerraformDevOpsCI/CDLinuxUnix
About this role
Roles & Responsibilities
- Generating response
- Copilot said:
- Site Reliability Engineer (SRE) Roles & Responsibilities
- Monitor application and infrastructure health 24/7.
- Ensure high availability, reliability, and performance of systems.
- Respond to incidents, outages, and production issues.
- Perform root cause analysis (RCA) and implement permanent fixes.
- Automate repetitive operational and deployment tasks.
- Develop and maintain monitoring, alerting, and logging solutions.
- Configure and manage CI/CD pipelines.
- Support application deployments and release management.
- Manage cloud infrastructure (Azure, AWS, or GCP).
- Perform capacity planning and resource optimization.
- Implement Infrastructure as Code (IaC) using tools like Terraform or ARM templates.
- Collaborate with development teams to improve application reliability.
- Define and track SLAs, SLIs, and SLOs.
- Manage backups, disaster recovery, and business continuity plans.
- Ensure security compliance and vulnerability remediation.
- Troubleshoot Linux, networking, database, and application issues.
- Maintain Kubernetes and containerized environments.
- Create operational documentation and runbooks.
- Conduct post-incident reviews and recommend improvements.
- Continuously improve system scalability, efficiency, and stability.
- Key Skills
- Linux/Unix Administration
- Azure/AWS/GCP
- Kubernetes & Docker
- Python, Bash, PowerShell
- Terraform
- Jenkins, Azure DevOps, GitHub Actions
- Prometheus, Grafana, Splunk, Azure Monitor
- Networking & Security
- Incident Management and Problem Solving
How to apply
Use the official application link to contact the employer and submit your application.