Site Reliability Engineering (SRE) Foundation: Building Resilient and Scalable Systems
Join our Observability Foundation course to master data-driven insights and system performance monitoring. Enhance your IT skills with advanced observability techniques.
Course Fee
S$2000
Up to 70% funding available
Course Information
Why Choose Garranto Academy for Your Site Reliability Engineering (SRE) Foundation: Building Resilient and Scalable Systems Training?
Choose Garranto Academy for your Site Reliability Engineering (SRE) Foundation: Building Resilient and Scalable Systems training and benefit from expert-led instruction, hands-on exercises, and real-world case studies tailored to today's industry demands. Our curriculum is designed by practitioners with deep Ops & IT Operations expertise, ensuring you gain immediately applicable skills backed by globally recognised frameworks.
Course Overview
Join our Observability Foundation course to master data-driven insights and system performance monitoring. Enhance your IT skills with advanced observability techniques.
Course Objectives
- Understand IT operations frameworks including ITIL, SRE, and DevOps.
- Manage IT services, incidents, problems, and changes effectively.
- Apply monitoring, observability, and alerting best practices.
- Implement automation to improve operational efficiency and reliability.
- Build resilient services with high availability and disaster recovery.
- Measure and improve IT service quality using SLAs, SLOs, and SLIs.
Prerequisites
- Basic familiarity with IT systems, networks, or software development.
- Experience in an IT support or operations role is beneficial.
- No prior ITIL or SRE certification required.
Course Outlines
Module 1: IT Operations Fundamentals
- IT service management principles and ITIL 4 overview
- Key IT operations roles: SRE, NOC, and DevOps engineers
- Service catalogue, SLAs, and on-call management
Module 2: Incident & Problem Management
- Incident classification, escalation, and resolution
- Root cause analysis (RCA) and post-incident reviews
- Problem management: proactive vs reactive approaches
Module 3: Monitoring & Observability
- Monitoring architecture: metrics, logs, and traces
- Key tools: Prometheus, Grafana, Datadog, ELK stack
- Building alerting strategies and reducing alert fatigue
Module 4: Change & Release Management
- Change advisory board (CAB) and change types
- Release planning and deployment pipeline management
- Rollback strategies and change risk assessment
Module 5: Automation & Reliability Engineering
- Infrastructure as code (IaC) with Terraform/Ansible
- Chaos engineering and reliability testing
- Continuous improvement with error budgets and retrospectives
Course Outcomes
- Manage IT services, incidents, and changes using proven frameworks.
- Implement effective monitoring and observability for IT systems.
- Reduce mean time to detect (MTTD) and mean time to resolve (MTTR).
- Apply automation to improve IT operational efficiency.
- Build and maintain highly available, resilient IT services.
Benefits of Site Reliability Engineering (SRE) Foundation: Building Resilient and Scalable Systems
Completing the Site Reliability Engineering (SRE) Foundation: Building Resilient and Scalable Systems programme at Garranto Academy equips you with in-demand competencies, enhances your professional credibility, and accelerates career growth. Participants gain practical experience, access to a peer network of professionals, and the confidence to apply new knowledge directly in their organisations.
What You'll Learn
Facilities & Equipment
Virtual Training
- Electronic materials
- IT support for software & hardware
- Administrative support
Face-to-Face Training
- Air-conditioned classroom
- Meals & refreshments provided
- Projector & smart board
- Stationery provided
What Learners Say
Real experiences, real results
Share Your Feedback
Attended this course? Log in to let other learners know what to expect.
