When Things Break: Crashes, Traffic Spikes & Leaked Keys
Real apps break: a deploy crashes the site, 1,000 people arrive at once, an API key ends up on GitHub. Learn the calm, professional playbook — triage, read the logs, roll back, fix the root cause, handle traffic spikes, respond to a leaked secret in minutes, and write a blameless post-mortem so it never happens again.
8
units
20
lessons
72
challenges
~70
minutes
Start learning today
Create a free account, then unlock this and every other Pro course with Learnix Pro for KES 999/month.
Create free account I already have an accountOutcomes
What you'll be able to do
Triage an incident: severity, communication and first steps
Read logs and error messages to find the real cause
Recover from a crash: roll back, fix forward, verify
Keep an app alive under 1,000+ requests: caching, queues, rate limits, scaling
Respond to an exposed API key: revoke, rotate, audit, prevent
Write a blameless post-mortem and set up monitoring and backups
How it works
Learn by doing, a few minutes a day
Every lesson mixes short explanations with hands-on challenges. Get instant feedback, fix mistakes on the spot, and build a daily habit that actually sticks.
Concept cards
Short, clear explanations with real examples
58 in this course
Multiple choice
Pick the right answer, get instant feedback
32 in this course
Fill the blank
Complete formulas, commands and shortcuts
14 in this course
Build it
Arrange tiles into the right sequence
11 in this course
Match pairs
Connect shortcuts, terms and meanings
15 in this course
Daily streaks
Keep the flame alive by learning every day
XP & daily goals
Pick a goal from 10 to 50 XP a day
Hearts
Five per lesson; missed questions come back
Leaderboard
Climb the weekly rankings
Syllabus
Everything you'll encounter
What is an incident?
Detect, triage, communicate — then fix.
Reading logs and errors
The error message usually tells you where to look.
The first 15 minutes of a crash
Is it everyone? What changed? Roll back.
Common crash causes and fixes
Most crashes come from the same few things.
When a mobile app crashes
Crash reports, hotfixes and staged rollouts.
What happens under heavy traffic
Find the bottleneck before it finds you.
Handling 1,000+ requests
Cache, queue, index, limit, scale.
Real traffic or an attack?
Rate limits, WAFs and DDoS basics.
Your API key leaked: the first hour
Revoke, rotate, check the damage.
Never leak a secret again
.env, .gitignore, scanning and least privilege.
The database is slow or down
Most "the site is slow" incidents are really database incidents.
A bad deploy: roll back first
Your new release broke production. Restore service, then investigate.
The blameless post-mortem
Fix the system, not the person.
Monitoring, backups and prevention
Find out before your users do.
Disk full & expired SSL certificates
Two silent killers with simple fixes.
Emails not sending & stuck queues
Background jobs that silently stop.
Payments failing: M-Pesa callbacks
Customers paid, but your system says unpaid.
Communicating during an outage
What to tell users, the boss and the team.
Website or account hacked
Defacement, malware, spam and stolen admin accounts.
Write your own runbook
A one-page plan so the next incident is calm.
Finish the course, earn a verifiable certificate
Complete every lesson, then pass the 20-question final exam (70% to pass, unlimited retakes), to receive a Learnix certificate with a unique number anyone can verify online.