
August 24, 2026
13 min read
By Kokil Thapa | Last reviewed: September 2026
A production Laravel application rarely fails at a convenient hour. When it does fail, the gap between a five-minute fix and a four-hour outage is almost always documentation. An On-Call and Incident Response Runbook converts panic into procedure by putting exact commands, decision trees, and escalation paths in front of the engineer before the alert fires. Without that structure, minutes disappear while someone guesses server paths or scrolls Slack history. I have maintained production systems since 2010, and the pattern holds across every stack I have run.
Whether you run a high-traffic WooCommerce store or a legal-tech portal, fatigue erases mental notes during an emergency. The runbook bridges "I think I fixed this last time" and "here is the exact command that restores service." This guide covers building one for PHP and Laravel stacks, grounded in real deployment workflows using Deployer 7, GitLab CI, and Ubuntu servers. The same discipline applies to Laravel applications built for Nepal-based clients and to distributed teams running systems worldwide. Treat it as part of production support and maintenance, not as an optional extra.
What should an On-Call and Incident Response Runbook contain?
A working runbook is not a generic wiki page. It is an executable checklist designed for high-stress conditions. Every entry has to pass the "3 AM test": can a tired engineer run this safely without opening three other documents? For Laravel 13.x on PHP 8.3 or higher, five components are non-negotiable.
- Alert definition and severity matrix: Separate P1 events (site down, payment gateway failing) from P3 events (admin dashboard slow). Define which alerts wake someone up and which wait for morning triage.
- Immediate triage commands: Copy-pasteable checks for system health. Do not link out to another doc; embed the actual
artisanorsystemctlcommand. - Architecture context: A simplified diagram showing how Nginx, PHP-FPM, Redis, and MySQL interact. When queues back up, the engineer needs to know whether the bottleneck is worker count or the database connection limit.
- Rollback and mitigation procedures: Exact steps to revert the last deploy or disable a feature flag. Restoring service always outranks root cause analysis.
- Escalation contacts and access: Verified phone numbers for senior engineers, hosting providers, and API vendors. Link to the credential vault, never paste secrets into the document.
Teams running multiple sites on shared infrastructure need one more section: environment isolation. Confusing staging with production during an incident is a catastrophic but preventable error. Prefix every destructive command with an explicit environment check. On the sister sites I maintain — notarykathmandu.com, translationnepal.com and similar properties sharing one Deployer 7 pipeline — that check is the first line of every procedure.
How do you write runbook commands for Laravel production debugging?
The most common failure mode is vagueness. "Check the logs" is useless when the engineer does not know whether you use daily rotation, a single file, or syslog. On Laravel 13.x running Ubuntu 24.04 with PHP-FPM, precision saves minutes. Every command in your On-Call and Incident Response Runbook must be tested against the current production environment, not a local Valet or Sail setup.
Essential diagnostic commands
Embed these exact commands, adjusted for your Deployer release paths. On a standard zero-downtime Deployer release, the active symlink is typically /var/www/project/current.
# Verify current release path and deployment timestamp
ls -la /var/www/project/current
cat /var/www/project/current/.env | grep APP_ENV
# Check Laravel health endpoint if configured
curl -s https://example.com/up
# Tail application logs with context
tail -n 100 /var/www/project/current/storage/logs/laravel.log | grep -i "error\|exception"
# Inspect PHP-FPM pool status
systemctl status php8.4-fpm
journalctl -u php8.4-fpm --since "10 minutes ago" --no-pager
# Check queue worker health and stuck jobs
php /var/www/project/current/artisan queue:monitor --timeout=60
redis-cli llen queues:default
# Validate database connectivity without running migrations
php /var/www/project/current/artisan db:show --counts=users,orders A frequent mistake I find during CI/CD pipeline audits is stale runbook commands pointing at an old PHP version. If you upgraded from 8.3 to 8.4 last month but the runbook still says php8.3-fpm, the on-call engineer wastes minutes on a service that no longer exists. Version pinning in documentation matters as much as version pinning in composer.json. This is also where Linux system administration discipline shows up: know your PHP binary paths before you need them.
Queue and cache emergency procedures
Laravel queues are usually the first component to degrade under load. Your runbook needs restart procedures that respect graceful shutdown. Killing workers outright can corrupt job payloads.
# Gracefully restart all queue workers after a deploy
php /var/www/project/current/artisan queue:restart
# Force-stop stuck workers (last resort only)
pkill -f "queue:work"
supervisorctl restart laravel-worker:*
# Clear application caches without dropping sessions
php /var/www/project/current/artisan config:clear
php /var/www/project/current/artisan route:clear
php /var/www/project/current/artisan view:clear
# Check Redis memory pressure
redis-cli info memory | grep used_memory_human Document the expected output of healthy commands. If redis-cli llen queues:default normally returns 0–50 but currently reads 15,000, that is your smoking gun. Without baseline expectations, metrics are just noise. The Laravel queue documentation covers worker signals and timeouts in detail. It is worth reading before an outage, not during one. Background job behaviour is covered further in Laravel queues and background jobs.
How should incident severity levels be defined for PHP web applications?
Severity definitions must follow business impact, not technical symptoms. A 500 error on an admin report page is technically identical to a 500 error on the checkout endpoint, yet operationally they are worlds apart. For eCommerce clients running WooCommerce or a custom Laravel cart, I define severity by revenue exposure and user-facing functionality.
| Severity | Business impact | Response time | Example scenarios |
|---|---|---|---|
| P1 — Critical | Revenue loss, data breach, complete outage | Immediate, 24/7 wake-up | Payment gateway timeout, database corruption, expired SSL certificate, homepage 500s |
| P2 — High | Major feature broken, significant user friction | Under 1 hour in business hours | Search broken, email notifications failing, admin panel unreachable, page loads over 5s |
| P3 — Medium | Minor feature issue with a workaround | Under 4 hours or next business day | Image upload fails for one format, CSV export misaligned, non-critical webhook delays |
| P4 — Low | Cosmetic, internal tooling, no user impact | Scheduled maintenance window | Typos in admin UI, deprecation warnings, log verbosity changes |
This matrix prevents alert fatigue. If every notification triggers a P1 response, engineers stop responding to real emergencies. For Nepal-based businesses on tighter budgets, defining P2 against P3 also sets expectations about after-hours support costs. A client paying Rs 15,000/month (~USD 110) for maintenance should understand that cosmetic fixes wait until Monday, while payment failures get immediate attention. The same logic underpins meaningful SLIs and SLOs and the alerting rules that follow from them.
How do you handle incident communication and stakeholder updates?
Technical resolution is only half the job. Stakeholders do not care about your strace output. They want to know when service returns and what to tell customers. Pre-written templates remove the cognitive load of drafting messages while debugging. In my work on legal-tech portals, where trust is the product, transparent communication during outages preserves relationships far better than silence.
Status update templates
Keep these in the runbook with placeholders clearly marked. Adjust tone for the audience — internal teams get technical detail, clients get business impact.
## INITIAL ACKNOWLEDGMENT (within 15 min of P1)
Subject: [INCIDENT] Service degradation — {SERVICE_NAME}
Status: Investigating
Impact: {USER_FACING_SYMPTOM}
Started: {TIMESTAMP_NPT}
Next update: {TIMESTAMP + 30 MIN}
We are aware of {SYMPTOM} affecting {SCOPE}.
Engineering is actively investigating.
No ETA yet — we will update in 30 minutes.
## PROGRESS UPDATE (every 30-60 min)
Status: {INVESTIGATING | IDENTIFIED | MITIGATING | RECOVERING}
Current action: {WHAT_WE_ARE_DOING_NOW}
ETA: {ESTIMATE_OR_UNKNOWN}
Workaround: {IF_AVAILABLE}
## RESOLUTION NOTICE
Status: Resolved
Duration: {START_TIME} to {END_TIME} ({TOTAL_MINUTES} min)
Root cause: {ONE_SENTENCE_SUMMARY}
Prevention: {FOLLOW_UP_ACTION_PLANNED} For Nepal-based operations, consider bilingual updates when the customer base includes non-English speakers. Also account for local business hours and festivals. An incident during Dashain needs a different communication cadence than a regular Tuesday. Tie these templates into your server security and monitoring documentation so the on-call engineer has full context without hunting. The Google SRE book chapter on postmortem culture is a solid reference for the tone these updates should carry.
What belongs in a post-incident review for Laravel applications?
The post-incident review is where organisational learning happens. Without it, you fix the same bug three times in six months. A review is not a blame session. It is a structured analysis of systemic failure, and for Laravel applications it should target concrete improvements to code, configuration, or process.
Review document structure
- Timeline: Minute-by-minute account from first alert to full recovery. Pull timestamps from monitoring tools, not memory.
- Impact quantification: Orders lost, users affected, revenue estimate. "23 failed checkout attempts during the Valentine's peak" beats "checkout was down."
- Root cause analysis: Use the Five Whys. "Queue failed" → why? → "Redis ran out of memory" → why? → "Memory limit too low" → why? → "Never tuned after traffic doubled" → why? → "No capacity review process."
- Action items: Concrete tasks with owners and deadlines. "Raise Redis maxmemory to 2GB" beats "monitor Redis better." Link every item to a ticket.
- Runbook updates: What was missing or wrong in the current document? This section closes the loop.
Schedule reviews within 48 hours of any P1. For distributed teams, record the session for async review. The point is to update the On-Call and Incident Response Runbook so the next incident is faster, cheaper, and less stressful. Pair this with the patterns in blameless postmortems that actually help and the wider incident response playbook.
Key Takeaways
- Write each runbook entry so a tired engineer can run it at 3 AM without opening another document.
- Pin exact PHP and service names —
php8.4-fpm, notphp-fpm— and update them with every upgrade. - Define severity by business impact, then publish the matrix so alerting rules match it.
- Keep pre-written stakeholder templates in the runbook, with placeholders and an NPT timestamp format.
- Hold a post-incident review within 48 hours and treat every finding as a runbook edit.
- Test triage commands during quiet periods, not during the first real outage.
People Also Ask
What is the difference between a runbook and a playbook?
A runbook holds the operational detail for a specific alert: exact commands, expected output, and known fixes. A playbook describes the wider process — who leads, how you communicate, when you escalate. The runbook is what you execute at 3 AM. The playbook is what your organisation follows around it. Most teams need both, and the runbook is the one that saves minutes.
How often should an incident response runbook be updated?
Update it whenever infrastructure changes and after every P1 or P2 incident. A PHP version bump, a new queue driver, a hosting migration, or a renamed release path all invalidate commands. If your team has not touched the runbook in six months, assume parts of it are wrong. Stale commands are worse than no commands, because they cost time and confidence.
Who should own the on-call runbook?
The engineer who carries the pager owns the content, even if someone else edits the file. Ownership means testing commands, chasing missing escalation contacts, and closing action items from reviews. In small teams that is usually the senior developer. In larger teams it rotates with the on-call schedule. Either way, one name sits against each section.
What is a realistic MTTR for a Laravel application?
For a well-instrumented Laravel app with a maintained runbook, 15–30 minutes is achievable for common failures like queue backlogs, cache issues, or a bad deploy. First-time incidents in unfamiliar areas take longer, which is exactly why the runbook exists. Track MTTR per incident type rather than as a single average. That tells you where documentation is still thin.
Build the Runbook Before the Next Outage
An On-Call and Incident Response Runbook is not a document you write once and archive. It evolves with your application, your team, and your infrastructure. Start small this week: pick your three most frequent alerts and document them properly. Test every command in a quiet window. Hand the draft to whoever takes the next on-call shift and ask what is missing. Add the monitoring layer that feeds it — a Prometheus and Grafana monitoring stack with SLO-driven alerting and a proper on-call alerting tool — so the runbook starts from a real signal.
If you are weighing the cost of downtime against better tooling or a larger VPS, run the numbers with the Nepal EMI calculator before committing budget. For teams that want the full picture, the SRE vs DevOps roles comparison and how to write effective runbooks cover the surrounding practice. You can see the kind of portals this discipline supports in the secure client portal with document sharing and the legal-guide site with lead capture.
If your production environment has no structured incident response, or your runbook is outdated and untested, get in touch through the contact page. I help teams build practical, maintainable operational documentation grounded in real production work — not theoretical frameworks that collapse under pressure. Whether you need a runbook audit, deployment pipeline hardening, or hands-on incident response training for your developers, we can scope an engagement that fits your operational reality and budget.
Frequently Asked Questions
0 Comments
Leave a comment
Your email is not published. Comments appear once they have been read. Sign in to have your details filled in.

