Kokil Thapa - Professional Web Developer in Nepal
Freelancer Web Developer in Nepal with 15+ Years of Experience

Kokil Thapa is an experienced full-stack web developer focused on building fast, secure, and scalable web applications. He helps businesses and individuals create SEO-friendly, user-focused digital platforms designed for long-term growth.

On-Call and Incident Response Runbook

By Kokil Thapa | Last reviewed: September 2026

A production Laravel application rarely fails at a convenient hour. When it does fail, the gap between a five-minute fix and a four-hour outage is almost always documentation. An On-Call and Incident Response Runbook converts panic into procedure by putting exact commands, decision trees, and escalation paths in front of the engineer before the alert fires. Without that structure, minutes disappear while someone guesses server paths or scrolls Slack history. I have maintained production systems since 2010, and the pattern holds across every stack I have run.

Whether you run a high-traffic WooCommerce store or a legal-tech portal, fatigue erases mental notes during an emergency. The runbook bridges "I think I fixed this last time" and "here is the exact command that restores service." This guide covers building one for PHP and Laravel stacks, grounded in real deployment workflows using Deployer 7, GitLab CI, and Ubuntu servers. The same discipline applies to Laravel applications built for Nepal-based clients and to distributed teams running systems worldwide. Treat it as part of production support and maintenance, not as an optional extra.

What should an On-Call and Incident Response Runbook contain?

A working runbook is not a generic wiki page. It is an executable checklist designed for high-stress conditions. Every entry has to pass the "3 AM test": can a tired engineer run this safely without opening three other documents? For Laravel 13.x on PHP 8.3 or higher, five components are non-negotiable.

  1. Alert definition and severity matrix: Separate P1 events (site down, payment gateway failing) from P3 events (admin dashboard slow). Define which alerts wake someone up and which wait for morning triage.
  2. Immediate triage commands: Copy-pasteable checks for system health. Do not link out to another doc; embed the actual artisan or systemctl command.
  3. Architecture context: A simplified diagram showing how Nginx, PHP-FPM, Redis, and MySQL interact. When queues back up, the engineer needs to know whether the bottleneck is worker count or the database connection limit.
  4. Rollback and mitigation procedures: Exact steps to revert the last deploy or disable a feature flag. Restoring service always outranks root cause analysis.
  5. Escalation contacts and access: Verified phone numbers for senior engineers, hosting providers, and API vendors. Link to the credential vault, never paste secrets into the document.

Teams running multiple sites on shared infrastructure need one more section: environment isolation. Confusing staging with production during an incident is a catastrophic but preventable error. Prefix every destructive command with an explicit environment check. On the sister sites I maintain — notarykathmandu.com, translationnepal.com and similar properties sharing one Deployer 7 pipeline — that check is the first line of every procedure.

Incident LifecycleAlert FiresP1 or P2 pageTriageLogs and metricsMitigateRollback, restartVerifyConfirm recoveryPost-Incident Review Updates the RunbookEvery incident feeds the next response
Incident lifecycle in an On-Call and Incident Response Runbook: alert, triage, mitigation, verification, review

How do you write runbook commands for Laravel production debugging?

The most common failure mode is vagueness. "Check the logs" is useless when the engineer does not know whether you use daily rotation, a single file, or syslog. On Laravel 13.x running Ubuntu 24.04 with PHP-FPM, precision saves minutes. Every command in your On-Call and Incident Response Runbook must be tested against the current production environment, not a local Valet or Sail setup.

Essential diagnostic commands

Embed these exact commands, adjusted for your Deployer release paths. On a standard zero-downtime Deployer release, the active symlink is typically /var/www/project/current.

# Verify current release path and deployment timestamp
ls -la /var/www/project/current
cat /var/www/project/current/.env | grep APP_ENV

# Check Laravel health endpoint if configured
curl -s https://example.com/up

# Tail application logs with context
tail -n 100 /var/www/project/current/storage/logs/laravel.log | grep -i "error\|exception"

# Inspect PHP-FPM pool status
systemctl status php8.4-fpm
journalctl -u php8.4-fpm --since "10 minutes ago" --no-pager

# Check queue worker health and stuck jobs
php /var/www/project/current/artisan queue:monitor --timeout=60
redis-cli llen queues:default

# Validate database connectivity without running migrations
php /var/www/project/current/artisan db:show --counts=users,orders

A frequent mistake I find during CI/CD pipeline audits is stale runbook commands pointing at an old PHP version. If you upgraded from 8.3 to 8.4 last month but the runbook still says php8.3-fpm, the on-call engineer wastes minutes on a service that no longer exists. Version pinning in documentation matters as much as version pinning in composer.json. This is also where Linux system administration discipline shows up: know your PHP binary paths before you need them.

Queue and cache emergency procedures

Laravel queues are usually the first component to degrade under load. Your runbook needs restart procedures that respect graceful shutdown. Killing workers outright can corrupt job payloads.

# Gracefully restart all queue workers after a deploy
php /var/www/project/current/artisan queue:restart

# Force-stop stuck workers (last resort only)
pkill -f "queue:work"
supervisorctl restart laravel-worker:*

# Clear application caches without dropping sessions
php /var/www/project/current/artisan config:clear
php /var/www/project/current/artisan route:clear
php /var/www/project/current/artisan view:clear

# Check Redis memory pressure
redis-cli info memory | grep used_memory_human

Document the expected output of healthy commands. If redis-cli llen queues:default normally returns 0–50 but currently reads 15,000, that is your smoking gun. Without baseline expectations, metrics are just noise. The Laravel queue documentation covers worker signals and timeouts in detail. It is worth reading before an outage, not during one. Background job behaviour is covered further in Laravel queues and background jobs.

How should incident severity levels be defined for PHP web applications?

Severity definitions must follow business impact, not technical symptoms. A 500 error on an admin report page is technically identical to a 500 error on the checkout endpoint, yet operationally they are worlds apart. For eCommerce clients running WooCommerce or a custom Laravel cart, I define severity by revenue exposure and user-facing functionality.

SeverityBusiness impactResponse timeExample scenarios
P1 — CriticalRevenue loss, data breach, complete outageImmediate, 24/7 wake-upPayment gateway timeout, database corruption, expired SSL certificate, homepage 500s
P2 — HighMajor feature broken, significant user frictionUnder 1 hour in business hoursSearch broken, email notifications failing, admin panel unreachable, page loads over 5s
P3 — MediumMinor feature issue with a workaroundUnder 4 hours or next business dayImage upload fails for one format, CSV export misaligned, non-critical webhook delays
P4 — LowCosmetic, internal tooling, no user impactScheduled maintenance windowTypos in admin UI, deprecation warnings, log verbosity changes

This matrix prevents alert fatigue. If every notification triggers a P1 response, engineers stop responding to real emergencies. For Nepal-based businesses on tighter budgets, defining P2 against P3 also sets expectations about after-hours support costs. A client paying Rs 15,000/month (~USD 110) for maintenance should understand that cosmetic fixes wait until Monday, while payment failures get immediate attention. The same logic underpins meaningful SLIs and SLOs and the alerting rules that follow from them.

Severity Decision TreeAlert ReceivedIs revenue ordata at risk?YESNOP1 CriticalPage the engineerFeature broken?Core workflow downP2 HighRespond in 1 hourP3 / P4Next business day
Decision tree for classifying incident severity in an On-Call and Incident Response Runbook

How do you handle incident communication and stakeholder updates?

Technical resolution is only half the job. Stakeholders do not care about your strace output. They want to know when service returns and what to tell customers. Pre-written templates remove the cognitive load of drafting messages while debugging. In my work on legal-tech portals, where trust is the product, transparent communication during outages preserves relationships far better than silence.

Status update templates

Keep these in the runbook with placeholders clearly marked. Adjust tone for the audience — internal teams get technical detail, clients get business impact.

## INITIAL ACKNOWLEDGMENT (within 15 min of P1)
Subject: [INCIDENT] Service degradation — {SERVICE_NAME}
Status: Investigating
Impact: {USER_FACING_SYMPTOM}
Started: {TIMESTAMP_NPT}
Next update: {TIMESTAMP + 30 MIN}

We are aware of {SYMPTOM} affecting {SCOPE}.
Engineering is actively investigating.
No ETA yet — we will update in 30 minutes.

## PROGRESS UPDATE (every 30-60 min)
Status: {INVESTIGATING | IDENTIFIED | MITIGATING | RECOVERING}
Current action: {WHAT_WE_ARE_DOING_NOW}
ETA: {ESTIMATE_OR_UNKNOWN}
Workaround: {IF_AVAILABLE}

## RESOLUTION NOTICE
Status: Resolved
Duration: {START_TIME} to {END_TIME} ({TOTAL_MINUTES} min)
Root cause: {ONE_SENTENCE_SUMMARY}
Prevention: {FOLLOW_UP_ACTION_PLANNED}

For Nepal-based operations, consider bilingual updates when the customer base includes non-English speakers. Also account for local business hours and festivals. An incident during Dashain needs a different communication cadence than a regular Tuesday. Tie these templates into your server security and monitoring documentation so the on-call engineer has full context without hunting. The Google SRE book chapter on postmortem culture is a solid reference for the tone these updates should carry.

What belongs in a post-incident review for Laravel applications?

The post-incident review is where organisational learning happens. Without it, you fix the same bug three times in six months. A review is not a blame session. It is a structured analysis of systemic failure, and for Laravel applications it should target concrete improvements to code, configuration, or process.

Review document structure

  • Timeline: Minute-by-minute account from first alert to full recovery. Pull timestamps from monitoring tools, not memory.
  • Impact quantification: Orders lost, users affected, revenue estimate. "23 failed checkout attempts during the Valentine's peak" beats "checkout was down."
  • Root cause analysis: Use the Five Whys. "Queue failed" → why? → "Redis ran out of memory" → why? → "Memory limit too low" → why? → "Never tuned after traffic doubled" → why? → "No capacity review process."
  • Action items: Concrete tasks with owners and deadlines. "Raise Redis maxmemory to 2GB" beats "monitor Redis better." Link every item to a ticket.
  • Runbook updates: What was missing or wrong in the current document? This section closes the loop.

Schedule reviews within 48 hours of any P1. For distributed teams, record the session for async review. The point is to update the On-Call and Incident Response Runbook so the next incident is faster, cheaper, and less stressful. Pair this with the patterns in blameless postmortems that actually help and the wider incident response playbook.

MTTR With and Without a RunbookNo RunbookWith RunbookGuess server path — 12 minSearch Slack — 20 minRestart wrong service — 8 minRebuild config — 15 minTotal: ~55 minOpen runbook — 3 minRun triage commands — 5 minRoll back release — 6 minVerify recovery — 4 minTotal: ~18 minSame incident, same engineers, different documentation
How an On-Call and Incident Response Runbook reduces mean time to recovery on Laravel production incidents

Key Takeaways

  • Write each runbook entry so a tired engineer can run it at 3 AM without opening another document.
  • Pin exact PHP and service names — php8.4-fpm, not php-fpm — and update them with every upgrade.
  • Define severity by business impact, then publish the matrix so alerting rules match it.
  • Keep pre-written stakeholder templates in the runbook, with placeholders and an NPT timestamp format.
  • Hold a post-incident review within 48 hours and treat every finding as a runbook edit.
  • Test triage commands during quiet periods, not during the first real outage.
Feedback LoopIncident OccursP1 or P2 pageRunbook UsedTriage and fixReview HeldWithin 48 hoursRunbook UpdatedNew commands addedLoop closes every incidentFaster MTTR, lower stress, higher reliability
Post-incident feedback loop that keeps an On-Call and Incident Response Runbook current

People Also Ask

What is the difference between a runbook and a playbook?

A runbook holds the operational detail for a specific alert: exact commands, expected output, and known fixes. A playbook describes the wider process — who leads, how you communicate, when you escalate. The runbook is what you execute at 3 AM. The playbook is what your organisation follows around it. Most teams need both, and the runbook is the one that saves minutes.

How often should an incident response runbook be updated?

Update it whenever infrastructure changes and after every P1 or P2 incident. A PHP version bump, a new queue driver, a hosting migration, or a renamed release path all invalidate commands. If your team has not touched the runbook in six months, assume parts of it are wrong. Stale commands are worse than no commands, because they cost time and confidence.

Who should own the on-call runbook?

The engineer who carries the pager owns the content, even if someone else edits the file. Ownership means testing commands, chasing missing escalation contacts, and closing action items from reviews. In small teams that is usually the senior developer. In larger teams it rotates with the on-call schedule. Either way, one name sits against each section.

What is a realistic MTTR for a Laravel application?

For a well-instrumented Laravel app with a maintained runbook, 15–30 minutes is achievable for common failures like queue backlogs, cache issues, or a bad deploy. First-time incidents in unfamiliar areas take longer, which is exactly why the runbook exists. Track MTTR per incident type rather than as a single average. That tells you where documentation is still thin.

Build the Runbook Before the Next Outage

An On-Call and Incident Response Runbook is not a document you write once and archive. It evolves with your application, your team, and your infrastructure. Start small this week: pick your three most frequent alerts and document them properly. Test every command in a quiet window. Hand the draft to whoever takes the next on-call shift and ask what is missing. Add the monitoring layer that feeds it — a Prometheus and Grafana monitoring stack with SLO-driven alerting and a proper on-call alerting tool — so the runbook starts from a real signal.

If you are weighing the cost of downtime against better tooling or a larger VPS, run the numbers with the Nepal EMI calculator before committing budget. For teams that want the full picture, the SRE vs DevOps roles comparison and how to write effective runbooks cover the surrounding practice. You can see the kind of portals this discipline supports in the secure client portal with document sharing and the legal-guide site with lead capture.

If your production environment has no structured incident response, or your runbook is outdated and untested, get in touch through the contact page. I help teams build practical, maintainable operational documentation grounded in real production work — not theoretical frameworks that collapse under pressure. Whether you need a runbook audit, deployment pipeline hardening, or hands-on incident response training for your developers, we can scope an engagement that fits your operational reality and budget.

Frequently Asked Questions

A documented procedure guiding engineers through detecting, triaging, mitigating, and resolving production incidents. It includes escalation paths, diagnostic commands, rollback steps, and communication templates to reduce mean time to recovery during outages.

Small teams lack redundant staffing, so tribal knowledge fails when the primary engineer is unavailable. A runbook ensures any competent developer can stabilize production systems at 3 AM using verified steps rather than guessing under pressure or relying on memory.

Initial creation takes 20-40 hours for core services, roughly Rs 75,000-150,000 (USD 550-1,100) at senior Nepal rates. Ongoing maintenance requires 4-8 hours monthly as infrastructure evolves, costing Rs 15,000-30,000 (USD 110-220) per month.

Every runbook needs severity definitions, escalation contacts with phone numbers, initial triage checklists, service-specific diagnostic commands, approved mitigation procedures, rollback instructions, stakeholder communication templates, and post-incident review requirements. Missing any section creates gaps that cause delays during actual emergencies when stress is high and cognitive load exceeds normal capacity.

Test quarterly via tabletop exercises simulating real failures. Update immediately after every production incident, infrastructure change, or dependency upgrade. In my experience maintaining Laravel applications on shared EC2 infrastructure, runbooks stale for six months become unreliable because PHP versions, Deployer configs, and database schemas drift from documented state without deliberate synchronization efforts.

GitLab CI pipelines trigger alerts via webhook to Slack or email. PagerDuty or Grafana OnCall handle escalation scheduling. Store runbooks in Git repositories alongside code for version control. For Nepal Gift Card and similar projects, I use GitLab issues as incident tickets linked to runbook commits, ensuring documentation stays synchronized with actual deployment configurations and rollback procedures.

Document php-fpm restart commands, queue worker recovery via supervisorctl, Redis cache flush procedures, and database connection troubleshooting. Include artisan commands for checking scheduled tasks and failed jobs. Specify log file paths under storage/logs and common error patterns like token mismatch or migration failures. Always verify these commands work on your exact Ubuntu and PHP-FPM configuration before documenting them as authoritative recovery steps.

Use three tiers. SEV1 means complete service outage affecting all users requiring immediate response. SEV2 indicates degraded functionality impacting significant user segments needing resolution within four hours. SEV3 covers minor issues with workarounds available, resolved next business day. Avoid five-tier systems designed for enterprises; they create confusion and slow triage decisions when small teams face production pressure at odd hours.

Document provider-specific status page URLs, API health check endpoints, and webhook verification steps. Include manual reconciliation procedures for eSewa, Khalti, or ConnectIPS when callbacks fail. Specify when to switch to backup gateways versus waiting. On WooCommerce florist sites processing international payments, I maintain separate runbooks for each gateway because Stripe timeout handling differs completely from local Nepal payment provider error responses and retry logic.

Vague instructions like "check the server" instead of exact SSH commands. Outdated credentials or deprecated tool references. Missing rollback procedures after failed deployments. Assuming reader knows internal architecture. Runbooks written once then abandoned. Effective runbooks read like executable scripts with copy-pasteable commands, current version numbers, and explicit decision trees that guide exhausted engineers through recovery without requiring architectural knowledge they may lack.

Implement follow-the-sun rotation if possible, otherwise limit on-call to one week monthly maximum. Compensate with time off after overnight incidents. Automate repetitive diagnostics via monitoring alerts with contextual data. Set clear severity thresholds preventing unnecessary wake-ups. In my experience supporting legal-tech portals, defining SEV3 as next-business-day resolution prevented 80% of false alarms while maintaining genuine emergency coverage for critical client-facing services.

Track mean time to acknowledge, mean time to resolve, and incident recurrence rate. Measure percentage of incidents resolved using runbook versus ad-hoc troubleshooting. Monitor on-call page frequency and false positive rates. After implementing structured runbooks for Adventure Third Pole Trek booking system, MTTR dropped from 90 minutes to 25 minutes because engineers stopped rediscovering diagnostic steps during each outage and followed validated procedures instead.

WordPress runbooks focus on plugin conflicts, theme errors, wp-cli commands, and database repair via WP-CLI. Laravel runbooks emphasize queue workers, job failures, artisan commands, and framework-specific caching. WordPress incidents often resolve via plugin deactivation; Laravel requires understanding service container bindings and middleware stacks. Both need database backup restoration steps, but Laravel migrations add complexity absent in WordPress content management workflows and update cycles.

Yes. Pre-write status page updates, email notifications, and social media posts for each severity level. Legal-tech clients especially need compliant language avoiding liability admission while maintaining transparency. Templates reduce decision fatigue during crises when crafting careful wording competes with technical troubleshooting. Include placeholders for incident timeline, affected services, estimated resolution time, and post-mortem commitment. Review templates quarterly with stakeholders to ensure tone matches current brand voice and regulatory requirements.

Never store passwords, API keys, or tokens in runbook documents. Reference environment variables or secrets managers like HashiCorp Vault or AWS Secrets Manager. Use placeholder syntax indicating where credentials inject at runtime. For Deployer-based deployments, document how to access shared .env files securely via SSH rather than embedding connection strings. Audit runbook access logs quarterly. Rotate any credential accidentally committed to version control immediately and treat the exposure as a security incident requiring its own response procedure.

Share this article

0 Comments

Leave a comment

Your email is not published. Comments appear once they have been read. Sign in to have your details filled in.

Quick Contact Options
Choose how you want to connect me: