
August 25, 2026
12 min read
By Kokil Thapa | Last reviewed: September 2026
Production outages on Ubuntu rarely arrive without warning. They build quietly for hours inside metrics nobody is scraping. This Ubuntu server monitoring guide is the setup I run on production PHP and Laravel servers — the same two-layer approach behind legal-tech portals and multi-currency WooCommerce stores I maintain. It covers fast triage, long-term history, application-level metrics, and alert rules that do not wake anyone at 3 a.m. for nothing.
If you would rather not run any of this yourself, that work sits squarely inside Linux system administration services for Nepal-based businesses. The rest of this page assumes you are doing it yourself and want the numbers to be right.
How do you perform real-time Ubuntu server monitoring?
When a server is actively degrading, you do not have time to query a time-series database. You need low-overhead visibility into the current state, and the native CLI tools on Ubuntu 22.04 and 24.04 LTS are still the fastest path. They also tell you things a dashboard cannot, like which specific process is spinning.
CPU and process inspection
The stock top command is not enough for multi-core debugging. Use htop or the newer btop for per-core usage, thread counts, and sortable process tables.
# Install the essential diagnostic set once
sudo apt update
sudo apt install htop btop iotop sysstat
# Full resource view
btop
# Narrow the view to the PHP-FPM user
htop -u www-data A common mistake on production Laravel servers is misreading load average. On an 8-core instance, a load average of 8.0 means full saturation, not failure. If load is 8.0 while CPU usage sits at 20%, your bottleneck is disk or network, not compute. Always read that number against core count. For a deeper walkthrough, see diagnosing high CPU and memory usage on a Linux server.
Disk I/O, latency and sockets
Database-driven applications stall on storage throughput more often than on slow SQL. The iotop utility shows which process is actually burning disk bandwidth, which is frequently not the one you suspect.
# Only processes doing real I/O
sudo iotop -aoP
# Extended disk stats, 1s interval, 5 samples
iostat -xz 1 5
# Who is connected to your web and database ports
ss -tunap | grep -E ':80|:443|:3306' In iostat output, watch %util and await. If %util approaches 100% while await exceeds 10 ms on NVMe (or 50 ms on SATA SSD), storage is saturated. For stores processing Dashain-season order spikes, that pair predicts checkout timeouts far better than CPU ever will.
What is the best persistent Ubuntu server monitoring stack in 2026?
Terminal tools vanish when you close the session. Production systems need a time-series database that remembers what normal looked like last Tuesday. In 2026 the default self-hosted answer remains Prometheus paired with Grafana — see the fuller Prometheus and Grafana monitoring stack setup for the long version.
The cost argument matters in Nepal. A managed CloudWatch bill on a moderately busy site can reach Rs 15,000 per month (~USD 112). A Rs 3,000–6,000 (~USD 22–45) VPS running Prometheus, node_exporter and Grafana costs the same every month regardless of how many metrics you collect.
Install the stack as dedicated users
Every component runs as a systemd service on Ubuntu 24.04. None of them should run as root, and none should own files they do not need. The general hardening rules in my Ubuntu security hardening guide apply here too.
sudo useradd --no-create-home --shell /bin/false prometheus
sudo useradd --no-create-home --shell /bin/false node_exporter
sudo mkdir /etc/prometheus /var/lib/prometheus
sudo chown prometheus:prometheus /etc/prometheus /var/lib/prometheus Configure node_exporter with a narrow collector set
node_exporter exposes hundreds of metrics by default. Most of them are noise on a single application server, and every extra series costs memory in Prometheus. Disable the defaults and enable only what you will actually alert on. The collector list is documented in the official node_exporter repository.
# /etc/default/node_exporter
ARGS="--web.listen-address=127.0.0.1:9100 \
--collector.disable-defaults \
--collector.cpu \
--collector.meminfo \
--collector.diskstats \
--collector.filesystem \
--collector.netdev \
--collector.systemd" Binding to 127.0.0.1 keeps the endpoint off the public internet. Prometheus then scrapes over a private VPC address or an SSH tunnel, never plain HTTP on a public interface. That habit lines up with the exposure patterns I flagged in cybersecurity trends for 2026, and with the basics in securing your website and server in Nepal.
Which PHP-FPM and Laravel metrics should you scrape?
System metrics miss application-layer failure. A server can sit at 30% CPU while PHP-FPM has exhausted its worker pool and requests queue in the kernel socket backlog. On Laravel and WordPress deployments you have to instrument the runtime, not just the box.
PHP-FPM pool status
Enable the FPM status endpoint in your pool configuration. The path below assumes PHP 8.3 or 8.4, which most production servers still run; the same directives apply on PHP 8.5. The tuning background is in PHP-FPM tuning for high-traffic websites.
# /etc/php/8.4/fpm/pool.d/www.conf
pm.status_path = /status
ping.path = /ping
ping.response = pong Restrict both paths in Nginx to localhost or an internal range, then let a PHP-FPM exporter translate the status page into Prometheus format. Three series carry most of the value:
- active_processes / max_children: sustained values above 80% mean request queuing is imminent.
- slow_requests: a rising counter flags backend bottlenecks even while CPU is idle.
- listen_queue_len: anything non-zero means the socket backlog is filling and users are waiting.
Queue and job health
Queue-heavy applications fail quietly. The web tier answers fine while workers die and orders sit unprocessed. Laravel Horizon exposes Redis-backed metrics you can scrape directly — the setup is covered in monitoring and scaling Laravel queues with Horizon. Without Horizon, export the essentials yourself.
// AppServiceProvider or a dedicated collector
$failed = DB::table('failed_jobs')
->where('failed_at', '>', now()->subHour())
->count();
Statsd::gauge('laravel.queue.failed_last_hour', $failed);
Statsd::gauge('laravel.queue.depth', Queue::size('default')); I have watched eCommerce deployments where every uptime check passed while the queue worker had been dead for six hours. System metrics said the server was healthy. Customers said their orders never arrived. Only application-level monitoring catches that gap.
How do you configure alerts without causing on-call fatigue?
Alert fatigue destroys the value of monitoring. If more than one alert in twenty needs no human action, the channel is already being ignored. The discipline that fixes this is simple: alert on user-visible symptoms, not on internal causes. That is the core idea behind the four golden signals of monitoring.
| Alert type | Weak example | Better example | Why it matters |
|---|---|---|---|
| Cause-based | CPU > 80% for 5 minutes | HTTP 5xx rate > 1% for 3 minutes | High CPU is normal during batch jobs; errors hurt users |
| Saturation | Disk 90% full | Disk fills within 4 hours at current rate | Predictive alerts allow a planned fix instead of a panic |
| Availability | Process not running | Health endpoint failing for 2 minutes | Restarts are routine; a failed health check is not |
| Performance | Response time > 500 ms | p95 latency above SLA for 5 minutes | Averages hide outliers; percentiles match user experience |
Route by severity, not by convenience
Critical alerts — service down, data loss risk — go to a phone. Warnings go to a channel someone reads during working hours. Never send everything everywhere. The routing rules below follow the same shape as the Prometheus Alertmanager alerting guide.
# /etc/prometheus/alertmanager.yml
route:
receiver: 'slack-warnings'
routes:
- match:
severity: critical
receiver: 'pagerduty-critical'
continue: false
- match:
severity: warning
receiver: 'slack-warnings'
receivers:
- name: 'pagerduty-critical'
pagerduty_configs:
- service_key: '<your-key>'
- name: 'slack-warnings'
slack_configs:
- channel: '#infra-alerts' Keep every host, database and monitoring component on UTC. Nepal runs on NPT (UTC+5:45), and mixing local timestamps in logs with UTC timestamps in metrics makes incident correlation painful. Convert to NPT only in the dashboard, never at the point of collection.
What are the most common Ubuntu server monitoring mistakes?
Instrumented systems still fail when the monitoring itself becomes unreliable. These are the patterns I run into most often during production audits.
- Monitoring the monitor. Add a dead-man's switch. If Prometheus stops scraping or Alertmanager stops sending heartbeats, an external checker such as Healthchecks.io or UptimeRobot should tell you. Silent monitoring failure is worse than none.
- Cardinality explosions. Labels like
user_idorrequest_idwill crash Prometheus within hours. Keep label sets bounded and use logs for per-request tracing instead. - Mixing timezones. One host on NPT and the rest on UTC makes every incident investigation slower than it needs to be.
- Alerts with no runbook. An alert reading "disk full" without a remediation link wastes half an hour. Point it at your wiki, or at practical references such as log rotation and disk space management on Linux.
- Testing alerts only in staging. Staging load never matches production. Fire synthetic faults in production occasionally to prove the routing works.
- Letting retention eat the disk. Fifteen days of raw samples is plenty for forensics on most web applications.
Budget the resources monitoring consumes
On a 2 GB VPS, Prometheus and Grafana can take 30–40% of memory between them. Cap retention explicitly, and consider a lightweight agent like zero-config server monitoring with Netdata on very small instances. Never let monitoring starve the application it protects.
# /etc/default/prometheus
ARGS="--storage.tsdb.retention.time=15d \
--storage.tsdb.retention.size=8GB \
--query.max-samples=50000000" Retention is also a compliance question. If you are VAT-registered in Nepal, the Inland Revenue Department expects business records to be kept for years, not days, so check the current requirement at the Inland Revenue Department before you trim storage. Metrics retention and financial record retention are separate decisions — do not let one silently break the other. The same logic applies to the fiscal-year reporting cycles in 2082/83 BS that most Nepali finance teams reconcile against.
Start from a hardened server
Monitoring agents read sensitive system state, so the base image matters. Work through initial Ubuntu server setup for a fresh VPS first, then add the observability layer. If you are sizing the hardware or the cloud spend behind all of this, the Nepal EMI calculator is a quick way to see what a hardware or hosting commitment really costs per month.
I have also run this exact stack on legal-tech portals holding client documents, including the Notary Nepal platform, where audit trails and uptime both matter. When the operational side needs an owner rather than a hobby, that is exactly what DevOps automation work in Nepal covers — or you can get in touch directly about a specific setup.
Key Takeaways
- Keep two layers: CLI tools for triage, Prometheus for history.
- Read load average against core count before calling it saturation.
- Alert on symptoms — error rate, latency percentiles, health checks.
- Scrape PHP-FPM pool stats and queue depth; system metrics miss them.
- Bind every exporter to localhost or a private network range.
- Cap retention at 15 days and verify the monitor itself is alive.
People Also Ask
Is Prometheus overkill for a single Ubuntu VPS?
Not necessarily, but it is a trade. Prometheus plus Grafana costs roughly 400–700 MB of RAM on a small box. On a 1 GB instance, Netdata or a hosted exporter is the better call. On a 4 GB application server, the self-hosted stack is cheap and predictable.
What should I monitor first on a new server?
Start with four things: CPU load per core, memory and swap usage, disk space plus %util, and whether your health endpoint responds. Add application metrics only once those four are alerting correctly.
How do I monitor a server behind a firewall?
Never open exporter ports publicly. Scrape over a private VPC address, an SSH tunnel, or a mesh VPN such as WireGuard. Prometheus pulls; it does not need inbound access from the internet.
Do I still need log monitoring if I have metrics?
Yes. Metrics tell you that error rates rose; logs tell you why. A log pipeline with structured JSON output and a retention window of 7–14 days covers most incident investigations without becoming a storage problem.
Build Monitoring You Actually Trust
Good Ubuntu server monitoring is not a tool choice. It is the habit of measuring what users feel, storing enough history to compare against, and alerting only when a human needs to act. Start with the CLI tools, add Prometheus and Grafana once the basics are stable, instrument PHP-FPM and your queues, and review the dashboards every quarter.
The specific versions will keep moving — PHP 8.5, Laravel 13, PostgreSQL 18 — but the principles do not. If you want a production observability setup designed, audited or handed over properly, talk to me about your infrastructure.
Frequently Asked Questions
0 Comments
Leave a comment
Your email is not published. Comments appear once they have been read. Sign in to have your details filled in.

