Kokil Thapa - Professional Web Developer in Nepal
Freelancer Web Developer in Nepal with 15+ Years of Experience

Kokil Thapa is an experienced full-stack web developer focused on building fast, secure, and scalable web applications. He helps businesses and individuals create SEO-friendly, user-focused digital platforms designed for long-term growth.

Ubuntu Server Monitoring Guide

By Kokil Thapa | Last reviewed: September 2026

Production outages on Ubuntu rarely arrive without warning. They build quietly for hours inside metrics nobody is scraping. This Ubuntu server monitoring guide is the setup I run on production PHP and Laravel servers — the same two-layer approach behind legal-tech portals and multi-currency WooCommerce stores I maintain. It covers fast triage, long-term history, application-level metrics, and alert rules that do not wake anyone at 3 a.m. for nothing.

If you would rather not run any of this yourself, that work sits squarely inside Linux system administration services for Nepal-based businesses. The rest of this page assumes you are doing it yourself and want the numbers to be right.

How do you perform real-time Ubuntu server monitoring?

When a server is actively degrading, you do not have time to query a time-series database. You need low-overhead visibility into the current state, and the native CLI tools on Ubuntu 22.04 and 24.04 LTS are still the fastest path. They also tell you things a dashboard cannot, like which specific process is spinning.

CPU and process inspection

The stock top command is not enough for multi-core debugging. Use htop or the newer btop for per-core usage, thread counts, and sortable process tables.

# Install the essential diagnostic set once
sudo apt update
sudo apt install htop btop iotop sysstat

# Full resource view
btop

# Narrow the view to the PHP-FPM user
htop -u www-data

A common mistake on production Laravel servers is misreading load average. On an 8-core instance, a load average of 8.0 means full saturation, not failure. If load is 8.0 while CPU usage sits at 20%, your bottleneck is disk or network, not compute. Always read that number against core count. For a deeper walkthrough, see diagnosing high CPU and memory usage on a Linux server.

Disk I/O, latency and sockets

Database-driven applications stall on storage throughput more often than on slow SQL. The iotop utility shows which process is actually burning disk bandwidth, which is frequently not the one you suspect.

# Only processes doing real I/O
sudo iotop -aoP

# Extended disk stats, 1s interval, 5 samples
iostat -xz 1 5

# Who is connected to your web and database ports
ss -tunap | grep -E ':80|:443|:3306'

In iostat output, watch %util and await. If %util approaches 100% while await exceeds 10 ms on NVMe (or 50 ms on SATA SSD), storage is saturated. For stores processing Dashain-season order spikes, that pair predicts checkout timeouts far better than CPU ever will.

Real-Time Diagnostic Decision TreeSymptom detectedHigh load averageSlow responseshtop / btopiotop / ssCPU boundOptimise or scale outI/O or network waitCheck disk, DB, upstream
Diagnostic decision tree correlating Ubuntu server monitoring symptoms with the right command-line tool

What is the best persistent Ubuntu server monitoring stack in 2026?

Terminal tools vanish when you close the session. Production systems need a time-series database that remembers what normal looked like last Tuesday. In 2026 the default self-hosted answer remains Prometheus paired with Grafana — see the fuller Prometheus and Grafana monitoring stack setup for the long version.

The cost argument matters in Nepal. A managed CloudWatch bill on a moderately busy site can reach Rs 15,000 per month (~USD 112). A Rs 3,000–6,000 (~USD 22–45) VPS running Prometheus, node_exporter and Grafana costs the same every month regardless of how many metrics you collect.

Install the stack as dedicated users

Every component runs as a systemd service on Ubuntu 24.04. None of them should run as root, and none should own files they do not need. The general hardening rules in my Ubuntu security hardening guide apply here too.

sudo useradd --no-create-home --shell /bin/false prometheus
sudo useradd --no-create-home --shell /bin/false node_exporter

sudo mkdir /etc/prometheus /var/lib/prometheus
sudo chown prometheus:prometheus /etc/prometheus /var/lib/prometheus

Configure node_exporter with a narrow collector set

node_exporter exposes hundreds of metrics by default. Most of them are noise on a single application server, and every extra series costs memory in Prometheus. Disable the defaults and enable only what you will actually alert on. The collector list is documented in the official node_exporter repository.

# /etc/default/node_exporter
ARGS="--web.listen-address=127.0.0.1:9100 \
      --collector.disable-defaults \
      --collector.cpu \
      --collector.meminfo \
      --collector.diskstats \
      --collector.filesystem \
      --collector.netdev \
      --collector.systemd"

Binding to 127.0.0.1 keeps the endpoint off the public internet. Prometheus then scrapes over a private VPC address or an SSH tunnel, never plain HTTP on a public interface. That habit lines up with the exposure patterns I flagged in cybersecurity trends for 2026, and with the basics in securing your website and server in Nepal.

Persistent Monitoring Architecturenode_exporter:9100 localhostOS metricsphp-fpm-exporter:9253 internalPool metricsPrometheusTSDB storage15s scrapeGrafanaDashboardsAlert rulesHTTP scrapeHTTP scrapeQuery APIBound to localhost or VPC · UFW blocks 9100 and 9090 · no public metric endpoints
Secure Ubuntu server monitoring dashboard topology with isolated metric endpoints

Which PHP-FPM and Laravel metrics should you scrape?

System metrics miss application-layer failure. A server can sit at 30% CPU while PHP-FPM has exhausted its worker pool and requests queue in the kernel socket backlog. On Laravel and WordPress deployments you have to instrument the runtime, not just the box.

PHP-FPM pool status

Enable the FPM status endpoint in your pool configuration. The path below assumes PHP 8.3 or 8.4, which most production servers still run; the same directives apply on PHP 8.5. The tuning background is in PHP-FPM tuning for high-traffic websites.

# /etc/php/8.4/fpm/pool.d/www.conf
pm.status_path = /status
ping.path = /ping
ping.response = pong

Restrict both paths in Nginx to localhost or an internal range, then let a PHP-FPM exporter translate the status page into Prometheus format. Three series carry most of the value:

  • active_processes / max_children: sustained values above 80% mean request queuing is imminent.
  • slow_requests: a rising counter flags backend bottlenecks even while CPU is idle.
  • listen_queue_len: anything non-zero means the socket backlog is filling and users are waiting.

Queue and job health

Queue-heavy applications fail quietly. The web tier answers fine while workers die and orders sit unprocessed. Laravel Horizon exposes Redis-backed metrics you can scrape directly — the setup is covered in monitoring and scaling Laravel queues with Horizon. Without Horizon, export the essentials yourself.

// AppServiceProvider or a dedicated collector
$failed = DB::table('failed_jobs')
    ->where('failed_at', '>', now()->subHour())
    ->count();

Statsd::gauge('laravel.queue.failed_last_hour', $failed);
Statsd::gauge('laravel.queue.depth', Queue::size('default'));

I have watched eCommerce deployments where every uptime check passed while the queue worker had been dead for six hours. System metrics said the server was healthy. Customers said their orders never arrived. Only application-level monitoring catches that gap.

How do you configure alerts without causing on-call fatigue?

Alert fatigue destroys the value of monitoring. If more than one alert in twenty needs no human action, the channel is already being ignored. The discipline that fixes this is simple: alert on user-visible symptoms, not on internal causes. That is the core idea behind the four golden signals of monitoring.

Alert typeWeak exampleBetter exampleWhy it matters
Cause-basedCPU > 80% for 5 minutesHTTP 5xx rate > 1% for 3 minutesHigh CPU is normal during batch jobs; errors hurt users
SaturationDisk 90% fullDisk fills within 4 hours at current ratePredictive alerts allow a planned fix instead of a panic
AvailabilityProcess not runningHealth endpoint failing for 2 minutesRestarts are routine; a failed health check is not
PerformanceResponse time > 500 msp95 latency above SLA for 5 minutesAverages hide outliers; percentiles match user experience

Route by severity, not by convenience

Critical alerts — service down, data loss risk — go to a phone. Warnings go to a channel someone reads during working hours. Never send everything everywhere. The routing rules below follow the same shape as the Prometheus Alertmanager alerting guide.

# /etc/prometheus/alertmanager.yml
route:
  receiver: 'slack-warnings'
  routes:
    - match:
        severity: critical
      receiver: 'pagerduty-critical'
      continue: false
    - match:
        severity: warning
      receiver: 'slack-warnings'

receivers:
  - name: 'pagerduty-critical'
    pagerduty_configs:
      - service_key: '<your-key>'
  - name: 'slack-warnings'
    slack_configs:
      - channel: '#infra-alerts'

Keep every host, database and monitoring component on UTC. Nepal runs on NPT (UTC+5:45), and mixing local timestamps in logs with UTC timestamps in metrics makes incident correlation painful. Convert to NPT only in the dashboard, never at the point of collection.

Severity-Based Alert RoutingAlertmanagerRoute by severity labelCRITICALPagerDuty + SMSWARNINGSlack #infra-alertsINFOWeekly email digestImmediate responseBusiness hours reviewTrend audit
Alert severity routing prevents notification fatigue in production Ubuntu server monitoring

What are the most common Ubuntu server monitoring mistakes?

Instrumented systems still fail when the monitoring itself becomes unreliable. These are the patterns I run into most often during production audits.

  1. Monitoring the monitor. Add a dead-man's switch. If Prometheus stops scraping or Alertmanager stops sending heartbeats, an external checker such as Healthchecks.io or UptimeRobot should tell you. Silent monitoring failure is worse than none.
  2. Cardinality explosions. Labels like user_id or request_id will crash Prometheus within hours. Keep label sets bounded and use logs for per-request tracing instead.
  3. Mixing timezones. One host on NPT and the rest on UTC makes every incident investigation slower than it needs to be.
  4. Alerts with no runbook. An alert reading "disk full" without a remediation link wastes half an hour. Point it at your wiki, or at practical references such as log rotation and disk space management on Linux.
  5. Testing alerts only in staging. Staging load never matches production. Fire synthetic faults in production occasionally to prove the routing works.
  6. Letting retention eat the disk. Fifteen days of raw samples is plenty for forensics on most web applications.

Budget the resources monitoring consumes

On a 2 GB VPS, Prometheus and Grafana can take 30–40% of memory between them. Cap retention explicitly, and consider a lightweight agent like zero-config server monitoring with Netdata on very small instances. Never let monitoring starve the application it protects.

# /etc/default/prometheus
ARGS="--storage.tsdb.retention.time=15d \
      --storage.tsdb.retention.size=8GB \
      --query.max-samples=50000000"

Retention is also a compliance question. If you are VAT-registered in Nepal, the Inland Revenue Department expects business records to be kept for years, not days, so check the current requirement at the Inland Revenue Department before you trim storage. Metrics retention and financial record retention are separate decisions — do not let one silently break the other. The same logic applies to the fiscal-year reporting cycles in 2082/83 BS that most Nepali finance teams reconcile against.

Start from a hardened server

Monitoring agents read sensitive system state, so the base image matters. Work through initial Ubuntu server setup for a fresh VPS first, then add the observability layer. If you are sizing the hardware or the cloud spend behind all of this, the Nepal EMI calculator is a quick way to see what a hardware or hosting commitment really costs per month.

I have also run this exact stack on legal-tech portals holding client documents, including the Notary Nepal platform, where audit trails and uptime both matter. When the operational side needs an owner rather than a hobby, that is exactly what DevOps automation work in Nepal covers — or you can get in touch directly about a specific setup.

Key Takeaways

  • Keep two layers: CLI tools for triage, Prometheus for history.
  • Read load average against core count before calling it saturation.
  • Alert on symptoms — error rate, latency percentiles, health checks.
  • Scrape PHP-FPM pool stats and queue depth; system metrics miss them.
  • Bind every exporter to localhost or a private network range.
  • Cap retention at 15 days and verify the monitor itself is alive.

People Also Ask

Is Prometheus overkill for a single Ubuntu VPS?

Not necessarily, but it is a trade. Prometheus plus Grafana costs roughly 400–700 MB of RAM on a small box. On a 1 GB instance, Netdata or a hosted exporter is the better call. On a 4 GB application server, the self-hosted stack is cheap and predictable.

What should I monitor first on a new server?

Start with four things: CPU load per core, memory and swap usage, disk space plus %util, and whether your health endpoint responds. Add application metrics only once those four are alerting correctly.

How do I monitor a server behind a firewall?

Never open exporter ports publicly. Scrape over a private VPC address, an SSH tunnel, or a mesh VPN such as WireGuard. Prometheus pulls; it does not need inbound access from the internet.

Do I still need log monitoring if I have metrics?

Yes. Metrics tell you that error rates rose; logs tell you why. A log pipeline with structured JSON output and a retention window of 7–14 days covers most incident investigations without becoming a storage problem.

Build Monitoring You Actually Trust

Good Ubuntu server monitoring is not a tool choice. It is the habit of measuring what users feel, storing enough history to compare against, and alerting only when a human needs to act. Start with the CLI tools, add Prometheus and Grafana once the basics are stable, instrument PHP-FPM and your queues, and review the dashboards every quarter.

The specific versions will keep moving — PHP 8.5, Laravel 13, PostgreSQL 18 — but the principles do not. If you want a production observability setup designed, audited or handed over properly, talk to me about your infrastructure.

Frequently Asked Questions

Netdata, Prometheus with Node Exporter, and Glances are top choices. All run natively on Ubuntu 24.04 LTS without licensing fees or cloud dependencies.

Lightweight agents like Node Exporter use under 50MB. Full-stack tools like Netdata or Zabbix agent may require 200-400MB depending on metrics collected and retention settings.

Choose Prometheus for cloud-native, metric-driven environments using time-series data. Pick Nagios for traditional host/service checks, legacy infrastructure, or when state-based alerting is primary.

Run the official kickstart script which handles dependencies and service setup automatically. In my experience managing production servers, this method avoids package conflicts common with apt installs. Post-installation, edit netdata.conf to set memory limits and disable unused collectors. Configure streaming to a central node if monitoring multiple servers. The dashboard becomes available immediately on port 19999, but always restrict access via UFW or reverse proxy authentication in production environments to prevent information disclosure.

Beyond basic CPU and memory, monitor disk I/O wait, inode usage, TCP connection states, and PHP-FPM or Nginx worker saturation. On legal-tech portals I have built, tracking database connection pool exhaustion prevented more outages than raw CPU alerts. Always measure application-level latency alongside system metrics. Configure alerts for sustained high load averages rather than instantaneous spikes to reduce noise. Track swap activity as an early warning for memory pressure before OOM kills occur.

Never expose monitoring ports directly to the internet. Use Nginx as a reverse proxy with HTTP basic auth or OAuth2-proxy in front of Netdata, Grafana, or Prometheus. Restrict access via UFW to trusted IPs only. In production deployments I manage, monitoring interfaces sit behind VPNs or SSH tunnels. Rotate credentials regularly and audit access logs. Disable default admin accounts and enforce TLS even on internal networks to prevent credential sniffing during lateral movement.

This usually indicates single-threaded processes saturating one core while others remain idle. Check per-core utilization with htop or mpstat. Common causes include unoptimized PHP scripts, runaway cron jobs, or database queries lacking proper indexing. On Laravel applications I have debugged, synchronous queue workers often cause this pattern. Profile the specific process consuming CPU rather than relying on aggregate metrics. Consider cgroup limits to isolate noisy workloads from critical services.

Combine Promtail or Fluent Bit with Loki or Elasticsearch for log aggregation. Configure structured logging in your application to enable correlation between metrics and log events. On production systems I maintain, I tag logs with request IDs matching trace IDs in metrics. Set up alerting on error rate thresholds rather than individual log lines to avoid fatigue. Retain raw logs separately from aggregated metrics for forensic analysis. Ensure log rotation prevents disk exhaustion during traffic spikes.

Set warnings at 80% and critical alerts at 90% for root and data partitions. However, also monitor inode usage separately since many small files can exhaust inodes before space runs out. On eCommerce sites handling uploads, I have seen inode exhaustion crash systems with 40% disk space remaining. Configure alerts based on growth rate predictions, not just current usage. Exclude temporary directories from alerts but monitor them for cleanup failures. Test recovery procedures before relying on automated cleanup scripts.

Use certbot certificates command for Let's Encrypt managed certs or integrate ssl_exporter with Prometheus for comprehensive monitoring. Configure alerts thirty days before expiry to allow renewal buffer. On servers I manage with Deployer 7, certificate checks run as part of deployment validation. Monitor both leaf and intermediate chain validity. Test auto-renewal hooks monthly since silent failures are common. For wildcard certificates, track DNS propagation delays that can cause renewal timeouts during maintenance windows.

Yes, cAdvisor exposes container metrics via Prometheus endpoint without modifying containers. Alternatively, Docker daemon exposes stats API natively. In containerized Laravel deployments I have operated, cAdvisor provides sufficient visibility for most debugging. Avoid installing monitoring agents inside containers as it increases image size and attack surface. Use labels for service identification rather than container names which change on restart. Correlate container metrics with host metrics to distinguish application issues from resource contention.

Capture packet traces with tcpdump during incidents and analyze with Wireshark. Monitor TCP retransmission rates, connection resets, and DNS resolution times continuously. On production APIs I have maintained, intermittent timeouts often traced to conntrack table exhaustion or misconfigured keepalive settings. Check ethtool for NIC errors and driver issues. Verify MTU consistency across network path. Use mtr instead of ping to identify routing problems. Correlate network anomalies with application logs to distinguish infrastructure from code-level issues.

Centralize metrics with Prometheus federation or VictoriaMetrics and use Grafana for unified dashboards. Deploy consistent exporters via Ansible or Puppet to ensure configuration parity. On sister sites sharing Deployer 7 pipelines, standardized monitoring configs prevent drift across environments. Implement service discovery rather than static targets to handle dynamic scaling. Separate infrastructure metrics from business metrics in distinct dashboards. Establish baseline performance profiles per server role to detect anomalies faster than generic thresholds allow.

Audit alerts quarterly and remove those not requiring immediate human action. Group related alerts into single notifications using Alertmanager inhibition rules. On production systems I operate, fewer than five percent of configured alerts trigger pages. Use severity levels strictly: page only for user-impacting issues. Implement maintenance windows for planned work. Document runbooks linked directly from alert messages. Track mean time to acknowledge and resolve to identify poorly tuned alerts. Silence flapping metrics until root causes are fixed.

For teams under three engineers, managed services save operational overhead despite costing USD 15-30 per host monthly (NPR 2,000-4,000). Self-hosted stacks require ongoing maintenance time that small teams cannot spare. However, for Nepal-based projects with budget constraints, Prometheus and Grafana provide equivalent capability at zero licensing cost. I recommend starting self-hosted and migrating to managed only when monitoring maintenance exceeds development velocity. Evaluate total cost including engineer hours, not just subscription fees. Data residency requirements may also favor local self-hosted solutions.

Share this article

0 Comments

Leave a comment

Your email is not published. Comments appear once they have been read. Sign in to have your details filled in.

Quick Contact Options
Choose how you want to connect me: