
Monitoring advice is usually written for teams with an on-call rotation and a budget line for observability. If you're one person, the goal is narrower: know when your app breaks, know roughly why, and don't get paged for things that don't matter. Everything beyond that is a hobby, and hobbies are fine, but they shouldn't come before shipping.
Start with the three questions you'll actually ask during an incident. Is the app up? Is it fast enough? Did something change recently? If your tooling answers those quickly, you're in good shape. If it answers a hundred questions you never ask, you've built a dashboard nobody reads.
Uptime checks are the foundation. An external service pinging your health endpoint every minute tells you what your users experience, which is different from what your server thinks. Make the health endpoint check real dependencies, like the database connection, but keep it fast and don't let it do heavy work. Alert to a channel you actually look at, and set it to notify only after two or three consecutive failures so a single blip doesn't wake you.
For performance, start with request latency and error rate. Most frameworks expose these with a few lines of middleware, and you can ship them to a hosted service or your own metrics store. Track them as percentiles rather than averages, because averages hide the slow requests that annoy real users. Add a simple log aggregation path so you can search recent errors without SSHing into the box and tailing files by hand.
Resource metrics matter too, but less than people think. CPU, memory, and disk tell you when you're about to hit a wall, not when users are suffering. Disk is the one to alert on early, since a full disk takes everything down at once. Memory leaks show up as a slow climb, so a weekly glance at the trend catches them before they page you.
Every alert should be actionable and rare. If an alert fires and your response is to shrug, delete it, because it's training you to ignore the ones that matter. Route critical alerts to your phone and everything else to a channel you check during work hours. Write down what each alert means and the first thing to check, so future you isn't debugging the alerting system during an outage.
Finally, do a postmortem on yourself. After any incident, spend ten minutes noting what happened, how you found out, and what would have surfaced it sooner. Usually the answer is a missing check or a noisy one. Iterating on that short list over a few months builds monitoring that fits your app instead of a generic template. That's the whole game.