Nice overview. Uptime and response time are my top priorities.
What metrics do you actually track for website/server monitoring?
6 Comments
Great list. In production, I try to focus on metrics that lead to action rather than dashboards full of graphs. My core set is uptime, response time, 5xx error rate, CPU, memory, disk usage, SSL certificate expiry, and database health.
Beyond infrastructure, I also monitor application level metrics like queue backlogs, failed jobs, Redis health, slow queries, and API latency.
I've found that a few well tuned alerts with clear ownership are far more valuable than hundreds of noisy ones that everyone eventually ignores.
Monitoring should help you identify and resolve incidents quickly, not just collect data.
Please log in to add a comment.
Everything on that list assumes the failure trips something. The one that got me did not. My daily report quietly stopped counting new users and kept publishing - nothing down, no 5xx, no threshold broken, just a wrong number every morning for eight days before I noticed.
I run a second loop now that watches the first one's errors. It still is not what catches these. What catches them is me not believing a number.
Please log in to add a comment.
Most teams and organizations start with technical metrics. They look at CPUs, response times, status codes, page performance, and so on. A really great observability strategy tracks the business metrics. In retail, you might track completion of checkout steps, basket values, basket abandonment, and financial throughput. When those metrics flicker, you use the technical metrics to work out why, but "is it working" is best defined by the goals the business has for the website.
I spoke to one organization who do hotel bookings. When they flick a feature flag, they pay close attention to the business metric dashboards and will flick the feature flag back off if they don't like what they see. For example, they tried a new search result ordering based on the idea it would increase the number of bookings. They were correct, the number of bookings went up, but each booking was lower value, so their financial throughput started dropping rapidly. Imagine how much money they saved by flicking that one back and fixing the hole they were drilling below the waterline of their business!
Please log in to add a comment.
Please log in to comment on this post.
More Posts
- © 2026 Coder Legion
- Feedback / Bug
- Privacy
- About Us
- Contacts
- Premium Subscription
- Terms of Service
- Early Builders
More From ApogeeWatcherverified
Related Jobs
- Analyste principal(e) - Contrôles et indicateurs de cybersécurité | Cybersecurity Controls & MetricsNTT DATA, Inc. · Full time · Canada
- Software Engineer, Test & Infrastructure II (Bilingual Spanish)Vail Systems · Full time · Springfield, IL
- Metrics Data Analyst-Washington DCSerco · Full time · Chambersburg, PA
Commenters (This Week)
Contribute meaningful comments to climb the leaderboard and earn badges!