One thing we've been thinking about while building monitoring systems:
A website returning HTTP 200 doesn't necessarily mean it's working.
You can have:
- a homepage returning 200 while login is broken
- an API returning 200 with an invalid response
- checkout failing halfway through
- SSL working today but expiring tomorrow
- DNS working from one location but failing somewhere else
- a background job silently stopped
- a critical browser flow broken while the page itself loads perfectly
So what should an uptime monitor actually consider a healthy service?
Is HTTP + response time enough for you, or do you expect monitoring to understand the application itself?
Curious how others approach this – especially for production SaaS and e-commerce.
Source: r/Monitorion · by /u/monitorion