ObservabilityHealthTech

Full-stack visibility across 200+ microservices in 2 weeks

Instrumented a 200-service estate with an open observability stack on HealthStack’s own infrastructure, cutting MTTR by 85%.

HealthStack Enterprise

200+
Services instrumented
< 2min
Mean time to detect
85%
MTTR reduction
2 weeks
Full deployment

The challenge

HealthStack was operating 200+ microservices with zero observability. Incidents took hours to diagnose and the team had no insight into service dependencies, latency, or error patterns.

Our approach

  • Auto-instrumented services with OpenTelemetry and stood up a Grafana, Prometheus, Loki, and Tempo stack on their own infrastructure.
  • Built service dependency maps and correlated metrics, logs, and traces into a single view.
  • Defined SLOs and runbook-linked alerts so on-call gets actionable pages, not noise.

What we shipped

An OpenTelemetry-based observability stack with dependency maps, SLO dashboards, and runbook-linked alerting — owned entirely by HealthStack.

Stack

OpenTelemetryGrafanaPrometheusTempoKubernetes

Their observability work gave us visibility we never had before. We went from flying blind with 200 microservices to full coverage in under two weeks.

Priya SharmaVP Engineering, HealthStack

Have a similar challenge?

Tell us what you're building. You'll hear from a senior engineer within one business day.

Start a project →