Observability as evidence infrastructure for consequential systems

Practice · Observability & Evidence

Standard observability answers what happened right now; a consequential system, such as one under the HIPAA Security Rule held to a 99.9% availability SLO, also needs its logs, metrics, and traces kept as a tamper-evident audit trail that proves what happened to someone asking weeks later.

How is observability different for a consequential system?

Standard observability answers what happened right now, for whoever is debugging. A consequential system also needs its logs, metrics, and traces kept as a tamper-evident audit trail that proves what happened to someone asking weeks later.

Standard observability answers what happened right now, for whoever is debugging. A consequential system needs its observability data to answer a second, harder question later: can you prove what happened, to someone who wasn't there and is asking weeks afterward. That second question changes what gets logged, how long it's kept, and whether it can be queried at all.

Observability answers "what happened." Evidence requires "and can you prove it."

The standard three pillars of observability — logs, metrics, traces — are built to answer an engineer's question in the moment. A consequential system asks each pillar a second question: not just what happened, but whether that record is durable, tamper-evident, and queryable well after the moment has passed.

The three pillars, and the evidence question each one adds

PillarDebugging questionEvidence question a consequential system also needs
LogsWhat happened at this moment?Is the log tamper-evident, and retained long enough to matter when someone asks later?
MetricsIs the system healthy right now?Can a threshold breach be tied back to a specific incident record after the fact?
TracesWhy was this one request slow?Does the trace identifier survive into the incident/RCA record, closing the loop?

Observability data as root-cause-analysis input

For CAPA under ISO 13485, root-cause analysis in MedTech needs the observability data to already exist as an audit trail before the incident happens — see Production Incident Evidence, Root Cause Analysis & CAPA. Reconstructing what a system was doing from memory, after the fact, is not an evidence-based RCA regardless of how confident the reconstruction sounds.

Queryable, not just collected

"We have logs somewhere" and "we can answer a specific compliance question with a query in minutes" are different capabilities, and only the second one is actually evidence infrastructure. Data that exists but can't be queried on demand imposes the same cost as data that was never collected, the first time someone actually needs it.

Engineering reference only. Specific retention periods and query SLAs are a function of your own regulatory and contractual obligations, not a universal number this page prescribes.

Provenance & review state

Last reviewed
Sources
  • NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (AU — Audit and Accountability family) — National Institute of Standards and Technology
  • OpenTelemetry specification — OpenTelemetry Authors / Cloud Native Computing Foundation
Ingested from

Sign in or sign up

Enter your work email to receive a temporary sign-in link.

By continuing, you agree to our Terms of Service and Privacy Policy.