A growing number of reported incidents often signals improving system reliability rather than declining performance, according to a recent article from Great Circle. The analysis challenges one of engineering leadership's most widespread beliefs: that rising incident counts point to worsening system health. Instead, the report argues, higher incident volumes typically reflect stronger incident management cultures where teams are more willing to formally declare operational problems that previously would have been handled quietly or hidden entirely.
The shift happens as organizations invest in stronger processes, tooling, training, and operational discipline, the report explains. Engineers who once resolved degraded services or "spicy bugs" informally now choose to trigger formal incident declarations, launching coordinated response efforts, documentation, and postmortems. While dashboards display more incidents initially, the organization becomes more resilient by making operational knowledge visible and repeatable. The pattern mirrors what happens in other engineering domains: companies that strengthen vulnerability management report more security findings not because systems suddenly became less secure, but because detection improved. Similarly, better observability generates more alerts, and enhanced testing uncovers more defects before production—in each case, improved measurement exposes existing problems rather than creating new ones.
According to the report, incident count measures an organization's willingness to surface and manage operational problems rather than underlying system reliability. Mature incident management cultures encourage engineers to declare incidents early, involve appropriate stakeholders, and conduct structured post-incident reviews—practices that initially cause reported incidents to rise but also create opportunities to learn, improve processes, and prevent larger outages down the line. The article warns that if engineers believe they'll be judged on keeping incident counts low, they may delay declarations, attempt solo problem resolution, or avoid escalating emerging issues until they worsen significantly. Recent guidance from Sygnia similarly argues that traditional incident metrics including raw counts, ticket volumes, and mean time to respond when viewed alone can create false confidence because they measure operational activity rather than organizational readiness.
The report positions this perspective within broader reliability engineering thinking that has moved away from simplistic operational metrics toward measures reflecting customer impact, recovery effectiveness, and organizational learning. Modern Service Level Objective practices increasingly emphasize user-centric measures such as error budget burn, service level indicator degradation, and customer impact over infrastructure-centric statistics, the analysis notes. Understanding how users experience failures provides a far more accurate picture of service health than simply counting incidents or measuring infrastructure uptime alone. The recommended alternative metrics include containment effectiveness, escalation quality, post-incident improvements, and incident response process maturity—shifting questions from "How many incidents occurred?" to "How quickly were users affected?", "How rapidly was service restored?", "Did we identify the root cause?", and "Are similar failures becoming less frequent?"
Organizations should avoid discouraging incident declarations merely to improve dashboard appearances, the report concludes. Healthy engineering cultures reward transparency, recognizing that declaring an incident isn't an admission of failure but the start of a structured learning process. By encouraging early reporting, blameless collaboration, and continuous improvement, companies create environments where operational knowledge accumulates over time rather than remaining trapped within individual responders. The broader lesson: metrics should reflect the outcomes organizations actually care about—a falling incident count may indicate improving reliability, but it could equally suggest under-reporting, inconsistent classification, or a weakening incident culture, while a temporary rise in reported incidents may represent healthier operational practices and better organizational awareness. The distinction matters because the wrong metrics can inadvertently punish the very behaviors that drive long-term resilience, while the right ones create incentives that align individual decisions with collective system health.

