# How incidents work

An incident is the record of an outage: when a monitor started failing, how long it stayed down, and when it recovered. TunnelHQ opens and closes incidents automatically from check results.

## When an incident opens and closes

It depends on how many locations the monitor tests from.

### Monitors that test from one location

The first failed check opens an incident straight away, marked **Degraded**, while the monitor retries. If the failed checks in a row reach the monitor's **Retries** setting (2 by default), the incident becomes **Down**. If a check passes first, the incident closes as **Recovered**, so a short blip leaves a short Degraded incident behind.

### Monitors that test from several locations

An incident opens only when the [aggregate status](/docs/concepts/regions/#aggregate-status) goes **Down**. **Partial** doesn't open one, and neither does the aggregate **Degraded** status (a refused credential or certificate). The incident closes when the aggregate status goes back to **Healthy**, or drops from **Down** to **Partial**.

### Incident states

| State | Meaning |
| --- | --- |
| Degraded | A check failed and the monitor is retrying. It becomes **Down** once the failed checks in a row reach the Retries setting, or closes if the next check passes. Only monitors that test from one location open Degraded incidents. |
| Down | The monitor is confirmed down. For monitors with several locations, this follows the aggregate status, so the incident reflects all locations, not one. |
| Recovered | The monitor came back, and the incident is closed. |

:::note[Two kinds of degraded]
A **Degraded incident** means a check failed and the monitor is retrying; the `monitor.degraded` [webhook event](/docs/webhooks/#about-monitordegraded) means the same thing. The aggregate **Degraded status** is different: the server refused the monitor's credentials or certificate, and it never opens an incident.
:::

:::note[Start time is the current outage]
An incident's start time is when the *current* outage began. If a server flaps, recovering and then failing again, the next failure opens a new incident dated to that failure instead of reusing the earlier start time.
:::

## The Incidents page

**Incidents** in the sidebar lists every incident in the current project. The cards at the top show **Active Outages**, **Need Attention**, **Acknowledged**, and **Recovered**. Filter by time range or status, search with the **Search monitors...** box, and download a record with **Export CSV**.

## The activity log

Every incident has an activity log. It records when the incident opened, changed status, and was resolved, and who acknowledged it, and you can add comments to it. The API returns the same log at `GET /incidents/:id/events`.

## Acknowledging an incident

Acknowledge an incident to show that someone is handling it. The incident records who acknowledged it and an optional note, so an outage that's being handled is easy to tell apart from one nobody has looked at yet.

## Getting notified

Alerts follow a monitor's status, not its incidents. The [alert channels](/docs/alert-channels/) you set up under **Integrations → Add Integration**, such as email and Slack, alert when a monitor goes **Down** and when it recovers. Choose which integrations alert for each monitor in its [notification integrations](/docs/concepts/monitors/#notification-integrations) setting.

[Webhook](/docs/webhooks/) endpoints subscribed to `monitor.degraded` and `incident.created` also hear when a Degraded incident opens.
