---
title: "Monitoring & alerting"
description: "How we monitor servers, workplaces and network: agent-based, thresholds per device type, alerts as tickets, service hours Mon–Fri 08–18, on-call by agreement — plus digital employee experience analytics."
canonical: "https://corevatis-redesign.pages.dev/en/technology/monitoring/"
lang: en
schema_type: WebPage
hreflang:
  de: "https://corevatis-redesign.pages.dev/de/technik/"
  en: "https://corevatis-redesign.pages.dev/en/technology/"
---
# Monitoring & alerting

An alert is only as good as the person who receives it and the instruction that comes with it. That is why every threshold has a recipient and every alert becomes a ticket.

## In brief

- Agent on every device, network via SNMP/API, Microsoft 365 via service health and identity
- Thresholds per device type, tuned against alert floods
- Every alert automatically becomes a ticket in the service portal
- Service hours Mon–Fri 08:00–18:00; on-call outside agreed per client
- Digital employee experience: sign-in times, crashes, device health — analysed, not just collected

## What we do

Servers: CPU, RAM, storage, services, event logs, backup jobs, certificates. Workplaces: health, disk, EDR and patch status, pending reboots. Network: firewall, switches, Wi-Fi controllers, internet connection — reachability and utilisation. Microsoft 365: service health, sign-in anomalies via the identity platform. Plus the user experience: how long sign-in takes, which application crashes, which device is getting slow — before someone calls.

## How exactly

| Parameter | Procedure |
|---|---|
| Check interval | Agent values every minute; network devices via SNMP/ping at short intervals; M365 service health continuously. Exact intervals set per client and device type. |
| Thresholds | Per device type: warning and critical separated (e.g. disk space, service stopped, backup job failed, certificate expiring). Starting values from our standard, fine-tuned in the first weeks. |
| Alert path | Alert → ticket in the service portal with device, value, time. Critical alerts additionally push/call to the responsible person. Recipients and order are in the escalation matrix. |
| Service hours | Mon–Fri 08:00–18:00 response by us. Outside: on-call agreed per client — otherwise handled at the next service start, critical alerts first. |
| Noise | Maintenance windows muted, flapping protection (no alert on brief fluctuations), recurring false alerts are analysed and the threshold adjusted — not the alert ignored. |
| User experience (DEX) | Sign-in duration, application crashes, blue screens, battery and disk health per device in our OpenSearch analytics. Trend over weeks; conspicuous devices are addressed proactively. |

## What you receive: Availability and alerts in the monthly report

Outages with duration and cause, alert counts by category, devices with conspicuous user experience — and what we did about it.

- Outages: system, start, duration, cause, action
- Alerts by category and criticality
- Response time on critical alerts
- Devices with DEX findings and status
- Thresholds adjusted this month

## Tools

- **RMM-Plattform** — Agent monitoring, thresholds, automations, remote access. Data location: EU instance. Replaceable by: any RMM with monitoring.
- **Serviceportal** — Alert becomes ticket; escalation, history per device. Data location: operated by us, Germany. Replaceable by: data portable via export/API.
- **OpenSearch (eigene DEX-Auswertung)** — User experience per device: sign-in, crashes, health, trends. Data location: operated by us, EU. Replaceable by: open format, exportable.
- **Wazuh** — Log collection and detection (see Managed SIEM). Data location: operated by us, EU. Replaceable by: open source — rules and data are yours.

## What we deliberately do not do

- No "24/7 monitoring" as a promise without agreed on-call — watching is around the clock, responding is at the agreed times
- No alert without a recipient and an instruction
- No collection of user data beyond what operations need — DEX measures devices, not people

## Prerequisites on your side

- RMM agent on all devices
- SNMP or API access to firewall, switches, Wi-Fi
- Escalation matrix: who at the client is reachable when
- Agreement on on-call outside service hours — or deliberately none

## Questions from IT

**Who gets woken up at night when the server is down?**
That depends on the agreed on-call. Without on-call: nobody — the alert sits at the top as a critical ticket at service start. With on-call: the person on duty at our end, and per escalation matrix your contact too. We state this clearly in the quote instead of writing "24/7" on the website.

**How do you prevent alerts from becoming wallpaper?**
By follow-up: every false alert is reviewed and the threshold or condition adjusted. Maintenance windows are muted, brief fluctuations filtered. The goal is that every alert that arrives triggers an action.

**Can we see the monitoring ourselves?**
Yes — in the service portal: device status, open alerts, ticket history per device. Your IT staff get their own access, on request via SSO with your Entra ID.

**What is digital employee experience in concrete terms?**
Measurements that describe how a device feels to the person in front of it: sign-in duration, application start, crashes, blue screens, battery, disk. We analyse them over weeks and address conspicuous devices before someone complains — or suffers quietly.

