Anticipating a full disk and backups that fail silently
Forecast when a volume will fill up instead of suffering it, spot SQL Server databases with no backup or integrity check, and know whether last night's backup script really ran
Two failures arrive without warning even though they build up over weeks: the disk that fills up, and the backup that fails every night while nobody reads the report. The first stops a service on a Monday morning. The second is discovered on the day you need to restore.
The full disk
A threshold is not enough
An alert at 90 % comes too late for a volume that grows by 20 GB a day, and too early for a stable volume that has been 91 % full for two years. The right question is: at this rate, when will it be full?
Checking without a tool
- Record the free space of each fixed volume, at the same time, every day for two weeks.
- Work out the average growth per day, leaving out days of clean-up or extension.
- Divide the free space by that growth: you get the number of days left.
Get-Volume | Where-Object DriveType -eq 'Fixed' | Select-Object DriveLetter, FileSystemLabel, @{n='Free (GB)';e={[math]::Round($_.SizeRemaining/1GB,1)}}, @{n='Size (GB)';e={[math]::Round($_.Size/1GB,1)}}The most frequent causes: application logs that nothing purges, a local backup folder that piles up, user profiles on a remote desktop server, and the transaction log of a SQL Server database.
The transaction log
In the FULL recovery model, a database's transaction log only empties after a log backup. Without that backup, it grows until the disk is full. The log_reuse_wait_desc column says what prevents it from emptying: LOG_BACKUP means it is waiting for a log backup.
SELECT name, recovery_model_desc, log_reuse_wait_desc
FROM sys.databases
ORDER BY name;Choose knowingly. The FULL model lets you restore to a precise point in time and requires regular log backups. The second model empties the log by itself: it is enough if going back to the database's last backup suits you.
Backups that fail silently
The usual causes
- the SQL Server Agent job fails, and the failure email goes to a mailbox nobody reads;
- the SQL Server Agent service does not restart after an update, and jobs no longer run;
- the scheduled backup script stops on an error without reporting anything;
- a new database was created without being added to the backup plan.
Checking SQL Server databases
SELECT d.name, d.recovery_model_desc,
MAX(CASE WHEN b.type = 'D' THEN b.backup_finish_date END) AS last_full,
MAX(CASE WHEN b.type = 'L' THEN b.backup_finish_date END) AS last_log
FROM sys.databases d
LEFT JOIN msdb.dbo.backupset b ON b.database_name = d.name
WHERE d.name <> 'tempdb'
GROUP BY d.name, d.recovery_model_desc
ORDER BY last_full;A database with NULL in the last_full column has no backup known to this instance. A tool that backs up virtual machines may, however, copy the database without leaving a trace in msdb: check in that tool.
SELECT j.name, h.run_date, h.run_time, h.message
FROM msdb.dbo.sysjobhistory h
JOIN msdb.dbo.sysjobs j ON j.job_id = h.job_id
WHERE h.run_status = 0 AND h.step_id = 0
ORDER BY h.run_date DESC, h.run_time DESC;The integrity check (DBCC CHECKDB) matters as much as the backup: a backup of an already corrupted database restores a corrupted database. The date of the last successful check is in the dbi_dbccLastKnownGood field of DBCC DBINFO WITH TABLERESULTS.
Checking a backup script
For a scheduled script (file copy, export, backup of a device), reverse the logic. Instead of waiting for a failure message, wait for a success signal: the script calls an address at the end of its work, and an alert goes out if the call does not arrive within the expected time. A script that crashes, a disabled task or a server that is switched off then give the same result: no signal, so an alert.
Testing the restore
A backup is only worth something if it restores. RESTORE VERIFYONLY checks that the file is readable, not that the database comes back. Restore a database on a test instance at least once a quarter, and write down how long it takes: that is your real recovery time.
What FirstSI follows
| Question | Screen | What opens an incident |
|---|---|---|
| When will this volume be full? | Asset Monitor → Capacity (#/am/capacite), Volumes tab | full within 30 days (medium), 14 days (high), 7 days or already 95 % full (critical) |
| Will a database file reach its maximum size? | same screen, Databases tab | same thresholds, relative to the maximum size |
| Are backups and integrity checks up to date? | DB Monitor → Background jobs (#/dbm/taches) | database with no recent backup (2 days, differential included), log not backed up for 24 hours in FULL, CHECKDB older than 7 days |
| Do SQL Server Agent jobs succeed? | same screen | failure on the last run or 3 times in 7 days, SQL Server Agent service stopped |
| Did last night's script run? | HostMonitor Heartbeats | ping missing within the expected time |
The saturation date follows the method described above. Every day, FirstSI keeps the highest usage of the day, draws the line of the last 30 days and extends it to 90 % and 100 %. It needs 7 days of measurements. An extension or a big clean-up is recognised as a break, and the calculation restarts from the days that follow. See Capacity forecast.
If an external tool backs up your databases, tick Backed up by an external tool in the instance setting: the backup rule becomes a low finding, with no incident. See Database background jobs.
For a heartbeat, the script calls its Ping URL at the end of the work. With $ErrorActionPreference = 'Stop', a PowerShell error stops the script before the call:
$ErrorActionPreference = 'Stop'
# … backup commands …
Invoke-RestMethod -Uri "https://console.exemple.fr/api/heartbeat/<identifier>/ping" -Method PostAn external program (robocopy, for example) does not raise a PowerShell error: test its exit code before the call. See HostMonitor: service availability.
Volumes that will be full within 14 days, or are already 95 % full, also appear in the To handle list on the dashboard.
Further reading
- DBMonitor: databases: performance, sessions, blocking and business check queries.
- Alerts and notifications: being notified by email, Teams or Slack.
- Hardening your Windows servers: patches, antivirus and persistence on the same servers.
Overview: SQL Server and MariaDB performance monitoring
Source: · FirstSI Docs · updated 2026-10-11