Maintenance clock
Every job on this estate that is not a CI pipeline — a cron, a systemd timer, a backup server’s own scheduler, a scheduler inside a container — is declared in one file with its host, its host-local start time and its measured duration. This panel renders that file. It is not a diagram of the intended schedule; it is the record the estate’s own clash reconciler reads, so the public clock and the internal one cannot drift apart.
Maintenance clock — 98 declared jobs across one day
Snapshot taken — build-time, not live- 98declared jobs
- 4,123starts a day
- 70watched by a dead-man
- 09:00busiest hour
cronin-container schedulernative schedulerothersystemd timer
Midnight at the top, clockwise, Australia/Melbourne (host-local AEST/AEDT). Each arc is one job at its declared start; jobs sharing a minute stack inward, so a crowded slot reads as depth. The ten sub-hourly jobs are the inner rings — they run all day and are stated as a rate rather than plotted 4,123 times.
The outer amber band is quiet hours, 21:00–07:00: a declared decision (ADR-0264) that nothing may page a human in it. The consequence is accepted deliberately — a real emergency born at 02:00 is not seen until morning — so anything that can fail overnight has to fail safe rather than fail loudly. The shaded sector is the overnight backup window, 22:00–07:00, which is not typed in anywhere: it is derived from the 4 daily backup jobs' own start times and measured durations.
| Time | Job | Host | Runs | Watched |
|---|---|---|---|---|
| */20 min | ADO CI reconciler (trigger-gap detection + pipeline heartbeats)cron · 1m | control01 | 72×/day | dead-man |
| */15 | ADO pipeline metrics collectioncron · 1m | control01 | 96×/day | dead-man |
| */5 min | External dead-man snitch pingercron · 1m | control01 | 288×/day | dead-man |
| */5 min | HA nginx-proxy add-on log shippercron · 1m | control01 | 288×/day | dead-man |
| :* | HA ownership anomaly detectorsystemd timer · 1m | control01 | 1440×/day | dead-man |
| */15 min | HA press-without-action detectorcron · 1m | control01 | 96×/day | dead-man |
| :* | cc-pool drain-request watchsystemd timer · 1m | control01 | 1440×/day | dead-man |
| :03, :18, :33, :48 | cc-pool in-flight marker refreshsystemd timer · 1m | control01 | 96×/day | dead-man |
| */15 (offset :07) | control01 shared repo checkout synccron · 1m | control01 | 96×/day | dead-man |
| :10, :40 | docker01 config-as-code early warningsystemd timer · 1m | control01 | 48×/day | none |
| :17 | Reboot advance-notice detectionsystemd timer · 1m | pve01, control01 | Hourly | dead-man |
| :20 | Dawn triage metrics collectioncron · 1m | control01 | Hourly | none |
| :23 | cc-pool fork-resume regression testsystemd timer · 2m | control01 | Hourly | dead-man |
| 01:30 | PVE backup all guestsnative scheduler · 50m | pve01 | Daily | none |
| 02:00 | ADO PAT rotation gatecron · 5m | control01 | Sun | none |
| 02:00 | Docker image/build-cache prunecron · 10m | docker01 | daily | dead-man |
| 02:00 | Media cleanupcron · 15m | pve01 | Daily | none |
| 02:30 | xt035 backup to xv035native scheduler · 30m | pve01 | Daily | none |
| 02:45 | pvesnap-prunecron · 5m | pve01 | Daily | none |
| 03:00 | rclone files→B2cron · 60m | docker01 | Daily | dead-man |
| 03:20 | Memory eviction sweepcron · 5m | control01 | Daily | dead-man |
| 03:30 | Infisical Postgres backupcron · 15m | pve01 | Daily | dead-man |
| 04:00 | Fleet patch+reboot cyclesystemd timer · 120m | control01 (orchestrates fleet ex-pve01) | Wed | dead-man |
| 04:00 | Reboot coordinator (guests)systemd timer · 8m | fleet (ex-pve01) | Daily | none |
| 04:10 | ADO pool-auth reconcilecron · 2m | control01 | Daily | dead-man |
| 04:15 | Reboot coordinator (pve01)systemd timer · 5m | pve01 | Daily | none |
| 04:30 | Infisical backup retention prunecron · 5m | pve01 | Daily | none |
| 05:00 | PBS local GCnative scheduler · 155m | xt035 | Sun | none |
| 06:00 | KB draft agentcron · 10m | control01 | Daily | dead-man |
| 06:00 | PBS B2 syncnative scheduler · 60m | xt035 | Daily | none |
| 06:00 | bun3d restore checkcron · 5m | control01 | Daily | none |
| 06:30 | PVE patchingsystemd timer · 30m | pve01 | Sat | none |
| 06:50 | Dawn triage (autonomous actuator)cron · 25m | control01 | Daily | dead-man |
| 07:00 | Morning health digestcron · 25m | control01 | Daily | dead-man |
| 07:20 | Delivery snag review digestcron · 5m | control01 | Mon | dead-man |
| 07:30 | Reboot deadline countdownsystemd timer · 1m | pve01, control01 | Daily | dead-man |
| 07:45 | Schedule reconcilercron · 3m | control01 | Daily | dead-man |
| 07:50, 13:50, 17:50 | GitHub scheduled-workflow freshness watchcron · 2m | control01 | Daily | dead-man |
| 08:00 | vectormap embedding-map rendersystemd timer · 3m | control01 | Mon | none |
| 08:05 | Hypercare window countdowncron · 1m | control01 | Daily | dead-man |
| 08:10 | dated-review watchcron · 1m | control01 | Daily | dead-man |
| 08:15 | Blog post daily count sync (Zabbix)cron · 1m | control01 | Daily | dead-man |
| 08:15 | Monitoring-as-Code reconcilercron · 2m | control01 | Daily | dead-man |
| 08:15 | Pi-hole tunnel-origin pin auditcron · 2m | control01 | Daily | none |
| 08:15 | Standards adherence gap analysiscron · 2m | control01 | Monthly | dead-man |
| 08:20 | DT SBOM-freshness watchdogcron · 1m | control01 | Daily | dead-man |
| 08:25, 14:25, 20:25 | Host/LXC synthetic test-plan gatecron · 2m | control01 | Daily | dead-man |
| 08:30 | Disabled-trigger zombie reapercron · 2m | control01 | Daily | dead-man |
| 08:45 | Entra SSO client-secret expiry watchdogcron · 2m | control01 | Daily | dead-man |
| 08:55 | Claude OAuth refresh-token expiry watchdogcron · 1m | control01 | Daily | dead-man |
| 09:00 | HA integration inventory drift-checkcron · 2m | control01 | Daily | dead-man |
| 09:00 | Top-charts list rebuildin-container scheduler · 2m | docker01 | Daily | dead-man |
| 09:00 | vulnscan weekly DevSecOps loopcron · 120m | control01 | Mon | dead-man |
| 09:10 | HA long-lived token dead-mancron · 1m | control01 | Daily | dead-man |
| 09:12 | smtp STARTTLS cert renewalcron · 1m | docker01 | Daily | none |
| 09:15 | NetBox asset classificationcron · 3m | control01 | Sun | dead-man |
| 09:20 | Claude OAuth refresh-token regime verdict remindercron · 1m | control01 | Daily | none |
| 09:30 | HA release-watchercron · 4m | control01 | Wed | dead-man |
| 09:30 | Year-planner term-date feed refreshin-container scheduler · 1m | docker01 | Sun | dead-man |
| 09:35 | HA entity-name conformance checkcron · 1m | control01 | Thu | dead-man |
| 09:40 | Year-planner RMIT next-year availability checkin-container scheduler · 1m | docker01 | Daily | dead-man |
| 09:45 | External attack-surface scancron · 3m | control01 | Monthly | dead-man |
| 09:50 | Fleet auto-update conformance reconcilersystemd timer · 2m | control01 | Daily | dead-man |
| 10:00 | Decom Phase 2 destroycron · 5m | control01 | Daily | dead-man |
| 10:00 | KB draft review digestcron · 5m | control01 | Mon | dead-man |
| 10:00 | PBS B2 prunenative scheduler · 30m | xt035 | Sun | none |
| 10:00 | PBS verifynative scheduler · 30m | xt035 | Tue | none |
| 10:20 | *arr import-list drift checkcron · 2m | control01 | Sun | dead-man |
| 10:30 | PBS B2 offsite restore-test (monthly recoverability proof)cron · 15m | control01 | Thu | dead-man |
| 10:30 | Restore drill (monthly recoverability proof)cron · 20m | control01 | Wed | dead-man |
| 10:40 | Claude MCP config reconcilersystemd timer · 1m | control01 | Daily | none |
| 10:50 | Agent permission-set auditcron · 1m | control01 | Daily | dead-man |
| 11:00 | PBS B2 GCnative scheduler · 30m | xt035 | Sun | none |
| 11:00 | Standards gap-analysis report (agent judgement layer)cron · 15m | control01 | Monthly | dead-man |
| 11:00 | bun3d weekly backupnative scheduler · 55m | pve01 | Wednesday | none |
| 11:15 | HA proxy TLS leaf renewalcron · 1m | control01 | Daily | none |
| 11:15 | Weekly CI performance reportcron · 5m | control01 | Mon | dead-man |
| 11:25 | Upstream release watch (accepted-risk expiry)cron · 2m | control01 | Daily | dead-man |
| 11:35, 15:35, 19:35 | Host→repo stack-source reconcilesystemd timer · 1m | control01 | Daily | dead-man |
| 11:40 | *arr stalled-grab watchdogcron · 3m | control01 | Daily | dead-man |
| 11:50 | Maintainerr retention rule evaluationin-container scheduler · 10m | docker01 | Daily | dead-man |
| 12:10 | CI shared-state PBS backupother · 2m | pve01 | Daily | dead-man |
| 12:30 | PBS backup-root config drift-exportcron · 2m | control01 | Daily | dead-man |
| 12:50 | Maintainerr retention deletion passin-container scheduler · 20m | docker01 | Daily | dead-man |
| 13:00 | Sleep quiet-hours stuck-mute canarycron · 1m | control01 | Daily | dead-man |
| 13:00 | Snapshot-tier appliance config exportcron · 3m | control01 | Daily | dead-man |
| 13:20 | arronpitman.com build-freshness watchdogcron · 1m | control01 | Daily | none |
| 13:40 | Vulnerability SLA conformance checkcron · 5m | control01 | Daily | dead-man |
| 14:10 | Open dependency-PR age checkcron · 2m | control01 | Daily | dead-man |
| 14:40 | Claude config docs reconciliationcron · 5m | control01 | Daily | dead-man |
| 14:50 | Service catalog reconcilecron · 5m | control01 | Daily | none |
| 15:00 | servicemap dependency-graph rendersystemd timer · 3m | control01 | Daily | dead-man |
| 15:30 | Backlog grooming debt checkcron · 1m | control01 | Monthly (15th) | dead-man |
| 16:10 | waitfor synthetic + freshness dead-mancron · 2m | control01 | Daily | dead-man |
| 16:40 | Stuck-alert watchdogcron · 2m | control01 | Daily | dead-man |
| 17:20 | agent-access discovery sweepcron · 2m | control01 | Daily | dead-man |
| 20:00 | Claude-pod podcast episodecron · 10m | control01 | Mon, Thu | none |
| 22:00 | PBS local prunenative scheduler · 30m | xt035 | Daily | none |
No job matches that filter.
Every job in the estate that is not a CI pipeline is declared in one file, and this panel renders that file — so the public clock and the record the estate's own clash reconciler reads cannot drift apart. 70 of 98 carry a dead-man watch, which is the check that matters for scheduled work: a job that stops running produces no error and no alert, only silence that looks exactly like success. The busiest hour is 09:00 with 26 distinct jobs — counted by job rather than by start, because an every-five-minute job would otherwise declare every hour the busiest and answer a question nobody asked.
The two bands
Section titled “The two bands”Quiet hours are a declared decision. Between 21:00 and 07:00 nothing may page a human. That is not a measurement and it is not tuned — it is a recorded architecture decision, and the price is accepted knowingly: a real emergency born at 02:00 is not seen until morning. That price is only payable because the things which can fail overnight are built to fail safe — they halt, roll back, or refuse to promote, rather than degrading further while nobody is watching. The design obligation that follows is the useful part: anything new that could get worse unattended does not get a louder alarm, it gets an automated remediation or a safe stop.
The backup window is derived. It is not typed in anywhere. It is computed from the start times and measured durations of the jobs whose own declaration says they are backup work, taking daily work only — because one guest is deliberately backed up at 11:00 on a Wednesday precisely to keep it out of the nightly window, so it is exactly the job that must not define it. A schedule change moves the band with it, rather than leaving a hand-typed pair of times to go quietly stale.
What the crowding shows
Section titled “What the crowding shows”The face stacks jobs inward when they share a minute, so a busy slot reads as depth. The morning peak is real and it is the interesting part of the picture: the overnight window is reserved for backup and patching, so everything else — reconcilers, digests, drift audits, conformance sweeps — is pushed into the hours after it. When no clean slot exists for a new job, that pressure is itself a finding worth surfacing, not a reason to cram the job into the least-bad gap.
Why the dead-man column matters more than it looks
Section titled “Why the dead-man column matters more than it looks”A scheduled job that stops running produces no error, no exception and no alert. It produces silence, and silence is what success looks like too. That is why the record tracks a dead-man watch per job — a freshness check that fires when the job’s own evidence goes stale — and why the count of jobs without one is worth publishing rather than smoothing over.