Skip to main content
Monitoring must detect user-visible failure before routine clinical work is materially affected. Send production alerts to a staffed channel with an escalation path.

Required checks

Dokploy operational views

Use Dokploy service monitoring, deployment history, real-time logs, and per-service terminal access. Limit dashboard access to named operators and protect it with HTTPS and strong authentication.

Daily operations

  • Review overnight backup status and backup age.
  • Review container state, restarts, application errors, failed jobs, and scheduler execution.
  • Check database and disk capacity trends.
  • Review OIE failed or quarantined messages and radiology reconciliation queues.
  • Confirm production simulators remain stopped.
  • Record active incidents and unresolved warnings at shift handover.

Weekly operations

  • Review slow growth in queues, databases, logs, and imaging storage.
  • Confirm certificate renewal and DNS health.
  • Review failed login, privilege, token, and administrative events.
  • Test a synthetic critical workflow in the approved production-monitoring manner.
  • Confirm documentation, channel exports, contacts, and downtime materials remain current.

Monthly maintenance

1

Review capacity

Forecast CPU, memory, database, system disk, and imaging storage for at least the next 12 months.
2

Patch staging

Apply operating-system, Dokploy, container, application, and dependency updates in staging. Re-run acceptance tests and document compatibility.
3

Schedule production patching

Use the release runbook, a current recovery point, and a maintenance window. Do not use floating image upgrades without staging evidence.
4

Review access and secrets

Remove stale accounts and tokens, confirm owners, and rotate any credential that reached its policy date or was exposed.

Useful application checks

Run in the app service terminal:
Review logs in Dokploy rather than downloading broad production logs. Sanitize patient data and credentials before escalation.

Alert ownership

Every alert needs a severity, response time, primary owner, backup owner, runbook link, escalation contact, and closure evidence. Repeated alerts must produce a corrective action rather than permanent acknowledgement.