Skip to main content
Back to Read
AI Agents22 April 2026Updated 6 September 202611 min read

Hermes Agent Down? Monitoring, Logs, systemd & Production Recovery (2026)

A production recovery guide for Hermes Agent: gateway status, logs, systemd service checks, update receipts, backups, healthchecks and safe restart decisions.

A person inspecting a network cable beside a compact computer on a back-office shelf.
Illustrative scene.
  1. 01The recovery decision tree
  2. 02systemd: what to trust and what to avoid
  3. 03Monitoring that catches real failures
  4. 04Logs: keep enough, not everything
  5. 05Rollback and restore

If Hermes Agent is down, do not start by reinstalling it. Start by proving which layer is broken: the gateway process, the messaging bridge, the model provider, the update, the host or the workflow skill. Most outages become longer because someone restarts the wrong thing, wipes useful logs, or runs an update while the service is already unhealthy.

The order matters: status, logs, health check, update receipt, then restart. That keeps evidence intact and avoids turning a small gateway issue into a broken restore.

Last updated: 6 September 2026. Checked against Hermes messaging and updating docs on 6 September 2026, systemd service documentation, and the dated 40-day Ampliflow production measurement.

TL;DR:

  • First command: hermes gateway status. Second command: the correct journalctl command for the service type you installed.
  • Use the Hermes-installed gateway service. Do not add custom kill hooks or stale copied unit files unless you have inspected the generated service.
  • Restart=always helps a long-running gateway recover after process exit, but it does not fix bad credentials, broken config, exhausted quota, unavailable providers or failed conditions.
  • Current Hermes updates already support snapshots, --check, --plan, --backup, update receipts, validation and gateway restart checks. Build automation around those native receipts instead of a brittle custom updater.
  • Backups and migration should use hermes backup and hermes import, with restore tested on a separate host/profile.

If a business process depends on the agent, monitoring is not optional DevOps work. It is part of the product. For the full deployment path, read How to Deploy Hermes Agent, the Oracle Cloud guide, and Hermes security & GDPR.

Five-minute recovery sequence

Run these in order. Stop as soon as you identify the failing layer.

1. Check the gateway status

bashhermes gateway status

If you installed a system service:

bashsudo hermes gateway status --system

The official messaging docs expose these commands directly. Use them before reaching for raw process lists.

2. Read the gateway logs

For a user service:

bashjournalctl --user -u hermes-gateway -n 100 --no-pager
journalctl --user -u hermes-gateway -f

For a system service:

bashjournalctl -u hermes-gateway -n 100 --no-pager
journalctl -u hermes-gateway -f

You are looking for one of six signals:

Log signalLikely layerFirst action
Authentication or pairing failureMessaging bridgeRelink the channel or refresh the token
Model provider 401/403CredentialsRotate or repair the model-provider key
Model provider 429Rate limit/quotaWait, lower concurrency, or switch the configured route after testing
Import/config errorHermes config/updateRun `hermes doctor`, inspect config changes, then consider rollback
Port already in useDuplicate service/manual processStop the duplicate service cleanly
Service starts then exitsRuntime or configRead the first error above the final exit line

3. Run the health check

bashhermes doctor
hermes --version

hermes doctor checks the install and configuration. hermes --version tells you the CLI version on that shell; compare it with service logs or the update receipt if you need to prove what the running gateway loaded.

4. If it failed after an update, inspect the receipt

Current Hermes updates write machine-readable receipts:

bashls -lt ~/.hermes/logs/update_receipts/
cat ~/.hermes/logs/update_receipts/latest.json
tail -n 200 ~/.hermes/logs/update.log

Use the receipt to answer three questions:

  1. Which services did the updater detect?
  2. Did the gateway restart phase succeed?
  3. Is any gateway still serving old code?

That is better evidence than "it stopped around the time I updated".

5. Restart only after you understand what will restart

For a user service:

bashhermes gateway restart
hermes gateway status

For a system service:

bashsudo hermes gateway restart --system
sudo hermes gateway status --system

If you have both user and system services installed, clean up the duplicate before continuing. The Hermes docs warn that having both installed makes start/stop/status behaviour ambiguous.

The recovery decision tree

  1. 01systemd: what to trust and what to avoid
  2. 02Monitoring that catches real failures
  3. 03Logs: keep enough, not everything
  4. 04Updates: use the native safety rails first
  5. 05Rollback and restore
textHermes not responding
  |
  +-- hermes gateway status says active
  |     |
  |     +-- messaging logs failing -> relink WhatsApp/Telegram/Slack
  |     +-- model calls failing -> check provider status, key, quota
  |     +-- skill output missing -> check cron/schedule and skill logs
  |
  +-- gateway inactive or failed
  |     |
  |     +-- recent update -> inspect update receipt and update.log
  |     +-- config/import error -> run hermes doctor, fix config or rollback
  |     +-- port conflict -> stop duplicate gateway/manual backend
  |
  +-- host unreachable
        |
        +-- cloud console says stopped/reclaimed -> restore from backup
        +-- SSH blocked -> check firewall, security list, bastion/VPN
        +-- disk full -> free logs/snapshots carefully, then restart

Do not skip straight to reinstall. Reinstalling can overwrite evidence, complicate backups and mask the real cause.

systemd: what to trust and what to avoid

Hermes now installs and manages gateway services itself:

bashhermes gateway install
hermes gateway start
hermes gateway status

For a headless VPS, the current Hermes messaging docs recommend installing the gateway as a system service:

bashsudo hermes gateway install --system
sudo hermes gateway start --system
sudo hermes gateway status --system

If you choose a user service instead, enable lingering so the service can survive logout:

bashsudo loginctl enable-linger $USER

The current updating docs also note that user services can be restarted by hermes update without root prompts. Choose one service model deliberately and record it.

The older version of this page included a hand-written unit file. That is now the wrong default. The generated Hermes service already includes service-manager details such as restart behaviour and clean shutdown semantics. Copying an old unit can lose those updates.

The systemd point still matters: according to systemd's service documentation, Restart=always restarts a service after clean or abnormal process exit, while Restart=on-failure ignores a clean exit. Inspect the generated Hermes unit and use restart behaviour that matches your workload. It will not repair invalid credentials, a syntax error, missing environment variables, a failed condition, a stopped host, or an unavailable model provider.

Avoid custom ExecStopPost=/bin/kill -9 ... drop-ins. The Hermes docs specifically warn that this can fire during clean restarts and cause restart loops.

Monitoring that catches real failures

Monitor the layers a human needs to act on.

LayerCheckAlert whenNotes
HostCloud instance state, SSH reachability, disk usageHost stopped, SSH unavailable, disk lowEspecially important on free-eligible cloud capacity
Gateway`hermes gateway status` or systemd stateInactive, failed, restart loopUse the right user/system service command
MessagingChannel-specific bridge logs and test messagePairing expired or send failsWhatsApp relinks should be documented
Model providerKnown small prompt or provider status/API checkAuth, quota, rate-limit or 5xx failuresAvoid treating provider outages as Hermes bugs
Updates`update_receipts/latest.json` and `update.log`Update skipped, partial, mixed-version fleetCurrent Hermes gives you native evidence
BackupsRestore test timestampNo tested restore in the last review windowA created backup is only half the control

For small deployments, this can be a simple heartbeat plus a weekly runbook review. For regulated or customer-facing workflows, feed these checks into the monitoring system your team already uses.

Healthchecks without turning Hermes into a monitoring project

For a founder-operated setup, a heartbeat service is enough to start:

bashcurl -fsS -m 10 https://hc-ping.com/<uuid>

Run it independently from the host through a narrow cron job or systemd timer. A successful host heartbeat does not prove the Hermes gateway, model provider or business workflow is healthy; check those layers separately. Send failures to someone who owns recovery.

Test the alert by deliberately pausing the heartbeat for longer than the threshold. Document:

  • who receives it;
  • what message appears;
  • what command they run first;
  • who owns recovery when the first person is unavailable.

Untested monitoring is decoration.

Logs: keep enough, not everything

For a normal single-host setup, keep:

  • systemd journal for the gateway;
  • Hermes application logs;
  • messaging bridge logs;
  • update logs and receipts;
  • workflow-level output logs for scheduled jobs.

Rotate logs before they fill the disk. This example assumes the service account is named hermes: replace the absolute path for your host, choose retention deliberately and test the writer’s rotation behaviour. The copytruncate approach can lose lines during copying, so use application-supported reopen/rotation where available.

text/home/hermes/.hermes/logs/*.log {
    daily
    rotate 14
    compress
    delaycompress
    missingok
    notifempty
    copytruncate
}

Do not promise a universal retention period. Retention depends on the data, the business purpose, the lawful basis, the regulator and the contract. The security and GDPR guide covers the governance layer.

Updates: use the native safety rails first

The biggest change since the original article is the current hermes update behaviour.

The official updating docs now describe:

  • pre-update snapshots;
  • optional full HERMES_HOME backup with hermes update --backup;
  • hermes update --check for previewing whether an update is available;
  • hermes update --plan for a fleet/service restart plan without changing files;
  • syntax validation and auto-rollback for critical files after the code pull;
  • gateway restart checks;
  • update receipts under ~/.hermes/logs/update_receipts/;
  • hermes backup and hermes import for moving to a new machine.

So the default recommendation is simple:

bashhermes update --check
hermes update --plan
hermes update --backup
hermes doctor
hermes gateway status

For unattended updates, build a thin wrapper around those commands and the receipt file. Do not maintain a separate updater that guesses what Hermes already knows about profiles, services and gateway restarts.

Rollback and restore

There are two different recovery jobs.

JobUse thisDo not confuse it with
Bad code update`git checkout <commit-or-tag>`, reinstall dependencies, restart gateway, then run `hermes config check`Moving the whole agent to a new host
Bad state/config or host migration`hermes backup` and `hermes import`A quick update snapshot
Single profile export`hermes profile export`A full credential-bearing backup

The current Hermes docs explicitly say update backups protect the in-place update path, while moving to another machine should use hermes backup and hermes import.

For Oracle Free Tier or any low-cost VPS, practise this before you need it:

  1. Take a backup.
  2. Restore to an isolated test profile or throwaway VM with production schedules, credentials and outbound channels disabled.
  3. Inspect the restored state before starting a gateway. Do not run two copies of the same scheduled workload.
  4. Test with a non-production channel, then follow the backup and migration cutover guide.
  5. Record the exact commands in the runbook.

The Ampliflow 40-day incident, kept in context

Our dated measurement from 3 April to 13 May 2026 included a 62-hour outage. The gateway stopped after an exception path, the service behaviour did not bring it back, and nobody had an offsite alert watching the workflow. The full numbers are in the production cost teardown.

The lesson is not "Hermes is unreliable". The lesson is that a long-running agent needs the same boring controls as any production service:

  • managed service;
  • restart semantics that match the workload;
  • logs someone can read;
  • alert path someone receives;
  • backups that restore;
  • a runbook written before the outage.

Do not overfit to our incident. Your incident may be a dead WhatsApp pairing, a model-provider quota, a changed skill path, a full disk or a reclaimed free-tier VM.

One-page runbook

Keep this in the repo, the password manager note, or wherever the operator can find it during an outage.

textHermes recovery runbook

1. Check service:
   hermes gateway status
   sudo hermes gateway status --system   # if installed as system service

2. Read logs:
   journalctl --user -u hermes-gateway -n 100 --no-pager
   journalctl -u hermes-gateway -n 100 --no-pager   # system service

3. Check install:
   hermes doctor
   hermes --version

4. If update-related:
   cat ~/.hermes/logs/update_receipts/latest.json
   tail -n 200 ~/.hermes/logs/update.log

5. If service is stopped:
   hermes gateway restart
   hermes gateway status

6. If restart fails:
   check provider status, credentials, config, disk space and port conflict

7. If host/state is bad:
   restore from the latest tested hermes backup

8. After recovery:
   write down root cause, changed files/config, commands run and next prevention step

Frequently asked questions

Why is my Hermes gateway active but not replying?

Because the gateway process is only one layer. Check the messaging bridge, channel pairing, model-provider credentials, quota, tool failures and skill logs. An active service can still be unable to answer.

Where are Hermes gateway logs?

For a user service, start with journalctl --user -u hermes-gateway -f. For a system service, use journalctl -u hermes-gateway -f. Also check ~/.hermes/logs/update.log and ~/.hermes/logs/update_receipts/latest.json after updates.

Should I use Restart=always?

For the Hermes gateway, yes, use the generated Hermes service and verify it has the intended restart behaviour. Restart=always restarts after clean and abnormal exits. It does not fix bad config, provider failure, credentials or host-level problems.

Should I run automatic updates?

Only with preview, backup, receipt checks and a tested rollback path. Current Hermes gives you --check, --plan, --backup, receipts and post-update validation, so use those rather than an old custom script.

Can monitoring prevent every outage?

No. Monitoring reduces detection time. Recovery still depends on logs, backups, access, credentials and a person who knows the runbook.

What should I do if Oracle reclaims or stops the VM?

Confirm the host state in Oracle, then use an isolated replacement and the backup and migration guide. Inspect restored schedules and channels before starting the gateway, so the replacement cannot duplicate production jobs. Test cutover and update monitoring.

What should you do next?

If Hermes is down now, run the five-minute sequence at the top of this page and write down what failed. If you are planning a deployment, build the runbook before the first real workflow depends on the gateway.

Need help turning a Hermes experiment into an operated workflow? Get unstuck →

Hermes setup help

Deployment, skills and day-two reliability

Get help setting up your Hermes agent

We deploy, harden and maintain Hermes Agent for UK businesses — cloud hosting, gateways, skills, approvals, monitoring and recovery included.

Cloud deployment & hardening
WhatsApp, Slack & email
Safe, repeatable skills
Monitoring & recovery
Scope my Hermes setup

Bring the use case or the setup you already have. We will tell you the smallest sensible next step.