Hermes Agent Down? Monitoring, Logs, systemd & Production Recovery (2026)
A production recovery guide for Hermes Agent: gateway status, logs, systemd service checks, update receipts, backups, healthchecks and safe restart decisions.

- 01The recovery decision tree
- 02systemd: what to trust and what to avoid
- 03Monitoring that catches real failures
- 04Logs: keep enough, not everything
- 05Rollback and restore
If Hermes Agent is down, do not start by reinstalling it. Start by proving which layer is broken: the gateway process, the messaging bridge, the model provider, the update, the host or the workflow skill. Most outages become longer because someone restarts the wrong thing, wipes useful logs, or runs an update while the service is already unhealthy.
The order matters: status, logs, health check, update receipt, then restart. That keeps evidence intact and avoids turning a small gateway issue into a broken restore.
Last updated: 6 September 2026. Checked against Hermes messaging and updating docs on 6 September 2026, systemd service documentation, and the dated 40-day Ampliflow production measurement.
TL;DR:
- First command:
hermes gateway status. Second command: the correctjournalctlcommand for the service type you installed. - Use the Hermes-installed gateway service. Do not add custom kill hooks or stale copied unit files unless you have inspected the generated service.
Restart=alwayshelps a long-running gateway recover after process exit, but it does not fix bad credentials, broken config, exhausted quota, unavailable providers or failed conditions.- Current Hermes updates already support snapshots,
--check,--plan,--backup, update receipts, validation and gateway restart checks. Build automation around those native receipts instead of a brittle custom updater. - Backups and migration should use
hermes backupandhermes import, with restore tested on a separate host/profile.
If a business process depends on the agent, monitoring is not optional DevOps work. It is part of the product. For the full deployment path, read How to Deploy Hermes Agent, the Oracle Cloud guide, and Hermes security & GDPR.
Five-minute recovery sequence
Run these in order. Stop as soon as you identify the failing layer.
1. Check the gateway status
bashhermes gateway statusIf you installed a system service:
bashsudo hermes gateway status --systemThe official messaging docs expose these commands directly. Use them before reaching for raw process lists.
2. Read the gateway logs
For a user service:
bashjournalctl --user -u hermes-gateway -n 100 --no-pager
journalctl --user -u hermes-gateway -fFor a system service:
bashjournalctl -u hermes-gateway -n 100 --no-pager
journalctl -u hermes-gateway -fYou are looking for one of six signals:
| Log signal | Likely layer | First action |
|---|---|---|
| Authentication or pairing failure | Messaging bridge | Relink the channel or refresh the token |
| Model provider 401/403 | Credentials | Rotate or repair the model-provider key |
| Model provider 429 | Rate limit/quota | Wait, lower concurrency, or switch the configured route after testing |
| Import/config error | Hermes config/update | Run `hermes doctor`, inspect config changes, then consider rollback |
| Port already in use | Duplicate service/manual process | Stop the duplicate service cleanly |
| Service starts then exits | Runtime or config | Read the first error above the final exit line |
3. Run the health check
bashhermes doctor
hermes --versionhermes doctor checks the install and configuration. hermes --version tells you the CLI version on that shell; compare it with service logs or the update receipt if you need to prove what the running gateway loaded.
4. If it failed after an update, inspect the receipt
Current Hermes updates write machine-readable receipts:
bashls -lt ~/.hermes/logs/update_receipts/
cat ~/.hermes/logs/update_receipts/latest.json
tail -n 200 ~/.hermes/logs/update.logUse the receipt to answer three questions:
- Which services did the updater detect?
- Did the gateway restart phase succeed?
- Is any gateway still serving old code?
That is better evidence than "it stopped around the time I updated".
5. Restart only after you understand what will restart
For a user service:
bashhermes gateway restart
hermes gateway statusFor a system service:
bashsudo hermes gateway restart --system
sudo hermes gateway status --systemIf you have both user and system services installed, clean up the duplicate before continuing. The Hermes docs warn that having both installed makes start/stop/status behaviour ambiguous.
The recovery decision tree
- 01systemd: what to trust and what to avoid
- 02Monitoring that catches real failures
- 03Logs: keep enough, not everything
- 04Updates: use the native safety rails first
- 05Rollback and restore
textHermes not responding
|
+-- hermes gateway status says active
| |
| +-- messaging logs failing -> relink WhatsApp/Telegram/Slack
| +-- model calls failing -> check provider status, key, quota
| +-- skill output missing -> check cron/schedule and skill logs
|
+-- gateway inactive or failed
| |
| +-- recent update -> inspect update receipt and update.log
| +-- config/import error -> run hermes doctor, fix config or rollback
| +-- port conflict -> stop duplicate gateway/manual backend
|
+-- host unreachable
|
+-- cloud console says stopped/reclaimed -> restore from backup
+-- SSH blocked -> check firewall, security list, bastion/VPN
+-- disk full -> free logs/snapshots carefully, then restartDo not skip straight to reinstall. Reinstalling can overwrite evidence, complicate backups and mask the real cause.
systemd: what to trust and what to avoid
Hermes now installs and manages gateway services itself:
bashhermes gateway install
hermes gateway start
hermes gateway statusFor a headless VPS, the current Hermes messaging docs recommend installing the gateway as a system service:
bashsudo hermes gateway install --system
sudo hermes gateway start --system
sudo hermes gateway status --systemIf you choose a user service instead, enable lingering so the service can survive logout:
bashsudo loginctl enable-linger $USERThe current updating docs also note that user services can be restarted by hermes update without root prompts. Choose one service model deliberately and record it.
The older version of this page included a hand-written unit file. That is now the wrong default. The generated Hermes service already includes service-manager details such as restart behaviour and clean shutdown semantics. Copying an old unit can lose those updates.
The systemd point still matters: according to systemd's service documentation, Restart=always restarts a service after clean or abnormal process exit, while Restart=on-failure ignores a clean exit. Inspect the generated Hermes unit and use restart behaviour that matches your workload. It will not repair invalid credentials, a syntax error, missing environment variables, a failed condition, a stopped host, or an unavailable model provider.
Avoid custom ExecStopPost=/bin/kill -9 ... drop-ins. The Hermes docs specifically warn that this can fire during clean restarts and cause restart loops.
Monitoring that catches real failures
Monitor the layers a human needs to act on.
| Layer | Check | Alert when | Notes |
|---|---|---|---|
| Host | Cloud instance state, SSH reachability, disk usage | Host stopped, SSH unavailable, disk low | Especially important on free-eligible cloud capacity |
| Gateway | `hermes gateway status` or systemd state | Inactive, failed, restart loop | Use the right user/system service command |
| Messaging | Channel-specific bridge logs and test message | Pairing expired or send fails | WhatsApp relinks should be documented |
| Model provider | Known small prompt or provider status/API check | Auth, quota, rate-limit or 5xx failures | Avoid treating provider outages as Hermes bugs |
| Updates | `update_receipts/latest.json` and `update.log` | Update skipped, partial, mixed-version fleet | Current Hermes gives you native evidence |
| Backups | Restore test timestamp | No tested restore in the last review window | A created backup is only half the control |
For small deployments, this can be a simple heartbeat plus a weekly runbook review. For regulated or customer-facing workflows, feed these checks into the monitoring system your team already uses.
Healthchecks without turning Hermes into a monitoring project
For a founder-operated setup, a heartbeat service is enough to start:
bashcurl -fsS -m 10 https://hc-ping.com/<uuid>Run it independently from the host through a narrow cron job or systemd timer. A successful host heartbeat does not prove the Hermes gateway, model provider or business workflow is healthy; check those layers separately. Send failures to someone who owns recovery.
Test the alert by deliberately pausing the heartbeat for longer than the threshold. Document:
- who receives it;
- what message appears;
- what command they run first;
- who owns recovery when the first person is unavailable.
Untested monitoring is decoration.
Logs: keep enough, not everything
For a normal single-host setup, keep:
- systemd journal for the gateway;
- Hermes application logs;
- messaging bridge logs;
- update logs and receipts;
- workflow-level output logs for scheduled jobs.
Rotate logs before they fill the disk. This example assumes the service account is named hermes: replace the absolute path for your host, choose retention deliberately and test the writer’s rotation behaviour. The copytruncate approach can lose lines during copying, so use application-supported reopen/rotation where available.
text/home/hermes/.hermes/logs/*.log {
daily
rotate 14
compress
delaycompress
missingok
notifempty
copytruncate
}Do not promise a universal retention period. Retention depends on the data, the business purpose, the lawful basis, the regulator and the contract. The security and GDPR guide covers the governance layer.
Updates: use the native safety rails first
The biggest change since the original article is the current hermes update behaviour.
The official updating docs now describe:
- pre-update snapshots;
- optional full
HERMES_HOMEbackup withhermes update --backup; hermes update --checkfor previewing whether an update is available;hermes update --planfor a fleet/service restart plan without changing files;- syntax validation and auto-rollback for critical files after the code pull;
- gateway restart checks;
- update receipts under
~/.hermes/logs/update_receipts/; hermes backupandhermes importfor moving to a new machine.
So the default recommendation is simple:
bashhermes update --check
hermes update --plan
hermes update --backup
hermes doctor
hermes gateway statusFor unattended updates, build a thin wrapper around those commands and the receipt file. Do not maintain a separate updater that guesses what Hermes already knows about profiles, services and gateway restarts.
Rollback and restore
There are two different recovery jobs.
| Job | Use this | Do not confuse it with |
|---|---|---|
| Bad code update | `git checkout <commit-or-tag>`, reinstall dependencies, restart gateway, then run `hermes config check` | Moving the whole agent to a new host |
| Bad state/config or host migration | `hermes backup` and `hermes import` | A quick update snapshot |
| Single profile export | `hermes profile export` | A full credential-bearing backup |
The current Hermes docs explicitly say update backups protect the in-place update path, while moving to another machine should use hermes backup and hermes import.
For Oracle Free Tier or any low-cost VPS, practise this before you need it:
- Take a backup.
- Restore to an isolated test profile or throwaway VM with production schedules, credentials and outbound channels disabled.
- Inspect the restored state before starting a gateway. Do not run two copies of the same scheduled workload.
- Test with a non-production channel, then follow the backup and migration cutover guide.
- Record the exact commands in the runbook.
The Ampliflow 40-day incident, kept in context
Our dated measurement from 3 April to 13 May 2026 included a 62-hour outage. The gateway stopped after an exception path, the service behaviour did not bring it back, and nobody had an offsite alert watching the workflow. The full numbers are in the production cost teardown.
The lesson is not "Hermes is unreliable". The lesson is that a long-running agent needs the same boring controls as any production service:
- managed service;
- restart semantics that match the workload;
- logs someone can read;
- alert path someone receives;
- backups that restore;
- a runbook written before the outage.
Do not overfit to our incident. Your incident may be a dead WhatsApp pairing, a model-provider quota, a changed skill path, a full disk or a reclaimed free-tier VM.
One-page runbook
Keep this in the repo, the password manager note, or wherever the operator can find it during an outage.
textHermes recovery runbook
1. Check service:
hermes gateway status
sudo hermes gateway status --system # if installed as system service
2. Read logs:
journalctl --user -u hermes-gateway -n 100 --no-pager
journalctl -u hermes-gateway -n 100 --no-pager # system service
3. Check install:
hermes doctor
hermes --version
4. If update-related:
cat ~/.hermes/logs/update_receipts/latest.json
tail -n 200 ~/.hermes/logs/update.log
5. If service is stopped:
hermes gateway restart
hermes gateway status
6. If restart fails:
check provider status, credentials, config, disk space and port conflict
7. If host/state is bad:
restore from the latest tested hermes backup
8. After recovery:
write down root cause, changed files/config, commands run and next prevention stepFrequently asked questions
Why is my Hermes gateway active but not replying?
Because the gateway process is only one layer. Check the messaging bridge, channel pairing, model-provider credentials, quota, tool failures and skill logs. An active service can still be unable to answer.
Where are Hermes gateway logs?
For a user service, start with journalctl --user -u hermes-gateway -f. For a system service, use journalctl -u hermes-gateway -f. Also check ~/.hermes/logs/update.log and ~/.hermes/logs/update_receipts/latest.json after updates.
Should I use Restart=always?
For the Hermes gateway, yes, use the generated Hermes service and verify it has the intended restart behaviour. Restart=always restarts after clean and abnormal exits. It does not fix bad config, provider failure, credentials or host-level problems.
Should I run automatic updates?
Only with preview, backup, receipt checks and a tested rollback path. Current Hermes gives you --check, --plan, --backup, receipts and post-update validation, so use those rather than an old custom script.
Can monitoring prevent every outage?
No. Monitoring reduces detection time. Recovery still depends on logs, backups, access, credentials and a person who knows the runbook.
What should I do if Oracle reclaims or stops the VM?
Confirm the host state in Oracle, then use an isolated replacement and the backup and migration guide. Inspect restored schedules and channels before starting the gateway, so the replacement cannot duplicate production jobs. Test cutover and update monitoring.
Related reading
- ↑ How to Deploy Hermes Agent: UK Business Complete Guide — the deployment baseline this recovery page assumes.
- ↔ Hermes Agent on Oracle Cloud Free Tier — hosting choices and free-tier risk.
- ↔ Hermes Agent Security & GDPR — controls around logs, access and retention.
- ↔ Hermes Agent Production Cost Teardown — the dated incident and cost evidence.
- ↔ Hermes Agent Backup & Restore: Move Server Without Losing Memory — the migration companion article.
- External: Hermes messaging gateway docs, Hermes updating docs, systemd service documentation.
What should you do next?
If Hermes is down now, run the five-minute sequence at the top of this page and write down what failed. If you are planning a deployment, build the runbook before the first real workflow depends on the gateway.
Need help turning a Hermes experiment into an operated workflow? Get unstuck →