CVE-2026-101088 Overview
CVE-2026-101088 is a race condition vulnerability in Nezha, an open-source server and website monitoring tool. The flaw affects versions >= 2.2.11 and < 2.3.1 within the service sentinel worker at service/singleton/servicesentinel.go. It represents an incomplete fix for a previously reported nil dereference denial of service tracked as GHSA-qjpp-gffx-2wm9. An authenticated user holding the member role can trigger a panic that crashes the entire Nezha instance. The vulnerability is classified under [CWE-367: Time-of-Check Time-of-Use (TOCTOU) Race Condition].
Critical Impact
An authenticated member-role user who owns an agent can crash the entire Nezha monitoring instance by issuing a concurrent server delete request, resulting in a full service outage.
Affected Products
- Nezha versions >= 2.2.11
- Nezha versions < 2.3.1
- Nezha service sentinel worker component (service/singleton/servicesentinel.go)
Discovery Timeline
- 2026-09-27 - CVE-2026-101088 published to NVD
- 2026-09-30 - Last updated in NVD database
Technical Details for CVE-2026-101088
Vulnerability Analysis
The vulnerability stems from an incomplete remediation of a prior nil pointer dereference issue (GHSA-qjpp-gffx-2wm9). The 2026-07-21 fix attempted to re-validate the service lifecycle under serviceResponseDataStoreLock. However, the fix reused an already-captured, now stale reporter pointer and never re-validated the server. The lock also does not guard the ServerShared structure. This creates a narrow race window exploitable by a concurrent server deletion request.
When the sentinel worker reads the server list snapshot after a successful concurrent delete, it dereferences a missing entry. The sentinel workers and the gRPC server lack any recover() call or recovery interceptor. The resulting panic propagates unrecovered and crashes the entire instance, producing a denial-of-service condition.
Root Cause
The root cause is a classic time-of-check to time-of-use defect. The server pointer is captured, then validated at one point in time, then used later without a repeat check. During that interval, a legitimate authenticated API call can delete the referenced server. Combined with the absence of panic recovery in long-running goroutines, a single nil dereference brings down the full monitoring service.
Attack Vector
Exploitation requires an authenticated user with the member role who owns at least one monitored agent. The attacker issues a POST /api/v1/batch-delete/server request targeting their own server while the sentinel worker is actively processing a report for that server. Winning the race window causes the worker goroutine to panic on a nil map or slice lookup against the stale server list. See the GitHub Security Advisory GHSA-jx78-55p5-rwv5 and the VulnCheck Denial of Service Advisory for additional technical context.
Detection Methods for CVE-2026-101088
Indicators of Compromise
- Unexpected Nezha process restarts or crash loops, particularly immediately following POST /api/v1/batch-delete/server requests.
- Go runtime panic stack traces in Nezha logs referencing service/singleton/servicesentinel.go or nil pointer dereferences within sentinel worker goroutines.
- Member-role API tokens issuing repeated batch delete calls against servers they own within short time windows.
Detection Strategies
- Correlate HTTP access logs for /api/v1/batch-delete/server endpoints with process termination events on Nezha hosts.
- Alert on repeated gRPC agent reconnection storms, which can indicate the dashboard process has restarted after an induced panic.
- Monitor for authenticated member accounts generating abnormal rates of server create/delete churn within short intervals.
Monitoring Recommendations
- Enable structured logging on the Nezha dashboard and forward panic traces to a centralized log store for retention and search.
- Instrument the host with process uptime and crash metrics, alerting when the Nezha binary exits unexpectedly.
- Track audit events for the batch-delete/server endpoint and establish a baseline for normal administrative activity.
How to Mitigate CVE-2026-101088
Immediate Actions Required
- Upgrade all Nezha dashboard instances to version 2.3.1 or later, which contains the complete fix.
- Audit existing member-role accounts and revoke access for users who do not require agent ownership.
- Place the Nezha dashboard behind authenticated reverse-proxy controls that restrict exposure of the management API to trusted networks.
Patch Information
The issue is fixed in Nezha version 2.3.1. The corrected code re-validates both the reporter and the server pointer under the appropriate lock scope and addresses the unguarded ServerShared access path. Refer to the GitHub Security Advisory GHSA-jx78-55p5-rwv5 for upstream commit details and release notes.
Workarounds
- Restrict the member role so untrusted users cannot register agents or invoke batch-delete/server until the upgrade is applied.
- Deploy the Nezha process under a supervisor such as systemd with automatic restart on failure to reduce outage duration while patching.
- Rate-limit the /api/v1/batch-delete/server endpoint at the reverse proxy layer to narrow the exploitable race window.
# Example systemd override to auto-restart Nezha after an induced panic
# /etc/systemd/system/nezha-dashboard.service.d/override.conf
[Service]
Restart=always
RestartSec=5s
StartLimitBurst=10
StartLimitIntervalSec=60
Disclaimer: This content was generated using AI. While we strive for accuracy, please verify critical information with official sources.