2026.10.7 — 2026-10-01
Breaking
Keep each broker's management URL on its own node row (88519332)
Topology discovery matched a seed to a node row by NodeID and HA role and took the first by name. After a failover a primary and its backup share a NodeID and can both read PRIMARY, so a seed could land on its partner's row, or on a new row named after its management address that read the same broker as the connector-named one. The scrape then counted two live nodes in one pair and raised a false split-brain.
A management URL now identifies its row: discovery attaches a seed to the row that already holds the URL, else to a connector-named row without a URL (by host, then by NodeID and role), else creates one, and never moves a URL from one row to another. HA role and state stay observed and are written to that row.
Discovery also reads each seed on its own. A broker that does not answer is skipped (the scrape records its error) and the cluster's other brokers are still discovered, instead of the whole pass failing. The cached-name reachability check that let a down broker through is gone.
Fixed
Judge a cycle by its last HA read and never mark a node unreachable for Studio's own throttling (ecec365d)
A node whose HA read was held back by Studio's own per-node call ceiling no longer shows as unreachable: the node was never asked, so it keeps the state and error it had, and the held-back read is logged.
A broker that answers its HA read slower than a second now takes part in the split-brain check, the stream signals and the notification subscriptions of its own scrape cycle. Before, that cycle was judged without its reading, so a live node on a slow link could hide a split-brain. The pass still moves on after a second, so a stalled broker holds up nothing but its own cycle.
A node's repeated failure is logged again after its cluster has been handed to another instance and back, and the record of it no longer outlives the cluster.
Upgrading needs no action.
Refuse a node URL another node holds and keep an operator's pinned node on upgrade (bea57ce4)
Overriding a node's management URL to one that another node of the same cluster already uses is now refused with a 409 and the name of the URL, and no audit entry is left half open. Before, the request reached the database and failed as a server error.
The upgrade step that removes duplicate node rows now keeps the row an operator pinned when a duplicate group holds one, and only then prefers the row discovery named after the broker's connector. Before, a pinned row could be deleted in favour of a discovered one, losing the pinned URL.
Upgrading needs no action.
Keep storing broker events after a cluster is deleted (96cbaa88)
Deleting a cluster while some of its broker events were still waiting to be written stopped events from being stored for every cluster. The write was rejected because the cluster no longer existed, the batch was put back at the head of the buffer, and every later flush failed on the same batch until the buffer filled up and new events were dropped.
A deleted cluster's waiting events are now discarded as soon as the deletion is signalled. A batch the database rejects because a cluster or node it names no longer exists has those events dropped and counted in the per-cluster dropped total, and the rest of the batch is written on the next flush. The flush no longer fails the scheduled job for this. A batch that fails for any other reason, such as the database being down, is still kept in order and retried.
Keep each broker's management URL on its own node row (88519332)
Topology discovery matched a seed to a node row by NodeID and HA role and took the first by name. After a failover a primary and its backup share a NodeID and can both read PRIMARY, so a seed could land on its partner's row, or on a new row named after its management address that read the same broker as the connector-named one. The scrape then counted two live nodes in one pair and raised a false split-brain.
A management URL now identifies its row: discovery attaches a seed to the row that already holds the URL, else to a connector-named row without a URL (by host, then by NodeID and role), else creates one, and never moves a URL from one row to another. HA role and state stay observed and are written to that row.
Discovery also reads each seed on its own. A broker that does not answer is skipped (the scrape records its error) and the cluster's other brokers are still discovered, instead of the whole pass failing. The cached-name reachability check that let a down broker through is gone.
Say why a broker call failed and forget a broker's name when it changes (c7de065b)
A broker that was slow, refused the connection or did not resolve was always reported as "Nothing answered at this address", and the log gave no reason. The message now says which it was: the broker accepted the connection but did not answer within the read timeout, did not accept the connection in time, nothing is listening at host:port, the host name does not resolve, or the underlying error otherwise. The same text reaches the node's last error and the health report. A node that keeps failing the same way is logged once with its cause, and again when the reason changes, instead of on every tick.
Studio holding a call back because it was already calling a node at its configured rate is no longer reported as the node being unreachable. It has its own error kind, THROTTLED, and says that Studio throttled itself and the node was not asked.
The broker's JMX name, learned once per management URL, is now forgotten when a call to that URL fails to connect or the broker reports the MBean missing, so a restarted or replaced broker is resolved again instead of failing until Studio restarts.
Let only the reachability probe decide whether a node is reachable (8bab8b44)
A healthy broker could show as unreachable, and a cluster as degraded, because any failed scrape job wrote broker_node.last_error: a failed queue listing, a sweep page, or a database error while saving what was scraped. Only the HA read that runs every few seconds now sets or clears a node's reachability. Failures in the other jobs are logged and leave the node as it was.
A Jolokia answer that carries an error (for example InstanceNotFound, returned with HTTP 200) is now a failed observation. It no longer overwrites the node's HA state with unknown values or clears its last error. Topology discovery rejects such a seed the same way.
One stalled broker no longer holds up the HA read of the other nodes in its cluster: each node is probed on its own and never twice at once, so failover and failback are reported on time. The scrape tiers also run on a pool of their own instead of being installed as the scheduler for every scheduled job, so events flushing and other jobs no longer share threads with broker I/O. The shared jobs pool's threads are now named jobs-N instead of scrape-N.
Upgrading needs no action.