2026.10.0 — 2026-10-01
Added
List the replicas on the Studio health screen (eea468cb)
Settings → Studio health now has a Replicas section: each replica's host, version, state, how long ago it last checked in, and the clusters it owns, with the one that answered marked "This replica". Replicas seen in the last ten minutes are listed, including stopped ones and one that vanished without stopping, which reads "Gone: no heartbeat, presumed crashed" so a dead replica is visible. A replica that is gone or draining is marked Degraded, and so is the screen as a whole. The broker node figures, database pool and open streams are labelled as seen from the replica answering, since each replica measures its own.
GET /api/v1/system/healthgainsreplicasandansweringReplica. Nothing to configure.Stop a run from any replica and interrupt runs cleanly on shutdown (2c79df3d)
Stopping a bulk run or a transfer now works whichever replica receives the request. If the run is not executing on the replica that got the request, it tells every replica, and the one executing the run stops it. The stop endpoints answer 202 Accepted, with the same body as before, and still answer 409 when the run is not running.
When a replica shuts down it now gives the bulk runs and transfers it is executing up to 20 seconds to finish (
artemis-studio.ha.run-grace). A run still going then is asked to stop, and is recorded as interrupted with the progress it made, not as stopped by an operator. A transfer stays resumable, and a bulk run is not resumed. This happens after the replica has stopped taking traffic and before background jobs and the brokers are released.Recover only the runs of a replica that is gone (20472d6a)
A bulk run or a transfer now records the replica that executes it. Recovery used to mark every running run as interrupted whenever any Studio process started, which would cut off a run that another healthy replica was still executing. It is now a job that runs once at startup and every 30 seconds on one replica, and it interrupts only a run whose replica has stopped or has not sent a heartbeat within 15 seconds (a draining replica still counts as alive while it finishes its runs). What recovery does to an interrupted run is unchanged: its queues or messages keep their recorded state, its audit event is marked failed, and a transfer stays resumable. A replica that crashed mid-run therefore has its run recorded as interrupted within about a minute.
Nothing to configure. A run recorded before this change has no replica and is treated as orphaned, so it is interrupted at the first pass.
Share console query references and API rate windows across replicas (6397896b)
Two small stores that lived in one replica's memory now live in the database, so a round-robin load balancer with no sticky sessions behaves as one Studio. Both are UNLOGGED tables: a crash loses them, which costs a rerun of one console query or a fresh rate window and nothing else.
- SQL console: the reference the POST issues and the stream GET redeems (ADR-0064) was found only on the replica that issued it, so the stream often answered "no longer valid". It is now a row in sql_query_ticket, redeemed once with DELETE ... RETURNING. A reference is still single use and valid for a minute; one presented by another user or for another cluster is refused without being used up. Unredeemed rows are purged by the data lifecycle as "SQL query references".
- API tokens: the per-minute request limits of a token and of its owner are counted in api_request_window, one statement per request, so a limit of N per minute is N across all replicas, not N per replica. A key trims its own older minutes as it goes; what a key that stopped calling leaves is purged by the data lifecycle as "API request rate windows". The in-flight (concurrency) limit stays in memory and applies per replica. Each request now costs one database round trip.
The login throttle stays in memory (ADR-0144): the database account lock is the control that matters.
Run the same plugins on every replica (49f69124)
Installing, enabling, disabling or uninstalling a plugin on one replica now takes effect on all of them, within seconds and without a restart. Before, a plugin activated on one replica stayed unavailable on the others until they restarted, so a load balancer served a plugin's pages on some requests and 404 on the rest.
When a plugin's state changes, the replica announces it on the replica bus once the change has committed. Every other replica compares the plugin_install rows with what it runs:
- a plugin that is active and not running there, or running another artifact than the row names, is started;
- a plugin that is disabled, uninstalled, incompatible or purged is stopped;
- a plugin that is activating, needs a restart or failed is left as it is.
The replica that made the change already runs what the row says and does nothing. A replica that reconnects to the bus reconciles too, so a change announced while it was cut off is not missed. Starts on a peer show in the audit log as PLUGIN_REPLICA_START.
The manifest version a browser polls to notice a changed plugin set is now a hash of the active id@version set instead of a counter seeded at random, so a load balancer alternating between replicas that run the same plugins no longer makes the web UI see a new manifest on every other request.
Keep caches coherent across replicas (47aad415)
With several Studio replicas on one database, a change made on one replica is now served by every replica within two seconds. Nothing to configure.
Each write announces itself on the replica bus in the transaction that makes it, so peers hear about it only once it has committed:
- a setting changed or reset: every replica re-reads the overrides and pushes the new value to the components that hold it;
- a cluster deleted: every replica releases its broker sessions for it;
- a cluster or environment moved: every replica rebuilds the environment index that permission checks use;
- a governance rule written: every replica drops its cached policy;
- a session ended or an API token revoked: every replica closes that session's or token's event streams. The 10 s session check stays as the backstop.
A replica that reconnects to the bus after losing it re-reads the caches it owns, so a message sent in the gap is not lost for good.
The governance-policy-refresh job is gone: it only polled a version number every 30 s to catch rule changes made on another instance, and the policy signal replaces it. It no longer appears in the job list.
Take up cluster duties as soon as a replica is ready (92ac7d61)
A replica that becomes ready asks the others to rebalance at once, and checks for readiness every quarter second until then. A single replica starts scraping within a second of being ready, and a joining one takes its share without waiting for the others' next tick.
Share the split-brain verdict between replicas (85441fc9)
The split-brain verdict of each broker node is now stored on its row (
broker_node.split_brain: NONE, SUSPECTED or CRITICAL) by the replica that scrapes the cluster, in the same pass that records the node's HA state. Every replica reads it from there, so the topology and health views, the split-brain alert, and the guards that refuse a transfer or a configuration apply while a pair is split all agree whichever replica answers. It used to live in the memory of the replica that scraped, so a request served by another replica saw no split at all. A single replica behaves as before.Nothing to configure: the migration adds the column. After a cluster moves to another replica, the verdict needs one extra scrape cycle (about 5 seconds) to confirm again.
Hand a cluster's subscriptions and consumers to its new owner (d10614b8)
When a replica stops owning a cluster, whether it handed the cluster over, lost it or is shutting down, it now drops what it held for that cluster: Core notification subscriptions, capture and plugin-messaging consumers, message-index tails, and its scrape bookkeeping. Before, two replicas could hold the capture queue's single consumer against each other, or both subscribe to a broker's notifications and record every event twice. The replica that takes a cluster over now scrapes it straight away, so its state, events and subscriptions do not wait for the next tick.
Nothing to configure.
Run each cluster's scheduled broker work on its owner only (3d7f7cda)
With several replicas, every replica used to poll every cluster's brokers on its own timer, so the management requests per node multiplied with the replica count. The timed work now runs for a cluster only on the replica that owns it: scrape tiers A and B (and the alert evaluation, plugin metrics and Core subscription reconcile that follow them), configuration drift, setup review, request-reply sampling, plugin messaging and capture reconcile, message-index sampling, and flow sampling. Adding a replica no longer adds broker load, and the work spreads over the replicas. The slow tier C sweep and topology discovery already ran on one replica per tick and are unchanged, as is everything a user starts, which runs on the replica that got the request.
Nothing to configure. A single replica owns every cluster; after a start it takes them within one or two heartbeats (5 s each) of becoming ready.
Give every cluster one owner among the replicas (26d7b5f1)
With several Studio replicas on one database, each cluster now has exactly one owner, the replica that will run that cluster's broker duties (the next commits gate them). Owners are chosen by rendezvous hashing over the ready replicas, so clusters spread evenly and a replica joining or leaving moves only the clusters that must move. The owner holds a row of the new
cluster_leasetable and renews it everyartemis-studio.ha.heartbeat(5 s); a lease lastsartemis-studio.ha.ttl(15 s), so a crashed owner's clusters are taken over within 20 s. A replica that shuts down hands its clusters over at once, and a replica that has not renewed for two thirds of the ttl stops claiming ownership before its lease can be taken.Nothing to configure: the migration adds the table, and a single replica owns every cluster.
Resume the live stream after a reconnect or a rolling restart (064f6c91)
The console now remembers the id of the last event it received and presents it on every connection, so after a network drop or a restart of the replica it was connected to, it picks up at the right event on whichever replica answers, without gaps or repeats.
When the server announces that it is shutting down, the console reconnects at once, with no delay, and this does not count as a failed connection. When the server says frames were lost, or the console reconnects after a failure, it refetches the views of that cluster, since change signals sent while it was away are not replayed. The SQL console reruns its query on a new stream when its replica shuts down, instead of reporting a lost connection. A view that asks for no live topics, such as the events view with Live switched off, no longer opens a second stream.
Replay missed events on any replica and drain streams on shutdown (25bf5f85)
A stream client that reconnects now presents its last event id as a
lastEventIdquery parameter (theLast-Event-IDheader still works) and is served by whichever replica the load balancer picks. The replica registers the client, sends the persisted broker events it missed, then the live frames that arrived during the replay, without repeating an id. A client gone for longer than the replay cap (500 events) gets the next 500 events and then a namedresyncevent, meaning: refetch your views.A replica that is draining or stopped answers a new stream with 503, so the client retries elsewhere. On shutdown, before it closes the streams, a replica sends every stream client (including SQL console streams) a named
reconnectevent, telling it to reconnect at once.Fan stream events out to every Studio replica (accb8ad9)
With several replicas on one database, an event detected on one replica now reaches the stream clients of all of them within a second. Every frame is announced over the replicas' PostgreSQL notification bus and each replica delivers it to its own clients when it arrives, so a frame is never delivered twice on one replica.
Broker events (the events topic) are announced as their seq numbers only, after the write that stored them has committed; each replica loads the rows and sends them with the seq as the SSE id, so a client never sees an event before it is in the database. The derived topics (consumers, sessions, connections, queues) are still nudged only by the replica that wrote.
When a replica's bus connection comes back after a loss, it sends a named resync event to its stream clients, because frames sent in the gap are gone.
Plugins publish through SseHub exactly as before. A frame whose data is over about 7 KB arrives as the plain signal with no data, and the client refetches.
Add a high-availability compose template (0905f02e)
deploy/compose/compose.ha.yaml runs two Studio replicas behind HAProxy on one Postgres. HAProxy routes on /readyz with no session affinity, so a replica that starts or drains leaves the pool within about two seconds. The replicas get 60 s to drain on stop.
The dev and prod compose healthchecks now call /readyz instead of the aggregate /actuator/health, so a replica that is starting or draining reports unhealthy.
Serve /livez and /readyz and drain before shutting down (2419aff5)
Studio now answers
/livezand/readyzon the main port, for a load balancer or orchestrator. Readiness fails while the instance is starting (until the plugin boot sequence has finished), while it is draining, and once the bus to the other replicas has been gone for more than ten seconds. Neither probe depends on a broker, and the existingstudiohealth group stays out of both. Until migrations have run the port does not answer at all, so a startup probe fails as it should.Shutdown now begins with a drain: the instance reports not ready, tells the rest of Studio it is draining, and keeps serving for
artemis-studio.ha.drain-delay(default 5s, raise it above your load balancer's detection time) before anything is stopped. Each shutdown phase is allowed 30s (spring.lifecycle.timeout-per-shutdown-phase); if you stop Studio with a shorter grace period, such as astop_grace_periodunder 40s, lengthen it.Add a PostgreSQL notification bus between replicas (37c5f78d)
Replicas can now tell each other things. Each process keeps one dedicated connection to PostgreSQL that listens on the
studiochannel, and sends withpg_notifyinside the caller's own transaction, so a message goes out when that transaction commits and never when it rolls back. Astudio-busthread reads the notifications, pings the connection every 10 seconds and reconnects with a backoff; after a reconnect it publishes a localBusResumedevent so listeners reload what they cached.Messages are one of three shapes (a stream frame, a batch of broker event numbers, or a signal of kind and key) and the wire never names a class. A frame over 7,500 bytes is sent without its data and a warning is logged once per topic, because PostgreSQL caps a notification at 8,000 bytes.
The bus opens its connection from
spring.datasource.url,usernameandpasswordwith TCP keep-alive on; it takes one extra connection per Studio instance that the pool does not count.max_connectionson the database should allow for it. The PostgreSQL JDBC driver is now a compile dependency, which changes nothing in the image.Count crashes, not boots, toward the plugin crash-loop guard (f8d5e659)
The guard that starts Studio without plugins after repeated failures used to count boots in
studio_boot, so a second healthy instance looked like a crash. It now asks the replica registry: three replicas that ended without recording a stop, within fifteen minutes, trip safe mode. Running several instances side by side is no longer a trigger. A replica is Ready once the plugin boot sequence has finished, and it writes its stop as the very last step of shutdown.The migration drops
studio_boot; nothing to do when upgrading. The crash history starts empty, so the first upgrade never trips safe mode.Register every Studio process as a replica (507e6baa)
Each Studio process now records itself in a new
studio_replicatable with its host, version and state (starting, ready, draining, stopped), and renews a heartbeat every 5 seconds from database time, on a thread of its own so a busy job pool cannot make a healthy instance look dead. An instance whose heartbeat is older than 15 seconds is treated as gone. Tune both withartemis-studio.ha.heartbeatandartemis-studio.ha.ttl.Replicas that stopped or died more than a day ago are removed by the data lifecycle's new
replicasstore, which appears on the Data page with the other retention policies. Nothing to do when upgrading; the migration adds the table.
Fixed
Keep a replica that cannot start a plugin from overwriting its shared status (d819a917)
A replica that follows another replica's plugin change and fails to start the plugin now logs the failure and leaves the shared install record alone. Before, it marked the plugin failed for every replica, overwriting the outcome the replica that made the change had recorded.
Push a setting to its holder only when its value changed (48a8ca68)
A setting write refreshes the stored values on the writing replica and again when its announcement comes back on the bus. Every refresh pushed every value, and pushing the broker timeouts replaces every broker HTTP client, so a broker call already under way, such as registering a cluster, could fail with "Nothing answered at this address". A refresh now pushes only the values that changed.
Do not fail a caller when a stream signal cannot be published (6112bd14)
Since stream signals travel over the database, publishing one threw a DataAccessException into whatever asked for it (alert evaluation, reply correlation, connection control, message operations) whenever the database or its connection pool was unavailable, failing work that had nothing to do with the signal.
SseHub.publish now logs the failure, once per topic until a publish succeeds again, and carries on. A missed signal only delays a client's refresh: it refetches when it reconnects or is told to resync.
No savepoint guards the caller's transaction. pg_notify with a payload capped below the 8,000-byte limit and a fixed channel does not fail as a statement, so inside a transaction the only failure left is a lost connection, which ends that transaction whatever publish does, and its commit reports it. A savepoint would add two round trips to every publish inside a transaction to guard against something that cannot happen.
Nothing to do when upgrading.
Run one tier A pass per cluster at a time (9a6f5399)
When a replica took a cluster over it started a tier A pass on its own thread at once, which could overlap the scheduled pass for the same cluster. The two passes counted the cycle twice, scraped every node twice and reconciled the Core subscriptions twice, doubling the load on the brokers for that moment.
A tier A pass now runs only when none is already running for that cluster on the replica; the overlapping one is skipped, and the next scheduled pass runs as usual.
Nothing to do when upgrading.
Make stopping a run on another replica survive a lost signal (8d15755c)
Stopping a bulk run or a transfer that another replica executes was a bus signal and nothing else. While the executing replica's bus connection was down the signal was lost and the run kept going after the operator had stopped it.
The stop endpoint now also records the request on the run (stop_requested_at, added to bulk_run and transfer_run), still sending the signal as the fast path. The runner looks at the row before each queue or batch, and every couple of seconds while a transfer waits for room on its target, so the run stops within one queue or batch either way. A transfer that is resumed or returned starts without the stop asked of its previous segment.
A replica whose bus connection came back also releases the Core sessions, subscriptions and pools it still holds for a cluster that no longer exists, which a missed cluster-deleted signal would otherwise leave open.
Nothing to do when upgrading; the two columns are added on startup.
Stop a run that recovery took over instead of overwriting it (f88dae95)
Recovery interrupts a run whose replica looks gone, but the replica could still be executing it, for example after its heartbeat failed for longer than the ttl. The runner then kept calling the broker and overwrote the INTERRUPTED run and its items, and finished the run's audit event a second time.
Bulk runs and transfers now write their rows only while the run is still executing on this replica: each write takes the run's row lock first, in the same transaction, so it either commits before recovery interrupts the run or does not happen. When the run was taken over, the runner makes no further broker call, writes nothing and leaves the audit events as recovery recorded them. The queue or batch that was in flight when it happened stays recorded as unknown, so check the broker for it.
A replica also stops claiming liveness it cannot prove: a run checks, before each queue or batch, that this replica recorded a heartbeat within the ttl (artemis-studio.ha.ttl, 15 s by default), and stops if it has not. Recovery then marks it interrupted as for any other gone replica.
Nothing to do when upgrading.
Keep SSE writes and database reads off the studio bus thread (08dcbf1e)
A client that stopped reading its event stream could block the thread that reads PostgreSQL notifications. The notification queue then grew for the whole cluster, and lease, session-ended, run-stop and settings signals stalled on every replica.
The bus reader now only reads and parses notifications and hands them, in order, to a dispatch thread through an unbounded queue, so the connection always drains. Each SSE subscriber has its own outbound queue (1,000 frames) written by its own virtual thread, so a slow client delays nobody. A client that falls that far behind loses its queued frames to a single resync event and carries on; this is logged once. Buffered replay, frame order, the heartbeat and the drain-time reconnect behave as before. Announced broker events are loaded from the database on a single-thread executor, so batches still go out in seq order without the dispatch thread waiting for the database.
Nothing to do when upgrading.
Keep the replica marker whole and state a replica's condition in one word (58f025df)
The replicas table names the state in one word (Ready, Draining, Gone) with the explanation beneath it, and the replica and version columns no longer wrap, so the "This replica" marker is never cut short.
Answer a stop with 200 as the API contract states (a519c347)
A stop that reaches a run on another replica is still answered with the run as it stands and 200, as before, rather than 202.
Re-mask stored content once per installation, not on every replica (d6281c13)
Re-masking only rewrites rows in the shared database, so running it on every replica did the same batches several times over. It now runs on one replica per tick, like the other installation-wide jobs.
Download a support bundle from any replica (7d558943)
Preparing a support bundle and downloading it are two requests, and the preview was kept in the memory of the replica that prepared it. Behind a load balancer the download often reached another replica, which answered "support bundle not found". The prepared bundle is now stored in the database, in an unlogged table, for the ten minutes it can be downloaded, and every replica can serve it. Expired previews are purged by the data lifecycle as "Support bundle previews" (default 1 hour); the page that lists stores shows the new one.
The in-memory cap of 32 prepared bundles is gone: previews are bounded by their ten minute lifetime instead.
Announce broker events in the transaction that stores them, and replay the most recent (899ba498)
The announcement of a flushed batch is now part of the flush transaction, so Postgres sends it only if the rows commit. Before, it was sent after commit and reached the other replicas only because the connection happened to commit on release.
A client that reconnects after missing more events than the replay cap now gets the most recent ones, not the oldest, followed by a signal to refetch. The replay then runs straight on into live delivery with no gap.
Publish only the events a flush wrote, and only after it commits (09630ad7)
The live event stream could miss or repeat broker events when two Studio instances flushed at once, because each worked out "its" rows from the highest sequence number in the table. It now takes the sequence numbers from its own insert and sends them to the stream after the transaction commits, so a rolled-back flush publishes nothing. Nothing to do when upgrading.
Mint the instance id once when replicas first boot together (53038204)
Two Studio instances starting against an empty database could each generate a capture instance id and overwrite the other's, leaving one replica naming broker objects after an id the database no longer holds. The first insert now wins and every instance reads it back. Nothing to do when upgrading.