2026.09.78 — 2026-09-30
Breaking
Label jobs with job_id and alert on stalled jobs (75b91ffa)
The
jobtag on studio.job and studio.job.lag collided with Prometheus's reservedjobtarget label, which a real server renames toexported_job. The tag is nowjob.id, exposed asjob_id. A new studio.job.degraded gauge (1 while no run has finished within three of the job's intervals) backs a StudioJobStalled alert, and the dashboard and guide follow.Tag broker meters by host:port and rename the request-reply latency metric (50fa1fca)
The rate limiter's meters (studio.broker.requests, studio.broker.permit.wait, studio.broker.permit.timeouts) tagged
nodewith the full Jolokia URL, which can carry credentials. They now carry the node's host:port only. The request-reply latency timerartemisstudio.rr.latencyis nowstudio.rr.latency, like the other Studio meters.Export metrics, traces and logs over OTLP (cf4afcb2)
Studio can now ship its metrics, traces and logs to any OpenTelemetry collector. It is off by default and sends nothing until you switch it on with the standard OpenTelemetry variables, which Spring Boot maps itself:
- OTEL_EXPORTER_OTLP_ENDPOINT names the collector (for example http://collector:4318); /v1/metrics, /v1/traces and /v1/logs are appended. OTEL_EXPORTER_OTLP_HEADERS, OTEL_TRACES_SAMPLER and the other OTEL_* settings work too.
- OTEL_TRACES_EXPORTER, OTEL_METRICS_EXPORTER and OTEL_LOGS_EXPORTER set to otlp turn on each signal.
One request is one trace: the HTTP handler, each Jolokia call (studio.broker.management), each Core operation (studio.broker.core: send, browse, sample, pool borrow, relay open and commit) and each database query are spans, and every scheduled job run is a root span. Sampling is 1.0; lower it with OTEL_TRACES_SAMPLER_ARG. Spans name the node as host:port and never carry message content, headers, management arguments or query parameters.
Exported spans, span events and logs are redacted like the console log. Set LOGGING_STRUCTURED_FORMAT_CONSOLE=ecs for JSON log lines, which carry the traceId of the exported span.
The studio.job timer gains Micrometer's error tag beside job and feature.
Added
Label jobs with job_id and alert on stalled jobs (75b91ffa)
The
jobtag on studio.job and studio.job.lag collided with Prometheus's reservedjobtarget label, which a real server renames toexported_job. The tag is nowjob.id, exposed asjob_id. A new studio.job.degraded gauge (1 while no run has finished within three of the job's intervals) backs a StudioJobStalled alert, and the dashboard and guide follow.Alert on unbounded thread growth (8694a347)
The shipped rules gain StudioThreadsHigh: more than 500 live threads for 15 minutes. Studio's threads are bounded by configuration, not by time, so steady growth points at a leak.
Ship Grafana dashboards and Prometheus alert rules (a87a1f10)
deploy/observability/ holds a Grafana dashboard (studio-overview.json, with node and job filters) and Prometheus alert rules (studio-alerts.yml): a job past its interval or failing, a broker node failing management calls, slow management calls, permit timeouts and a waiting or saturated database pool. They use the renamed metrics, where
nodeis host:port. ObservabilityAssetsIntegrationTest fails when a metric they query is no longer served on /actuator/prometheus.Show Studio's own health in Settings (04b3d7cf)
Adds GET /api/v1/system/health (settings-read) and a "Studio health" section in Settings that polls it every 5 seconds. It shows each background job's status and lag, each broker node's last success and failure, p95 management latency and rate-limit wait, the database pool, and the open event streams, with a degraded mark per row and for the view. A figure that cannot be read is null and shown as unavailable, never as zero.
New metrics: studio.job.lag (per job, seconds past its interval) and studio.stream.clients. studio.broker.management now publishes its 95th percentile.
Tag broker meters by host:port and rename the request-reply latency metric (50fa1fca)
The rate limiter's meters (studio.broker.requests, studio.broker.permit.wait, studio.broker.permit.timeouts) tagged
nodewith the full Jolokia URL, which can carry credentials. They now carry the node's host:port only. The request-reply latency timerartemisstudio.rr.latencyis nowstudio.rr.latency, like the other Studio meters.Export metrics, traces and logs over OTLP (cf4afcb2)
Studio can now ship its metrics, traces and logs to any OpenTelemetry collector. It is off by default and sends nothing until you switch it on with the standard OpenTelemetry variables, which Spring Boot maps itself:
- OTEL_EXPORTER_OTLP_ENDPOINT names the collector (for example http://collector:4318); /v1/metrics, /v1/traces and /v1/logs are appended. OTEL_EXPORTER_OTLP_HEADERS, OTEL_TRACES_SAMPLER and the other OTEL_* settings work too.
- OTEL_TRACES_EXPORTER, OTEL_METRICS_EXPORTER and OTEL_LOGS_EXPORTER set to otlp turn on each signal.
One request is one trace: the HTTP handler, each Jolokia call (studio.broker.management), each Core operation (studio.broker.core: send, browse, sample, pool borrow, relay open and commit) and each database query are spans, and every scheduled job run is a root span. Sampling is 1.0; lower it with OTEL_TRACES_SAMPLER_ARG. Spans name the node as host:port and never carry message content, headers, management arguments or query parameters.
Exported spans, span events and logs are redacted like the console log. Set LOGGING_STRUCTURED_FORMAT_CONSOLE=ecs for JSON log lines, which carry the traceId of the exported span.
The studio.job timer gains Micrometer's error tag beside job and feature.
Fixed
Show a job's and a node's health word in full and say a last run succeeded (d341da77)
The Studio health view cut its Healthy and Degraded badges to three letters in a narrow column, and called a job whose last run succeeded Running, which reads as in progress. The badges now keep their width and the status reads Last run succeeded.
Give Studio room for its telemetry libraries and plugins (901e16af)
Studio now carries the OpenTelemetry SDK, its log appender and JDBC observation, which put metaspace at about 150 MB before any plugin loads, against a 256 MB cap. Installing plugins could then exhaust metaspace and restart Studio.
- compose.prod.yaml: Studio's container limit rises from 1 GB to 2 GB and the metaspace cap from 256 MB to 384 MB (also in .env.example). If you set JAVA_OPTS yourself, raise -XX:MaxMetaspaceSize to 384m.
- compose.dev.yaml: heap 1.5 GB and metaspace 384 MB.
The cap stays, so a plugin that does not unload still fails loudly.
Tag Core calls with the node's management address and keep credentials out of it (093a6ed4)
Core observations were tagged with the Core port, so the
nodelabel differed from every other Studio meter. They now carry the node's management host:port, found from the Core URL among the registered nodes; the Core address stays as a span-only attribute. A URL whose user info held a comma, parenthesis or question mark leaked part of the credentials into the node label; the user info is now removed first. The Studio health table shows a missing p95 as unavailable.