A service is not production-ready merely because it returns 200 on your laptop. Operators need to know whether the process is alive, whether it can receive traffic, where latency is accumulating, and which log lines belong to one request across several services.

Spring Boot Actuator and Micrometer provide those signals, but exposing every management endpoint publicly is not observability—it is an incident waiting to happen. The goal is useful signals with a deliberately small attack surface.

This architecture-and-lab guide builds that baseline: health probes, Prometheus metrics, W3C traceparent propagation, a human-friendly request ID, correlated logs, and locked-down Actuator endpoints.

Four signals, four different jobs

Do not collapse every operational question into /actuator/health.

Question                         Signal
------------------------------   ------------------------------
Is the JVM process alive?        Liveness health
Should traffic reach it now?     Readiness health
Is it fast and error-free?       Metrics
What happened to one request?    Traces + correlated logs

These signals work together:

Client
  |
  | traceparent + X-Request-Id
  v
Gateway ---- request duration / status metrics
  |
  v
Order service ---- structured logs with traceId + requestId
  |                    |
  |                    +---- trace backend
  +---- /actuator/prometheus ---- Prometheus
  +---- /actuator/health/*  ---- orchestrator

Metrics tell you that the checkout error rate increased. A trace identifies the slow downstream span. Correlated logs explain the validation or business failure inside that request.

Add the observability dependencies

Start with a Spring Boot 3 application on Java 17 or newer. This lab uses Spring MVC, Actuator, Prometheus, and Micrometer Tracing with Brave:

<dependencies>
  <dependency>
    <groupId>org.springframework.boot</groupId>
    <artifactId>spring-boot-starter-web</artifactId>
  </dependency>

  <dependency>
    <groupId>org.springframework.boot</groupId>
    <artifactId>spring-boot-starter-actuator</artifactId>
  </dependency>

  <dependency>
    <groupId>io.micrometer</groupId>
    <artifactId>micrometer-registry-prometheus</artifactId>
  </dependency>

  <dependency>
    <groupId>io.micrometer</groupId>
    <artifactId>micrometer-tracing-bridge-brave</artifactId>
  </dependency>
</dependencies>

Use the versions managed by Spring Boot’s dependency management. A tracing bridge gives the application a tracer and W3C propagation support; add the reporter/exporter for your chosen backend separately when you are ready to ship spans.

Note: Actuator uses Micrometer Observation underneath. You do not need to hand-code timers for standard Spring MVC requests—Boot already records http.server.requests.

Configure a small management surface

Use explicit endpoint exposure rather than include: "*". A practical local configuration is:

spring:
  application:
    name: orders

management:
  endpoints:
    web:
      exposure:
        include: health,info,prometheus
  endpoint:
    health:
      show-details: when_authorized
      probes:
        enabled: true
        add-additional-paths: true
  tracing:
    sampling:
      probability: 1.0

logging:
  pattern:
    correlation: "[${spring.application.name:},%X{traceId:-},%X{spanId:-},%X{requestId:-}] "

This exposes:

  • /actuator/health for aggregate health.
  • /actuator/health/liveness and /livez for liveness.
  • /actuator/health/readiness and /readyz for readiness.
  • /actuator/prometheus for Prometheus scraping.
  • /actuator/info for intentionally published build/application information.

The additional /livez and /readyz paths use the main server port. That matters when management endpoints run on a separate port: the main HTTP connector can fail while the management connector remains healthy.

Sampling every trace is a lab setting. In production, choose a probability that matches traffic volume and cost, or use tail-based sampling in your telemetry pipeline.

Liveness is not readiness

Liveness answers: “Can this process recover without a restart?” If it reports down, Kubernetes or another supervisor may restart the process.

Readiness answers: “Can this instance serve traffic now?” If it reports down, the orchestrator should stop routing new requests to it without necessarily restarting it.

Do not put databases, queues, or remote APIs in liveness. A database outage would restart every healthy application instance, increasing load during an already bad event.

A useful policy is:

Liveness  -> internal process state only
Readiness -> dependencies required to serve this instance
Health    -> broader diagnostic view for authorized operators

Spring Boot publishes liveness and readiness application availability states. Add a custom readiness contributor only when a dependency truly determines whether this instance can accept work.

Add a focused health indicator

Suppose order creation depends on writable disk space in a spool directory. A focused indicator can report that capability:

package com.geekmonks.orders;

import java.nio.file.Files;
import java.nio.file.Path;
import org.springframework.boot.actuate.health.Health;
import org.springframework.boot.actuate.health.HealthIndicator;
import org.springframework.stereotype.Component;

@Component
class OrderSpoolHealthIndicator implements HealthIndicator {

    private final Path spool = Path.of("/tmp/orders-spool");

    @Override
    public Health health() {
        boolean writable = Files.isDirectory(spool) && Files.isWritable(spool);

        return writable
                ? Health.up().withDetail("spool", "writable").build()
                : Health.down().withDetail("spool", "not writable").build();
    }
}

Its component endpoint is /actuator/health/orderSpool. Health details can expose hostnames, dependency names, or failure messages, so keep show-details: when_authorized.

Note: A custom HealthIndicator joins the aggregate health endpoint. It does not automatically join a probe group. Add it to readiness only after deciding that losing this capability should remove the instance from service:

management:
  endpoint:
    health:
      group:
        readiness:
          include: readinessState,orderSpool

Metrics describe system behavior

With Actuator and the Prometheus registry present, Boot publishes JVM, process, system, datasource, and HTTP server metrics. Run the application and inspect a few:

curl -sS http://localhost:8080/actuator/metrics/http.server.requests
curl -sS http://localhost:8080/actuator/prometheus \
  | grep '^http_server_requests_seconds'

The HTTP metric includes tags such as method, status, outcome, and the templated route. Those dimensions support questions like:

  • What is the 5xx rate?
  • Which route has the worst p95 latency?
  • Did latency change after the latest deployment?

Add one business metric

Infrastructure metrics say whether the service is healthy. A low-cardinality business counter says whether it is doing useful work:

package com.geekmonks.orders;

import io.micrometer.core.instrument.Counter;
import io.micrometer.core.instrument.MeterRegistry;
import org.springframework.stereotype.Service;

@Service
class OrderService {

    private final Counter createdOrders;

    OrderService(MeterRegistry registry) {
        this.createdOrders = Counter.builder("orders.created")
                .description("Orders accepted by the service")
                .tag("channel", "api")
                .register(registry);
    }

    void create() {
        // Validate and persist the order first.
        createdOrders.increment();
    }
}

After the service creates an order:

curl -sS http://localhost:8080/actuator/metrics/orders.created
curl -sS http://localhost:8080/actuator/prometheus \
  | grep '^orders_created_total'

Use tags for bounded sets such as channel=api|batch or result=accepted|rejected.

Never tag metrics with request IDs, user IDs, email addresses, or raw URLs. Every unique tag combination creates another time series. That cardinality can exhaust the monitoring system and your budget. Put per-request values in traces and logs instead.

W3C Trace Context uses a traceparent header:

traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01

The long value is the trace ID, the shorter value is the current parent span ID, and the final flags include the sampling decision. Tracing instrumentation creates a new span for each hop while retaining the trace ID.

An X-Request-Id header is an application convention. It is convenient for support tickets and can remain stable at the edge, but it does not encode span relationships or sampling. A sound setup carries both:

traceparent   -> machine-readable distributed trace context
X-Request-Id -> human-friendly correlation token

The Spring Cloud Gateway guide already places cross-cutting edge concerns at the gateway. Generate or validate X-Request-Id there, then forward it. Each service should preserve an incoming safe value or generate one when called directly.

Put the request ID in the response and MDC

This Servlet filter makes the request ID available in logs and echoes it to the caller:

package com.geekmonks.orders;

import jakarta.servlet.FilterChain;
import jakarta.servlet.ServletException;
import jakarta.servlet.http.HttpServletRequest;
import jakarta.servlet.http.HttpServletResponse;
import java.io.IOException;
import java.util.UUID;
import org.slf4j.MDC;
import org.springframework.stereotype.Component;
import org.springframework.web.filter.OncePerRequestFilter;

@Component
class RequestIdFilter extends OncePerRequestFilter {

    static final String HEADER = "X-Request-Id";

    @Override
    protected void doFilterInternal(
            HttpServletRequest request,
            HttpServletResponse response,
            FilterChain chain) throws ServletException, IOException {

        String requestId = normalize(request.getHeader(HEADER));
        response.setHeader(HEADER, requestId);

        try (MDC.MDCCloseable ignored = MDC.putCloseable("requestId", requestId)) {
            chain.doFilter(request, response);
        }
    }

    private static String normalize(String candidate) {
        if (candidate != null && candidate.matches("[A-Za-z0-9._-]{1,128}")) {
            return candidate;
        }
        return UUID.randomUUID().toString();
    }
}

Validation is important because a caller-controlled header will enter logs and responses. Restrict its length and character set rather than copying arbitrary input.

The try block removes the MDC value even when request handling fails. For work moved to another executor or reactive pipeline, MDC does not automatically follow the task; use the framework’s context-propagation facilities instead of manually copying thread locals everywhere.

Let instrumented clients propagate traceparent

Spring Boot configures observations for web requests and for HTTP clients built from its managed builders. Inject RestClient.Builder instead of calling RestClient.create():

package com.geekmonks.orders;

import org.springframework.stereotype.Component;
import org.springframework.web.client.RestClient;

@Component
class InventoryClient {

    private final RestClient client;

    InventoryClient(RestClient.Builder builder) {
        this.client = builder
                .baseUrl("http://inventory:8080")
                .build();
    }

    boolean isAvailable(String sku) {
        return Boolean.TRUE.equals(client.get()
                .uri("/api/inventory/{sku}", sku)
                .retrieve()
                .body(Boolean.class));
    }
}

Using the auto-configured builder lets Boot attach observation interceptors. The outbound call receives the active trace context, so its traceparent continues the inbound trace.

Note: Creating a tracing library is not the same as exporting spans. Connect an OTLP, Zipkin, or other supported exporter to your tracing backend. Keep credentials and collector URLs in environment-specific configuration.

Lock Actuator down in production

Actuator can reveal environment details, configuration properties, thread dumps, loggers, mappings, and operational controls. A separate management port is useful segmentation, but it is not authentication.

Use layers:

  1. Expose only required endpoints.
  2. Bind or firewall the management port to a private operations network.
  3. Require authentication and an operations role for everything except the minimal health probe.
  4. Keep health details hidden from anonymous callers.
  5. Protect Prometheus with network policy, authentication, or both.

With Spring Security on the classpath, define an explicit management security chain:

package com.geekmonks.orders;

import static org.springframework.boot.actuate.autoconfigure.security.servlet.EndpointRequest.toAnyEndpoint;
import static org.springframework.boot.actuate.autoconfigure.security.servlet.EndpointRequest.toLinks;
import static org.springframework.boot.actuate.autoconfigure.security.servlet.EndpointRequest.to;

import org.springframework.boot.actuate.health.HealthEndpoint;
import org.springframework.context.annotation.Bean;
import org.springframework.context.annotation.Configuration;
import org.springframework.core.annotation.Order;
import org.springframework.security.config.Customizer;
import org.springframework.security.config.annotation.web.builders.HttpSecurity;
import org.springframework.security.web.SecurityFilterChain;

@Configuration
class ManagementSecurityConfig {

    @Bean
    @Order(1)
    SecurityFilterChain actuatorSecurity(HttpSecurity http) throws Exception {
        http.securityMatcher(toAnyEndpoint())
                .authorizeHttpRequests(auth -> auth
                        .requestMatchers(to(HealthEndpoint.class)).permitAll()
                        .requestMatchers(toLinks()).hasRole("OPS")
                        .anyRequest().hasRole("OPS"))
                .httpBasic(Customizer.withDefaults());
        return http.build();
    }
}

Add a second SecurityFilterChain for application routes; defining a management chain does not secure the rest of the application for you.

If probes use /livez and /readyz on the main port, permit those exact paths in the application chain. Do not permit /actuator/** broadly just to make a probe work.

For stronger isolation, move management traffic:

management:
  server:
    port: 9090
    address: 127.0.0.1

That example is suitable when a local sidecar or host agent performs scraping. In containers, 127.0.0.1 is usually invisible to a scraper in another pod or container; bind to the pod interface and enforce a network policy instead.

Run the lab

Create the spool directory, start the application, and exercise the endpoints:

mkdir -p /tmp/orders-spool
./mvnw spring-boot:run

Check the two probe semantics:

curl -i http://localhost:8080/livez
curl -i http://localhost:8080/readyz

Send a request with correlation headers:

curl -i http://localhost:8080/api/orders/42 \
  -H 'X-Request-Id: support-7f31' \
  -H 'traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01'

Confirm that:

  • the response includes X-Request-Id: support-7f31;
  • application logs include requestId, traceId, and spanId;
  • http.server.requests records the route, status, and duration;
  • an instrumented downstream client continues the trace.

Remove write access from the spool directory to test the custom health indicator, then restore it:

chmod -w /tmp/orders-spool
curl -sS http://localhost:8080/actuator/health
chmod +w /tmp/orders-spool

When running as a privileged user, the writable check may still succeed; use a non-root application user, as you should in production.

Production cheat sheet

Health:
  liveness  = process can recover without restart
  readiness = instance can accept traffic
  details   = when_authorized

Metrics:
  alert on rate, errors, duration, saturation
  use bounded tags
  never tag with request/user IDs

Correlation:
  traceparent  = distributed trace hierarchy
  X-Request-Id = support-friendly request token
  MDC          = always clear after the request
  HTTP clients = build from Boot-managed builders

Security:
  expose an allowlist, never "*"
  isolate the management network/port
  authenticate non-probe endpoints
  treat Prometheus and health details as protected data

Wrap-up

Actuator is the operational boundary of a Spring Boot service. Health tells an orchestrator what action to take, metrics reveal behavior over time, and trace plus request IDs connect one user’s path across gateway, services, and logs.

Keep those signals purposeful. Liveness must not depend on the database, metric tags must stay bounded, tracing must propagate through instrumented clients, and management endpoints must be allowlisted and protected. That baseline turns “the API is slow” into a question your system can actually answer.

Next optional step Turn validation 400s and domain misses into ProblemDetail bodies clients can parse. Problem Details and @ControllerAdvice