03-data-access

HikariCP: Why a Bigger Pool Is Usually Slower

What each Hikari timeout controls, how to size a pool with arithmetic rather than guesswork, and how to see exhaustion coming before it becomes an outage.

October 9, 2026
spring-boothikaricpconnection-pooldatasourcetuningperformancedatabase

The Resource Everything Contends For

Opening a database connection is expensive — a TCP connection, authentication, TLS, session setup. Tens of milliseconds, for work you'd otherwise repeat on every query.

So connections are pooled: a fixed set is opened at startup and handed out on request. HikariCP is Spring Boot's default and needs no configuration to work. It needs configuration to work well, and the defaults are the reason most pool problems look like mysterious latency rather than obvious errors.

What Each Setting Does

yaml
spring:
  datasource:
    url: jdbc:postgresql://db.internal:5432/shop
    username: ${DATABASE_USER}
    password: ${DATABASE_PASSWORD}
    hikari:
      maximum-pool-size: 10        # default 10
      minimum-idle: 10             # default = maximum-pool-size
      connection-timeout: 30000    # 30s — how long a thread waits for a connection
      idle-timeout: 600000         # 10m — when idle connections are retired
      max-lifetime: 1800000        # 30m — hard cap on any connection's age
      keepalive-time: 0            # disabled by default
      pool-name: shop-pool
SettingControlsGuidance
maximum-pool-sizeTotal connectionsSee the arithmetic below — smaller than you'd think
minimum-idleConnections kept warmLeave equal to max for steady load
connection-timeoutWait before giving upLower it. 30s is far too long
idle-timeoutRetire idle connectionsOnly relevant if minimum-idle < max
max-lifetimeMaximum connection ageMust be below any network/DB idle timeout
keepalive-timeProbe idle connectionsUseful behind aggressive firewalls

Two of these deserve immediate attention.

connection-timeout: 30000 is too high. It means a request thread waits half a minute for a connection before failing. Under exhaustion, that converts a fast failure into 30 seconds of held request threads — the thread pool fills, and now every endpoint is affected. A timeout of 2–5 seconds fails fast and sheds load, which is what you want:

yaml
connection-timeout: 3000

max-lifetime must be shorter than the shortest idle timeout between you and the database. Cloud load balancers, firewalls and Postgres's own idle_session_timeout all close connections unilaterally. If Hikari believes a connection is alive and the network has dropped it, the next query fails with a broken-pipe error. Setting max-lifetime a minute or two below that external limit means Hikari retires connections before anything else can.

⚠️

AWS RDS Proxy, many Kubernetes ingress setups and most cloud NAT gateways close idle TCP connections after a few minutes. The symptom is intermittent connection-reset errors with no pattern, usually after quiet periods. Check what idle timeout sits between your app and your database, then set max-lifetime below it. This one setting explains a large share of "random" database errors.

Sizing: Smaller Than You Think

The instinct is that more connections mean more throughput. Past a point, the opposite is true.

A database executes queries on a finite number of cores. Connections beyond that don't execute in parallel — they queue, and the queueing has cost: context switching, lock contention, and on Postgres, a process per connection with its own memory. A pool of 100 against an 8-core database produces worse throughput than a pool of 20, plus much worse latency variance.

HikariCP's own guidance, from PostgreSQL's:

text
connections = ((core_count × 2) + effective_spindle_count)

For an 8-core database on SSD, that's roughly 17. For most services, a pool of 10–20 is correct, and Spring Boot's default of 10 is a reasonable starting point rather than a placeholder to raise.

Count every consumer

The database has a global connection ceiling (max_connections, often 100 on a small instance). Your pool is not the only claimant:

text
8 app instances × 20 connections  = 160
+ 2 background workers × 10       =  20
+ migrations, admin tools, monitoring ≈ 10
                                    ----
                                    190  ← against max_connections = 100

This fails during a rolling deployment, when old and new instances are briefly both running — the worst possible moment. Total pool size across all instances must fit within max_connections, with headroom for deployment overlap and a human needing to connect.

🚨

Reserve capacity for administrative access. A database at its connection limit cannot be connected to in order to diagnose why — including by you, during the incident. Postgres's superuser_reserved_connections keeps a few back; make sure your arithmetic doesn't consume them.

Pool Exhaustion

The failure mode to recognise:

text
HikariPool-1 - Connection is not available, request timed out after 30000ms.

Every connection is checked out and nothing was returned within the timeout. The important point: this is almost never solved by a bigger pool. It's a symptom, and the causes are specific:

1. A connection held across network I/O. The transaction anti-pattern from the transactions guide — an HTTP call inside @Transactional pins a connection for the call's duration. A few slow calls starve the pool.

2. A leak. A connection obtained and never closed. Hikari can find these for you:

yaml
spring:
  datasource:
    hikari:
      leak-detection-threshold: 20000    # warn if held > 20s

It logs a stack trace of the acquiring code — the fastest route to the offending line. Keep it on in non-production; the cost is small enough for production too.

3. Genuinely slow queries. A query taking 5 seconds holds its connection for 5 seconds. The fix is the query (or an index), not the pool.

4. More concurrency than the pool supports. The real case for a bigger pool — but check the first three before assuming it, and remember the database's ceiling.

⚠️

Virtual threads make this sharper. spring.threads.virtual.enabled: true removes the thread-pool ceiling that was implicitly limiting your database concurrency, so a service that previously queued at 200 threads now sends thousands of requests at a pool of 10. The bottleneck moves to the pool, and connection-timeout becomes the setting that decides whether you shed load or collapse. Enabling virtual threads is a reason to revisit pool configuration, not to ignore it.

Watching the Pool

With Actuator and Micrometer, Hikari publishes metrics that tell you everything:

MetricMeaning
hikaricp.connections.activeCurrently in use
hikaricp.connections.idleAvailable
hikaricp.connections.pendingThreads waiting
hikaricp.connections.usageHow long connections are held
hikaricp.connections.acquireHow long acquisition takes
hikaricp.connections.timeoutAcquisition failures

pending is the one to alert on. Sustained above zero means threads are waiting for connections — the leading indicator of exhaustion, visible well before the timeouts start. Alert on it rather than on the resulting errors.

usage is the diagnostic companion: if connections are held for seconds, something is doing I/O or slow work inside a transaction, and that's the cause to fix.

Actuator's health endpoint includes a DataSourceHealthIndicator that runs a validation query:

yaml
management:
  endpoint:
    health:
      show-details: when-authorized
  health:
    db:
      enabled: true
✅

Think about whether your readiness probe should include the database. If it does, a brief database blip marks every instance unready and takes the whole service out of the load balancer — turning a partial degradation into a total outage. For an app that cannot function without its database, that may be correct; for one that serves cached or static responses, it isn't. Phase 6 covers the liveness/readiness distinction properly.

Check yourself

An app on 8 instances with maximum-pool-size: 25 runs against Postgres with max_connections = 100. It works normally but fails during deployments with 'too many clients already'. Why?

Multiple DataSources

When an app talks to two databases, auto-configuration can't guess which is primary, so you declare both:

java
@Configuration
public class DataSourceConfig {
 
    @Bean
    @Primary
    @ConfigurationProperties("app.datasource.orders")
    DataSourceProperties ordersProperties() {
        return new DataSourceProperties();
    }
 
    @Bean
    @Primary
    @ConfigurationProperties("app.datasource.orders.hikari")
    DataSource ordersDataSource(DataSourceProperties ordersProperties) {
        return ordersProperties.initializeDataSourceBuilder()
                .type(HikariDataSource.class).build();
    }
 
    @Bean
    @ConfigurationProperties("app.datasource.reporting")
    DataSourceProperties reportingProperties() {
        return new DataSourceProperties();
    }
 
    @Bean
    @ConfigurationProperties("app.datasource.reporting.hikari")
    DataSource reportingDataSource(DataSourceProperties reportingProperties) {
        return reportingProperties.initializeDataSourceBuilder()
                .type(HikariDataSource.class).build();
    }
}

@Primary is what keeps the rest of auto-configuration working — JPA, Flyway and Actuator all need to know which DataSource is the default. This is the @Primary/@Qualifier disambiguation from phase 1 in its most common real use.

Each needs its own pool budget, and both count against the server's limit. Also note that a single @Transactional cannot span two DataSources atomically — that needs XA transactions, which are a significant complication and usually a sign the operation should be redesigned.

A Starting Configuration

Reasonable defaults for a typical service, to be adjusted from measurements:

yaml
spring:
  datasource:
    hikari:
      maximum-pool-size: 15
      minimum-idle: 15
      connection-timeout: 3000
      max-lifetime: 1500000          # 25m — below typical infra idle timeouts
      idle-timeout: 600000
      leak-detection-threshold: 20000
      pool-name: shop-pool
  jpa:
    open-in-view: false              # from the transactions guide
    hibernate:
      ddl-auto: validate             # from the Flyway guide

Then watch hikaricp.connections.pending and hikaricp.connections.usage under real load, and change one thing at a time.

The Mental Model, Restated

  1. Pooling exists because connecting is expensive. Hikari is the default and needs tuning, not replacing.
  2. Lower connection-timeout to a few seconds. 30s turns exhaustion into a full outage.
  3. Keep max-lifetime below any infrastructure idle timeout — this explains most "random" connection resets.
  4. Size at roughly cores × 2. Bigger pools are usually slower.
  5. Count every instance against max_connections, including deployment overlap and admin access.
  6. Exhaustion is a symptom — I/O in transactions, leaks, or slow queries. Use leak-detection-threshold.
  7. Alert on pending, the leading indicator.

Phase 3 in Four Sentences

JPA maps entities and gives you repositories, and its defaults — eager @ManyToOne, ordinal enums, N+1 on every loop — need overriding deliberately. Transactions are proxies, so self-calls don't apply and checked exceptions commit; boundaries belong in the service layer with no network I/O inside them. For reports, bulk work and database-specific SQL, JdbcClient is the simpler tool, and both styles share one transaction. Flyway owns the schema with ddl-auto: validate, caching trades correctness for latency and always needs a TTL, and every query ultimately contends for a connection pool that should be smaller than instinct suggests.

What's Next

Phase 4 moves beyond the request/data path to the infrastructure around it: Spring Security's filter chain, the testing pyramid and Spring's test slices, Actuator and observability, AOP as the general case of the proxying you've now seen three times, application events and messaging, and scheduled and asynchronous execution.