HikariCP: Why a Bigger Pool Is Usually Slower
What each Hikari timeout controls, how to size a pool with arithmetic rather than guesswork, and how to see exhaustion coming before it becomes an outage.
The Resource Everything Contends For
Opening a database connection is expensive — a TCP connection, authentication, TLS, session setup. Tens of milliseconds, for work you'd otherwise repeat on every query.
So connections are pooled: a fixed set is opened at startup and handed out on request. HikariCP is Spring Boot's default and needs no configuration to work. It needs configuration to work well, and the defaults are the reason most pool problems look like mysterious latency rather than obvious errors.
What Each Setting Does
spring:
datasource:
url: jdbc:postgresql://db.internal:5432/shop
username: ${DATABASE_USER}
password: ${DATABASE_PASSWORD}
hikari:
maximum-pool-size: 10 # default 10
minimum-idle: 10 # default = maximum-pool-size
connection-timeout: 30000 # 30s — how long a thread waits for a connection
idle-timeout: 600000 # 10m — when idle connections are retired
max-lifetime: 1800000 # 30m — hard cap on any connection's age
keepalive-time: 0 # disabled by default
pool-name: shop-pool| Setting | Controls | Guidance |
|---|---|---|
maximum-pool-size | Total connections | See the arithmetic below — smaller than you'd think |
minimum-idle | Connections kept warm | Leave equal to max for steady load |
connection-timeout | Wait before giving up | Lower it. 30s is far too long |
idle-timeout | Retire idle connections | Only relevant if minimum-idle < max |
max-lifetime | Maximum connection age | Must be below any network/DB idle timeout |
keepalive-time | Probe idle connections | Useful behind aggressive firewalls |
Two of these deserve immediate attention.
connection-timeout: 30000 is too high. It means a request thread waits half a minute for a connection before failing. Under exhaustion, that converts a fast failure into 30 seconds of held request threads — the thread pool fills, and now every endpoint is affected. A timeout of 2–5 seconds fails fast and sheds load, which is what you want:
connection-timeout: 3000max-lifetime must be shorter than the shortest idle timeout between you and the database. Cloud load balancers, firewalls and Postgres's own idle_session_timeout all close connections unilaterally. If Hikari believes a connection is alive and the network has dropped it, the next query fails with a broken-pipe error. Setting max-lifetime a minute or two below that external limit means Hikari retires connections before anything else can.
AWS RDS Proxy, many Kubernetes ingress setups and most cloud NAT gateways close idle TCP connections after a few minutes. The symptom is intermittent connection-reset errors with no pattern, usually after quiet periods. Check what idle timeout sits between your app and your database, then set max-lifetime below it. This one setting explains a large share of "random" database errors.
Sizing: Smaller Than You Think
The instinct is that more connections mean more throughput. Past a point, the opposite is true.
A database executes queries on a finite number of cores. Connections beyond that don't execute in parallel — they queue, and the queueing has cost: context switching, lock contention, and on Postgres, a process per connection with its own memory. A pool of 100 against an 8-core database produces worse throughput than a pool of 20, plus much worse latency variance.
HikariCP's own guidance, from PostgreSQL's:
connections = ((core_count × 2) + effective_spindle_count)
For an 8-core database on SSD, that's roughly 17. For most services, a pool of 10–20 is correct, and Spring Boot's default of 10 is a reasonable starting point rather than a placeholder to raise.
Count every consumer
The database has a global connection ceiling (max_connections, often 100 on a small instance). Your pool is not the only claimant:
8 app instances × 20 connections = 160
+ 2 background workers × 10 = 20
+ migrations, admin tools, monitoring ≈ 10
----
190 ← against max_connections = 100
This fails during a rolling deployment, when old and new instances are briefly both running — the worst possible moment. Total pool size across all instances must fit within max_connections, with headroom for deployment overlap and a human needing to connect.
Reserve capacity for administrative access. A database at its connection limit cannot be connected to in order to diagnose why — including by you, during the incident. Postgres's superuser_reserved_connections keeps a few back; make sure your arithmetic doesn't consume them.
Pool Exhaustion
The failure mode to recognise:
HikariPool-1 - Connection is not available, request timed out after 30000ms.
Every connection is checked out and nothing was returned within the timeout. The important point: this is almost never solved by a bigger pool. It's a symptom, and the causes are specific:
1. A connection held across network I/O. The transaction anti-pattern from the transactions guide — an HTTP call inside @Transactional pins a connection for the call's duration. A few slow calls starve the pool.
2. A leak. A connection obtained and never closed. Hikari can find these for you:
spring:
datasource:
hikari:
leak-detection-threshold: 20000 # warn if held > 20sIt logs a stack trace of the acquiring code — the fastest route to the offending line. Keep it on in non-production; the cost is small enough for production too.
3. Genuinely slow queries. A query taking 5 seconds holds its connection for 5 seconds. The fix is the query (or an index), not the pool.
4. More concurrency than the pool supports. The real case for a bigger pool — but check the first three before assuming it, and remember the database's ceiling.
Virtual threads make this sharper. spring.threads.virtual.enabled: true removes the thread-pool ceiling that was implicitly limiting your database concurrency, so a service that previously queued at 200 threads now sends thousands of requests at a pool of 10. The bottleneck moves to the pool, and connection-timeout becomes the setting that decides whether you shed load or collapse. Enabling virtual threads is a reason to revisit pool configuration, not to ignore it.
Watching the Pool
With Actuator and Micrometer, Hikari publishes metrics that tell you everything:
| Metric | Meaning |
|---|---|
hikaricp.connections.active | Currently in use |
hikaricp.connections.idle | Available |
hikaricp.connections.pending | Threads waiting |
hikaricp.connections.usage | How long connections are held |
hikaricp.connections.acquire | How long acquisition takes |
hikaricp.connections.timeout | Acquisition failures |
pending is the one to alert on. Sustained above zero means threads are waiting for connections — the leading indicator of exhaustion, visible well before the timeouts start. Alert on it rather than on the resulting errors.
usage is the diagnostic companion: if connections are held for seconds, something is doing I/O or slow work inside a transaction, and that's the cause to fix.
Actuator's health endpoint includes a DataSourceHealthIndicator that runs a validation query:
management:
endpoint:
health:
show-details: when-authorized
health:
db:
enabled: trueThink about whether your readiness probe should include the database. If it does, a brief database blip marks every instance unready and takes the whole service out of the load balancer — turning a partial degradation into a total outage. For an app that cannot function without its database, that may be correct; for one that serves cached or static responses, it isn't. Phase 6 covers the liveness/readiness distinction properly.
Check yourself
An app on 8 instances with maximum-pool-size: 25 runs against Postgres with max_connections = 100. It works normally but fails during deployments with 'too many clients already'. Why?
Multiple DataSources
When an app talks to two databases, auto-configuration can't guess which is primary, so you declare both:
@Configuration
public class DataSourceConfig {
@Bean
@Primary
@ConfigurationProperties("app.datasource.orders")
DataSourceProperties ordersProperties() {
return new DataSourceProperties();
}
@Bean
@Primary
@ConfigurationProperties("app.datasource.orders.hikari")
DataSource ordersDataSource(DataSourceProperties ordersProperties) {
return ordersProperties.initializeDataSourceBuilder()
.type(HikariDataSource.class).build();
}
@Bean
@ConfigurationProperties("app.datasource.reporting")
DataSourceProperties reportingProperties() {
return new DataSourceProperties();
}
@Bean
@ConfigurationProperties("app.datasource.reporting.hikari")
DataSource reportingDataSource(DataSourceProperties reportingProperties) {
return reportingProperties.initializeDataSourceBuilder()
.type(HikariDataSource.class).build();
}
}@Primary is what keeps the rest of auto-configuration working — JPA, Flyway and Actuator all need to know which DataSource is the default. This is the @Primary/@Qualifier disambiguation from phase 1 in its most common real use.
Each needs its own pool budget, and both count against the server's limit. Also note that a single @Transactional cannot span two DataSources atomically — that needs XA transactions, which are a significant complication and usually a sign the operation should be redesigned.
A Starting Configuration
Reasonable defaults for a typical service, to be adjusted from measurements:
spring:
datasource:
hikari:
maximum-pool-size: 15
minimum-idle: 15
connection-timeout: 3000
max-lifetime: 1500000 # 25m — below typical infra idle timeouts
idle-timeout: 600000
leak-detection-threshold: 20000
pool-name: shop-pool
jpa:
open-in-view: false # from the transactions guide
hibernate:
ddl-auto: validate # from the Flyway guideThen watch hikaricp.connections.pending and hikaricp.connections.usage under real load, and change one thing at a time.
The Mental Model, Restated
- Pooling exists because connecting is expensive. Hikari is the default and needs tuning, not replacing.
- Lower
connection-timeoutto a few seconds. 30s turns exhaustion into a full outage. - Keep
max-lifetimebelow any infrastructure idle timeout — this explains most "random" connection resets. - Size at roughly
cores × 2. Bigger pools are usually slower. - Count every instance against
max_connections, including deployment overlap and admin access. - Exhaustion is a symptom — I/O in transactions, leaks, or slow queries. Use
leak-detection-threshold. - Alert on
pending, the leading indicator.
Phase 3 in Four Sentences
JPA maps entities and gives you repositories, and its defaults — eager @ManyToOne, ordinal enums, N+1 on every loop — need overriding deliberately. Transactions are proxies, so self-calls don't apply and checked exceptions commit; boundaries belong in the service layer with no network I/O inside them. For reports, bulk work and database-specific SQL, JdbcClient is the simpler tool, and both styles share one transaction. Flyway owns the schema with ddl-auto: validate, caching trades correctness for latency and always needs a TTL, and every query ultimately contends for a connection pool that should be smaller than instinct suggests.
What's Next
Phase 4 moves beyond the request/data path to the infrastructure around it: Spring Security's filter chain, the testing pyramid and Spring's test slices, Actuator and observability, AOP as the general case of the proxying you've now seen three times, application events and messaging, and scheduled and asynchronous execution.