Spring Boot and Postgres: where the limits actually are
One Spring Boot app, one Postgres: where a request really queues, what the connection pool gates, and how much search the database can do on its own.
A Spring Boot app spends its life in front of a Postgres instance. The two articles here work that pairing from opposite ends: what runs out first under load, and where the line is that you are not allowed to cross. One measures where a request actually queues once virtual threads are on. The other pushes search work into Postgres until the JPA Criteria API runs out of ways to express it.
They are independent, so read either one first. Both ship with a companion repo that runs with one command, and every number comes out of that repo rather than someone else's benchmark.
In this guide
Virtual Threads in Spring Boot: Where Your Real Limit LivesSpring Boot virtual threads move Tomcat's concurrency gate from 200 to 8192 without telling you. Where your real limit lives, measured, plus the fix.23 min read
Search that finds typos, synonyms and exact terms: one Postgres table, one Spring Boot appBuilt in plain PostgreSQL 18 first, then driven from Spring Boot 4.1 through the JPA Criteria API, up to the exact line where Criteria runs out and one native query takes over. Runnable repo, 36 tests against real Postgres.23 min read
Key ideas
- Cheap threads relocate your limit, they don't remove it.
spring.threads.virtual.enabled=truemakesserver.tomcat.threads.maxinert and moves the gate toserver.tomcat.max-connections, default 8192. Throughput barely changes; the queue leaves the connector and reappears inside the app. - The scarce resource is usually the pool. About 7,000 concurrent requests piled onto a 10-connection HikariCP pool and roughly 15% failed on the 30-second connection timeout, with zero pinning events.
@ConcurrencyLimit(limit = 10, policy = REJECT)turns those slow failures into fast 503s. - Postgres already does three kinds of search. Exact terms through a stored
tsvector, typos throughpg_trgmtrigrams, and the vocabulary gap through a native synonym dictionary, all over one table with nothing to keep in sync. - The Criteria API reaches further than it looks, then stops: an unregistered
cb.functionrenders straight through, sots_rankneeds no setup, and only@@and%need aFunctionContributor. Reciprocal Rank Fusion needs CTEs,ROW_NUMBER()and aFULL OUTER JOIN, so that one stays native SQL. - Every layer stops somewhere specific.
ts_rankscores a row from that row alone, with no corpus-wide statistics and therefore no inverse-document-frequency term, and Reciprocal Rank Fusion is rank arithmetic that knows nothing about relevance beyond position.
Frequently asked questions
Do virtual threads remove the need for a concurrency limit in Spring Boot?
No. Setting spring.threads.virtual.enabled=true makes server.tomcat.threads.max inert and moves the gate to server.tomcat.max-connections, which defaults to 8192. Throughput stays roughly the same, but requests that used to queue at the connector now queue inside the app and fail late on whatever is actually scarce, usually the connection pool. The fix is a limit you chose, such as @ConcurrencyLimit(limit = N, policy = REJECT).
Can Postgres do full-text search without a separate search engine?
For exact terms, typos and synonyms, yes. A stored tsvector handles lexical matching, pg_trgm trigrams reach misspellings, and a native synonym dictionary bridges vocabulary, all over one table with nothing to keep in sync. What it is not: ts_rank keeps no corpus-wide statistics, so there is no inverse-document-frequency term, and there is no reranker in the database. Semantic search is a separate leg, and that one needs pgvector.
What runs out first in a Spring Boot app under load, threads or database connections?
Usually the database connections. Threads are the cheap resource once virtual threads are on; the pool is not. In a measured run, roughly 7,000 concurrent requests piled onto a 10-connection HikariCP pool and about 15% of them failed on the 30-second connection timeout, with zero pinning events recorded. Whichever layer is scarcest is the one that should carry the deliberate limit.
More guides
- Java Foundations: the primitives, runnableReference articles for the Java primitives, each explained and backed by a runnable brick: data modeling, then the memory model.2 articles 20 min read
- Building ralphctl: an agent harness, release by releaseThe build log of one agent harness: how ralphctl went from a v0.0.4 sprint CLI to a cross-provider loop I trust to run coding agents overnight.5 articles 70 min read
- Spring Boot and OAuth2: a field guide for every surfaceSecure a Spring Boot app with OAuth2 across every surface: HTTP APIs, the gateway, and the message broker, all on one Keycloak and the same scopes.5 articles 102 min read
- AI agent harnesses: a field guideWhat an AI agent harness is, how the generator-evaluator loop and check gates decide when a change is done, and why the harness is the part you trust.3 articles 35 min read