03-data-access

Spring Data JPA: Entities, Repositories and the N+1 Problem

Map entities and relationships, get repositories for free, and learn to see the query count — because the N+1 problem is invisible until you look.

October 9, 2026
spring-bootjpahibernatespring-datan+1entityorm

The Layer Where Production Bugs Live

An ORM makes database access look like working with objects. That is its value and its hazard: order.getCustomer().getName() is a field access in Java and may be a network round trip to Postgres. The syntax hides the cost.

Most JPA problems in production are not mapping mistakes. They are cost-visibility problems — code that reads fine and issues 400 queries. So this guide covers the mapping you need, then spends its second half on seeing what Hibernate actually does.

Entities

java
@Entity
@Table(name = "orders")
public class Order {
 
    @Id
    @GeneratedValue(strategy = GenerationType.IDENTITY)
    private Long id;
 
    @Column(nullable = false, length = 100)
    private String customerName;
 
    @Column(nullable = false, precision = 19, scale = 2)
    private BigDecimal total;
 
    @Enumerated(EnumType.STRING)          // never ORDINAL
    private OrderStatus status;
 
    @Column(nullable = false)
    private Instant placedAt;
 
    protected Order() { }                 // JPA needs a no-arg constructor
 
    public Order(String customerName, BigDecimal total) {
        this.customerName = customerName;
        this.total = total;
        this.status = OrderStatus.NEW;
        this.placedAt = Instant.now();
    }
    // getters; setters only where mutation is legitimate
}

Four decisions in there are worth defending.

@Table(name = "orders") because order is a reserved word in SQL. Naming the table explicitly avoids a class of quoting bug.

@Enumerated(EnumType.STRING), always. The default is ORDINAL, which stores the enum's position — so inserting a new constant in the middle silently reinterprets every existing row. This is a genuine data-corruption default. Store the name.

BigDecimal with explicit precision and scale for money. double cannot represent 0.10 exactly, and errors accumulate across a sum.

A protected no-arg constructor, required by JPA for proxying and reflection, kept non-public so application code uses the real one.

⚠️

Records cannot be JPA entities. They're final with final fields, and JPA requires a mutable class with a no-arg constructor it can proxy. Records are excellent for DTOs — which is the pairing from the Jackson guide: mutable entities inside the persistence layer, immutable records on the API boundary.

ID generation

StrategyMechanismNotes
IDENTITYAuto-increment columnSimple; disables JDBC batch inserts
SEQUENCEA database sequencePreferred on Postgres/Oracle; batches fine
AUTOProvider choosesPicks a sequence-table strategy on some databases
UUIDGenerated UUIDGood for distributed writes; wider index

On Postgres, prefer SEQUENCE with an allocation size:

java
@Id
@GeneratedValue(strategy = GenerationType.SEQUENCE, generator = "order_seq")
@SequenceGenerator(name = "order_seq", sequenceName = "order_seq", allocationSize = 50)
private Long id;

IDENTITY forces Hibernate to execute each insert immediately to learn the generated ID, which defeats batching. If you bulk-insert, this single choice can be the difference between one round trip and a thousand.

Relationships

java
@Entity
public class Order {
 
    @OneToMany(mappedBy = "order", cascade = CascadeType.ALL, orphanRemoval = true)
    private List<OrderLine> lines = new ArrayList<>();
 
    @ManyToOne(fetch = FetchType.LAZY)        // explicit, not the default
    @JoinColumn(name = "customer_id", nullable = false)
    private Customer customer;
 
    public void addLine(OrderLine line) {     // keep both sides consistent
        lines.add(line);
        line.setOrder(this);
    }
}

mappedBy marks the non-owning side. The side with the foreign-key column owns the relationship. Omit mappedBy on a bidirectional association and Hibernate assumes two separate unidirectional ones, producing a surprise join table.

Maintain both sides in code. Hibernate writes based on the owning side, so setting only order.lines without line.order can leave a null foreign key. A helper method like addLine is the standard remedy.

Default fetch types are a trap:

AssociationDefault
@ManyToOneEAGER
@OneToOneEAGER
@OneToManyLAZY
@ManyToManyLAZY

The two EAGER defaults mean loading one Order also loads its Customer, and whatever that eagerly loads, transitively. A chain of @ManyToOne relations can pull half your schema for a single findById.

🚨

Set fetch = FetchType.LAZY on every @ManyToOne and @OneToOne. This is the highest-value habit in JPA. EAGER is not a performance optimisation — it is an unconditional join on every query that touches the entity, whether you need the association or not. Make loading explicit per query instead, with the fetch joins below.

Repositories

Declare an interface; Spring Data implements it:

java
public interface OrderRepository extends JpaRepository<Order, Long> {
 
    List<Order> findByStatus(OrderStatus status);
    List<Order> findByCustomerNameContainingIgnoreCase(String fragment);
    Optional<Order> findByIdAndStatus(Long id, OrderStatus status);
    long countByStatus(OrderStatus status);
    boolean existsByCustomerName(String name);
    Page<Order> findByStatus(OrderStatus status, Pageable pageable);
}

The hierarchy: Repository → CrudRepository → PagingAndSortingRepository → JpaRepository. JpaRepository adds JPA-specific operations like flush and saveAllAndFlush. Extending it is fine and conventional; exposing a narrower interface is a reasonable discipline if you want to restrict what callers can do.

Derived queries are parsed from the method name — findBy, countBy, existsBy, deleteBy, plus keywords like Between, LessThan, In, OrderBy, IgnoreCase. Parsing happens at startup, so a typo'd property name fails fast rather than at first call.

They stop paying off past about three conditions. findByStatusAndCustomerNameContainingIgnoreCaseAndPlacedAtBetweenOrderByTotalDesc is not better than JPQL:

java
@Query("""
       select o from Order o
       where o.status = :status
         and o.placedAt between :from and :to
       order by o.total desc
       """)
List<Order> findRecentByStatus(@Param("status") OrderStatus status,
                               @Param("from") Instant from,
                               @Param("to") Instant to);

JPQL queries entities and their fields, not tables and columns. For database-specific SQL, drop down:

java
@Query(value = "select * from orders where total > :threshold", nativeQuery = true)
List<Order> findExpensive(@Param("threshold") BigDecimal threshold);

Native queries cost you provider portability and aren't validated at startup. Worth it for a window function; not worth it to avoid learning JPQL.

Projections: stop fetching what you don't use

To display an order list you rarely need whole entities. Ask for the columns you want:

java
public interface OrderSummary {           // interface projection
    Long getId();
    String getCustomerName();
    BigDecimal getTotal();
}
 
List<OrderSummary> findByStatus(OrderStatus status);
java
@Query("select new com.example.shop.order.OrderSummaryDto(o.id, o.customerName, o.total) from Order o")
List<OrderSummaryDto> summaries();        // DTO projection

Projections sidestep most lazy-loading problems entirely: there is no managed entity, so nothing can be lazily dereferenced later. For read-only endpoints they are usually the right answer, and the fastest one.

The N+1 Problem

The defining performance bug of ORMs, and the one worth being able to recognise instantly.

java
List<Order> orders = orderRepository.findAll();         // 1 query
for (Order order : orders) {
    System.out.println(order.getCustomer().getName());  // 1 query each
}

100 orders means 101 queries. Each is fast, which is what makes this so hard to spot — no single slow query appears in your database's slow-query log. The endpoint is just inexplicably slow, and it degrades linearly with data volume.

Fixes, in order of preference

1. A fetch join — load the association in the same query:

java
@Query("select o from Order o join fetch o.customer where o.status = :status")
List<Order> findByStatusWithCustomer(@Param("status") OrderStatus status);

2. An entity graph — declarative, reusable, works with derived queries and paging:

java
@EntityGraph(attributePaths = { "customer", "lines" })
List<Order> findByStatus(OrderStatus status);

3. Batch fetching — turn N queries into N/size queries:

yaml
spring:
  jpa:
    properties:
      hibernate:
        default_batch_fetch_size: 50

This is a good global default. It won't match a fetch join, but it turns a pathological 101 queries into 3 with no code change — useful insurance for the paths you haven't audited.

4. A projection — don't load entities at all. Often the real answer.

⚠️

join fetch on a collection breaks pagination. Fetch-joining a @OneToMany multiplies rows (one per child), so LIMIT applies to the joined rows, not the parents. Hibernate detects this and pages in memory after loading everything — with a warning, and catastrophic behaviour on a large table. For paging plus collections, use an @EntityGraph, batch fetching, or two queries: page the IDs, then fetch the collections for that page.

Check yourself

An endpoint returns 20 orders, each with its lines. The log shows 21 queries. The fix applied is @Query('select o from Order o join fetch o.lines') with Pageable. What goes wrong?

Seeing the Queries

You cannot tune what you cannot see. Make queries visible in development:

yaml
spring:
  jpa:
    show-sql: false            # unformatted, no parameters — prefer the logger
    properties:
      hibernate:
        format_sql: true
logging:
  level:
    org.hibernate.SQL: DEBUG
    org.hibernate.orm.jdbc.bind: TRACE        # parameter values

Better still, count them. A query log tells you what ran; a count tells you whether it regressed. Libraries like datasource-proxy or Hypersistence Utils can assert a query count in a test:

java
@Test
void listingOrdersIssuesOneQuery() {
    SQLStatementCountValidator.reset();
    orderService.list();
    SQLStatementCountValidator.assertSelectCount(1);
}

That test is how an N+1 stops coming back. Without it, someone adds an innocuous order.getCustomer() to a loop six months from now and the regression ships silently.

✅

show-sql: true writes to stdout unformatted and without parameter values. Use the logger configuration instead — it respects your log framework, formats the SQL, and org.hibernate.orm.jdbc.bind at TRACE shows the actual bound parameters, which is usually the part you need. Never enable these in production: the volume is enormous and bound parameters may contain personal data.

Two Traps Worth Naming

findAll() with no limit. Fine on a 50-row reference table, an outage on a 50-million-row one. Prefer Pageable on anything that grows.

equals/hashCode on entities. The default identity-based implementations break when an entity is detached and reattached; naive ID-based ones break for unsaved entities whose ID is still null. If entities go into a HashSet, implement equals/hashCode on a business key — something immutable and unique that isn't the generated ID. If they don't, leave the defaults alone.

Check yourself

A @ManyToOne association is left at its default fetch type. A repository method loads 500 entities for a report that never touches that association. What is the cost?

The Mental Model, Restated

  1. An ORM hides cost, not complexity. A getter can be a network round trip.
  2. EnumType.STRING, BigDecimal for money, explicit table names. The defaults are wrong in ways that corrupt data.
  3. fetch = LAZY on every @ManyToOne and @OneToOne.
  4. Derived queries for simple cases, JPQL past three conditions, projections for read-only endpoints.
  5. N+1 is the default failure mode. Fix with fetch joins or entity graphs; set default_batch_fetch_size as insurance.
  6. join fetch plus pagination is a trap.
  7. Log queries in development and assert query counts in tests.

What's Next

Every write so far has been implicitly transactional — save works because Spring Data wraps it. Once you have multiple writes that must succeed or fail together, that implicit behaviour isn't enough. The next guide covers @Transactional properly: where boundaries belong, what propagation and isolation actually change, the rollback rule that surprises people, and the self-invocation problem that makes the annotation silently do nothing.