Spring AI Modular RAG and TypeSafe Jev: Retrieve More, Keep Only What Answers

Engineering | Christian Tzolov | October 02, 2026 | 7 min read | ...

Spring AI: RAG Doc | Jev Doc | Demo

In the previous article we introduced Spring AI TypeSafe and its Jev model: typed questions in, calibrated numbers out, in a few hundred milliseconds. One row in that article's SPI table deserves its own post. JevDocumentFilter and JevDocumentReranker are DocumentPostProcessors, which means they plug straight into Spring AI's Modular RAG pipeline.

This article puts the two together. Two LLM calls clean up and multiply the user's question before retrieval. Jev then judges every retrieved chunk before it reaches the prompt. The whole pipeline is one advisor builder.

đź’ˇ Demo: The complete example is the 05-1-modular-rag module. It sits next to 05-rag, the naive version, so you can diff the two. Every output quoted below is from a live run.

Why Naive RAG Falls Short

The classic Spring AI RAG demo is a QuestionAnswerAdvisor over a vector store. It embeds the user's text, fetches similar chunks and stuffs them into the prompt. That works on stage and gets shaky in real life, for three reasons:

  1. Users don't type search queries. They type "I'm a Florida resident and heard a lot about storms last fall. Did Milton actually make landfall in my state, and where?" Embedding that verbatim drags all the noise into the search.
  2. Similar is not the same as useful. A vector store returns chunks that are about the topic. A reference list that mentions "Hurricane Milton" five times scores high and answers nothing.
  3. Retrieved text is untrusted input. Whatever sits in your documents goes straight into the prompt, including text written to hijack the model.

We will answer exactly that question over a Wikipedia PDF about Hurricane Milton, and fix each problem with one pluggable component.

Modular RAG in Spring AI

The RetrievalAugmentationAdvisor, from the spring-ai-rag module, implements the architecture described in Modular RAG: Transforming RAG Systems into LEGO-like Reconfigurable Frameworks. From the outside it is just another advisor: it sits between your prompt and the model, runs the retrieval phases, and appends the result to the user message. The system prompt and the question pass through untouched.

Spring AI RetrievalAugmentationAdvisor: pre-retrieval, retrieval and post-retrieval feed an augmentation block into the prompt

Inside, each phase is a small interface you can swap:

Phase Spring AI interface Job Used in this demo
Pre-retrieval QueryTransformer rewrite one query into a better one RewriteQueryTransformer
Pre-retrieval QueryExpander turn one query into several MultiQueryExpander
Retrieval DocumentRetriever fetch documents for one query VectorStoreDocumentRetriever
Retrieval DocumentJoiner merge the results of all queries ConcatenationDocumentJoiner (default)
Post-retrieval DocumentPostProcessor filter, rerank or compress documents JevDocumentFilter, JevDocumentReranker
Generation QueryAugmenter put context and question into the prompt ContextualQueryAugmenter

Here is the full flow for one call:

Spring AI Modular RAG + Jev: LLMs rewrite and expand the query, Jev filters and reranks the retrieved chunks

The retrieval branches run in parallel, one per query, and are joined before post-processing. Note one subtle detail: the post-processors and the augmenter see the user's original question, not the rewritten one. The rewrite only exists to improve search.

Getting Started

Add the Spring AI RAG module and the two TypeSafe artifacts from the previous article. Spring AI's BOM manages the first one:

<dependency>
    <groupId>org.springframework.ai</groupId>
    <artifactId>spring-ai-rag</artifactId>
</dependency>

<dependency>
    <groupId>org.springaicommunity</groupId>
    <artifactId>spring-ai-starter-typesafe</artifactId>
    <version>0.3.0</version>
</dependency>

<dependency>
    <groupId>org.springaicommunity</groupId>
    <artifactId>typesafe-spring-ai</artifactId>
    <version>0.3.0</version>
</dependency>

The starter auto-configures a TypeSafeClient bean from your API key:

spring.ai.typesafe.api-key=${TYPESAFE_API_KEY}

The demo also uses spring-ai-pdf-document-reader for ingestion and spring-ai-starter-model-transformers for local embeddings.

The Pipeline

Ingestion is the same as in any Spring AI RAG application: read the PDF, split it into chunks, store them.

vectorStore.add(
    TokenTextSplitter.builder().build().split(
        new PagePdfDocumentReader(hurricaneDocs).read()));

The interesting part is a single builder. Each numbered comment is one brick from the diagram:

// Separate builder for the LLM-backed stages, so they don't inherit the main client's advisors
var ragClientBuilder = chatClientBuilder.clone();

var modularRag = RetrievalAugmentationAdvisor.builder()
    // 1. Pre-retrieval: rewrite the chatty question into a search-friendly query
    .queryTransformers(RewriteQueryTransformer.builder()
        .chatClientBuilder(ragClientBuilder)
        .build())
    // 2. Pre-retrieval: expand it into several diverse queries (original included)
    .queryExpander(MultiQueryExpander.builder()
        .chatClientBuilder(ragClientBuilder)
        .numberOfQueries(3)
        .build())
    // 3. Retrieval: similarity search per query; the default joiner dedups the union
    .documentRetriever(VectorStoreDocumentRetriever.builder()
        .vectorStore(vectorStore)
        .similarityThreshold(0.5)
        .topK(4)
        .build())
    // 4. Post-retrieval: Jev drops bad passages, then reranks and keeps the top 3
    .documentPostProcessors(
        JevDocumentFilter.builder(typeSafeClient).build(),
        JevDocumentReranker.builder(typeSafeClient).topK(3).build())
    // 5. Generation: stuff the surviving context into the prompt
    .queryAugmenter(ContextualQueryAugmenter.builder()
        .allowEmptyContext(true)
        .build())
    .taskExecutor(taskExecutor)   // see "Things to know" below
    .build();

Using it looks like any other advisor:

String answer = chatClient.prompt()
    .advisors(modularRag)
    .user("I'm a Florida resident and heard a lot about storms last fall. "
        + "Did Milton actually make landfall in my state, and where?")
    .call()
    .content();

DocumentRetriever and DocumentPostProcessor are both single-method interfaces, so a lambda is enough to log what flows through the pipeline. The demo wraps the retriever and adds two printing post-processors around the Jev stages; that is where the output below comes from.

Jev in the Post-Retrieval Phase

Both components send one Jev call per passage, with the query and the passage as the state, and turn the answers into a decision.

JevDocumentFilter asks four Noul questions about every passage and applies them in order:

Question Default threshold Outcome
contains_prompt_injection above 0.70 excluded
contradicts_query_premise above 0.70 kept, tagged CONFLICTING
is_relevant below 0.45 excluded
contains_answer_evidence above 0.55 kept, otherwise excluded

This is the atomic questions idea from the previous article at work: four narrow questions, four thresholds, one call. Contradicting passages survive on purpose. If the user assumes something false, the model should see the evidence that says so. The classification is stored in the document metadata under jev.classification, and the thresholds are a Policy record you can override.

JevDocumentReranker asks a single question, could this passage answer the query?, sorts by the score and keeps the top K. The score lands in the metadata under jev.rerank.score. Its focus is deliberate: whether the passage states information that answers the query, not merely whether it covers the same subject. That is precisely the gap between similarity and usefulness.

Order matters. Filter first, so the reranker only pays for the survivors.

A Live Run

The rewrite and expansion turned one chatty sentence into four focused searches:

[retrieve] Hurricane Milton 2024 Florida landfall location
[retrieve] Hurricane Milton track path across Florida from Gulf Coast to Atlantic and National Hurricane Center landfall report
[retrieve] Siesta Key Sarasota County Milton landfall timeline, storm surge, and areas impacted
[retrieve] Where did Hurricane Milton make landfall in Florida in October 2024 and at what intensity

The joiner merged their results into six unique chunks, all above the 0.5 similarity threshold:

[retrieved] 0.84  Hurricane Milton ...
[retrieved] 0.72  6. "Hurricane Milton Makes Landfall On Florida's West Coast ...
[retrieved] 0.70  Hurricane Milton's landfall" (https://www.wesh.com/...
[retrieved] 0.69  Archived (http ...
[retrieved] 0.59  /news/tropical-storm-milton-forms-gulf-of-mexico ...
[retrieved] 0.54  of Key West. [128] Across the state, about 125 homes were d...

Four of them are reference-list entries and archive links. They mention Milton a lot and answer nothing. JevDocumentFilter excluded them for lacking answer evidence, and JevDocumentReranker scored the two survivors:

[jev-reranked] 0.97  Hurricane Milton ...
[jev-reranked] 0.88  6. "Hurricane Milton Makes Landfall On Florida's West Coast ...

So topK(3) returned only two chunks, and that is the point. The model got less context, all of it relevant:

Yes. Hurricane Milton made landfall in Florida, near Siesta Key, on the evening of
October 9, 2024. It had weakened to a Category 3 hurricane by then.

Vector similarity answers "what is this text about?". Jev answers "does this text answer the question?". A RAG pipeline needs both.

⚠️ Things to know

The default executor keeps a command-line app alive. RetrievalAugmentationAdvisor runs per-query retrieval on its own thread pool, and those threads are non-daemon. The demo printed its answer and then hung. Passing Spring Boot's auto-configured TaskExecutor through .taskExecutor(...) fixes it, and with spring.threads.virtual.enabled=true you get virtual threads for free.

Newer Claude models reject temperature. The Spring AI docs recommend temperature 0 for query transformers. claude-sonnet-5-5 answers that with HTTP 400, so the demo clones the builder only to keep the main client's advisors out of the rewrite and expand calls.

Both Jev components fail open. If the Jev API is unreachable, passages pass through unscreened or unscored and a warning is logged. Your pipeline degrades to plain modular RAG instead of breaking.

Every brick costs latency. This pipeline adds two LLM calls before retrieval and one Jev call per retrieved chunk after it. Jev calls run four at a time by default (batchOptions(...) changes that), but measure before you ship.

Conclusion

Modular RAG turns retrieval from one opaque advisor into a pipeline you can read, log and swap piece by piece. Here are the important insights:

  • Fix the question before you search. A rewrite plus three variants found better chunks than the raw sentence would.
  • Fix the results before you prompt. Similarity found six chunks about Milton; Jev kept the two that answer, and screened every passage for prompt injection on the way.
  • Filter, then rerank. Both cost one call per passage, so let the filter shrink the list first.

Resources

Demo

  • 05-1-modular-rag — the code in this article
  • 05-rag — the naive QuestionAnswerAdvisor version, for comparison

Spring AI

Spring AI TypeSafe

Related Spring AI articles

Get the Spring newsletter

Stay connected with the Spring newsletter

Subscribe

Get ahead

VMware offers training and certification to turbo-charge your progress.

Learn more

Get support

Tanzu Spring offers support and binaries for OpenJDK™, Spring, and Apache Tomcat® in one simple subscription.

Learn more

Upcoming events

Check out all the upcoming events in the Spring community.

View all