Skip to main content
In a retrieval-augmented generation pipeline, a general-purpose LLM can grade each retrieved passage before generation. With GLiDE, each passage receives a probability that it is relevant to the query. Use the probability directly for ranking and filtering.

Score and rerank passages

Each score is the Noul probability that the true criterion holds. Rank descending to put the strongest matches first, and apply a threshold when downstream generation should only receive sufficiently relevant context.
All Noul questions in this example share one state, so GLiDE scores the shortlist in one request. The complete state is included in each internally rendered question; keep it within the 40,000-token per-question limit. Split larger shortlists into bounded concurrent requests and retry 425, 429, and 503 with backoff.

Score one passage per request

Use this simpler shape when passages arrive independently:

What changes from LLM reranking

  • Binary relevance is a probability from 0 to 1, not a generated integer.
  • You can rank and threshold the same value without parsing model text.
  • The criteria state exactly what “relevant” means for your application.
Evaluate ranking quality and choose thresholds on a labeled retrieval set before deploying the reranker. See GLiDE Inference for authentication, limits, errors, and the complete response contract.