← The RAG seriesEnglish日本語中文
RAG Practice 03 · NAIVE RAG · DOCUMENT QA · Readable
What to think about when RAG gives an incomplete answer
26 pages of English Cloudflare documentation, 336 chunks, a basic RAG built on them, and questions asked in Chinese. One of the 9 test questions came back incomplete, and I followed it through Top-k, deduplication, a reranker, splitting the question, rewriting it and swapping the embedding model.
- Documents
- Chunking
- Embedding
- Retrieval
- Rerankexperiments only
- Answer
- Testing
Cloudflare Docs (English): 26 pages, 336 chunksQuestions in Chinese; the English is only for diagnosisBaseline: structure-aware chunks, bge-m3, cosine Top-53 embedding models compared9 questions; each experiment that calls a model ran onceIntermediate level, about 22 minutes to read through; for the conclusions only, read 07; starting point
Read in order: build a working basic version → run real tests → one question answers incompletely → decide which layer fails → run experiments → test again → where it stands → look back
First see how this baseline works, then how Q7 gets an incomplete answer → 01
Starting point, level, how to read (skippable)
Guide · Starting point
Why this experiment, and who it is for
- Starting point
- Most RAG tutorials go technique by technique: chunk first, then embed, then rerank… each step looks useful. This one goes the other way: build a small but complete RAG, ask real questions of real documents, then follow one real failure all the way down and see which failure each mechanism actually answers.
- Who it is for
- Engineers and product people who build, or are about to build, “answer questions from documents”; people who have run a basic RAG and met “everything retrieved is relevant, yet the answer is incomplete”.
- What you learn
- How a real baseline works
- “Relevant” and “enough to answer” are not the same thing
- How to tell “the corpus has no evidence” from “retrieval missed it”
- Which failure each of Top-k, deduplication, a reranker, splitting, rewriting and a different embedding answers, and when it does nothing
- What a system can still do for the user when retrieval cannot cover the question
- Not for
- An install-from-scratch tutorial; a search for the “best RAG configuration” or a model leaderboard; Graph RAG, agents or fine-tuning; a production guide you can ship as it stands.
Guide · What this piece is
What it is, and how long it takes
- Nature
- The investigation log of a real engineering experiment: not a tutorial, and not an official Cloudflare guide. The conclusions rest on this snapshot (26 pages, taken 2026-10-04), this benchmark and a few cases. Questions are asked in Chinese, the documents are in English, and the goal is that every sentence of the answer has evidence behind it.
- Level
- Intermediate. You should roughly know what a chunk, an embedding and vector search are (01 explains each in a sentence); you do not need to write code.
- Reading time
- About 22 minutes read through; for the conclusions only, read 07, about 3 minutes.
- Shallow to deep
- 01–03 Basics: a baseline, one case that works, one case with no evidence
- 04–05 Going further: a question that answers incompletely, and six hypotheses ruled out in turn
- 06 Putting it together: what to do when retrieval cannot cover it
- 07 Conclusions, 08 Reproduction notes
- How to read
- Reading in order is smoothest. If the basics are familiar, start at 04; for the conclusions read 07; to reproduce it yourself read 08.
01 · Baseline · Basics
A RAG kept deliberately simple
What follows investigates a failure, not tuning, so every link is as simple as it can be. After this section you only need to remember this line:
- Official docs
- Structure-aware chunks
- embedding
- cosine retrieval
- Top-5 as evidence
- Qwen generates
- Citation check
- Corpus
- 26 pages of official Cloudflare documentation (Pages, Workers, R2, D1, Wrangler, bindings, variables and secrets, limits and pricing), snapshot
v1-2026-10-04 - Chunking
- 336 chunks, of which 334 have body text that can be embedded (the other 2 are headings only); median length 827 characters
- Retrieval
bge-m3 (1024 dimensions) + flat cosine; the Top-5 go to generation- Generation
qwen3.8-flash: may only use the evidence it is given, must cite chunks and quote the source word for word; the system checks that each citation really exists- Enough?
- Whether the evidence is enough to answer is judged by a hand-written “evidence checklist”; if it is not enough, the system declines and the model is not called (the checklist was written for one case only and is not general)
New to these words?
- chunk
- A small piece cut from a document; the smallest unit of retrieval and citation
- embedding
- Maps questions and chunks to vectors, so that distance finds content with a similar meaning
- cosine
- A score for how close two vectors point; higher is more similar
- Top-k
- Take the k highest-scoring chunks as evidence
- Citation check
- Every citation in the answer must match a chunk that was really retrieved, and the quoted text must appear word for word in that chunk
How the chunks are cut, which models were used and what the environment was are in the “Reproduction notes” at the end; they are not expanded here.
02 · What working looks like · Basics
A successful case: Q5
Question怎么给 Pages Functions 配置 R2 bucket 绑定,并在代码里访问它?Translation: How do I configure an R2 bucket binding for Pages Functions, and access it in code?
First, what it looks like when it works, so that when the failures come you know the whole system is running.
The single chunk at rank 1 covers “how to configure an R2 binding for a Pages project”, “redeploy” and “access it with context.env” together, so the evidence status is sufficient (sufficient). Qwen cited the chunks at ranks 1, 2, 3, 4 in 4 places; the quoted text exists word for word in each, and the check passed.
Even the successful case has a gap: the chunk with the Wrangler configuration syntax for Pages sits at rank 8, outside the Top-5. We recorded the gap and did not patch it.
03 · When there is no evidence · Basics
Q9a: asking about something the corpus does not contain
QuestionWorkers KV 的免费额度是多少?超出后如何计费?Translation: How large is the Workers KV free allowance? How is usage beyond it billed?
Check the corpus first: 34 chunks (13 pages) in this snapshot mention KV, but all of them are about bindings or configuration; none gives the free allowance or the billing. This is not a retrieval failure: no chunk was “missed” by retrieval.
Retrieval still returned 5 chunks, scoring 0.574–0.593, about the limits and free allowances of Workers, Pages and R2. Q5’s rank 5 scored 0.608: the score alone cannot tell “has evidence” from “has none”.
- Normal path: the evidence status is insufficient, and the system declines without calling the model.
- In the experiment we bypassed that gate on purpose and let Qwen face the same 5 distractor chunks. Its reply was: “提供的证据中未包含 Workers KV 的免费额度及超出后的计费信息。Translation: The evidence provided does not include the Workers KV free allowance or how usage beyond it is billed.” No number, no other product’s allowance passed off as KV’s, and no citation.
“The evidence provided does not include it” must not be written as “the Cloudflare docs do not have it”.Not in the corpus ≠ retrieval missed it ≠ the vendor does not have it. The model was called once, which shows it did not overstep this time, not that it is stable; and a thing missing entirely is also the easiest kind to decline.
04 · The question that is really trouble · Going further
Q7: everything looks relevant, yet the answer is incomplete
Question环境变量、secret 和普通配置应该分别怎么设置?本地开发和线上部署有什么区别?Translation: How should environment variables, secrets and plain configuration each be set? What is the difference between local development and a production deployment?
A complete answer needs these four pieces:
Look at the Top-10 first. Reading the headings, you would think retrieval did well:
All of it is secrets, local development and environment variables; nothing is off topic. But it covers only 1/4 of the four pieces needed: just C4. C1, C2 and C3 are in the corpus, within the first 60, but ranked 52, 36 and 36.
semantic similarity ≠ evidence sufficiency.“Related to the question” and “enough to answer it” are two things. Retrieval does the former.
One open labelling point: chunks like “Compare secrets and environment variables” are marked irrelevant in the benchmark, yet to a human reader they help with C1/C2. Whether they count as evidence is a Gold question the author has yet to review; nothing was changed here, and no score was adjusted for it.
05 · Decide which layer, then experiment · Going further
Six hypotheses, and every one really ran
First decide which layer fails. Is the evidence missing from the corpus? No: C1–C3 are in the first 60. Does it fall just outside the Top-5? No either: they are at ranks 36–52. So it is “there, but ranked deep”. Each hypothesis below stands for a common mechanism, and every step is: hypothesis → experiment → result → judgement. Experiments that did not help are written here too, because a technique does not work just by being added: each mechanism answers a different failure.
- Hypothesis
- The chunks we need are at ranks 6–10 and the Top-5 cuts them off.
- Experiment
- Change only k: 3, 5, 6, 10, as prefixes of the same ranking.
- Result
- Q8 (are the Pages and R2 free allowances shared?) gets its missing piece only at k=6: coverage 1/2 → 2/2. Q7 stays at 1/4 from k=3 to 10. Raising k to 10 grows the irrelevant chunks across the four questions from 8 to 24, and the prompt to about 1.6 times as long.
- Judgement
- Q8 is a real cutoff problem; Q7 is not: what it lacks is at ranks 36–52. Keep k at 5.
- Hypothesis
- The same passage appears on several pages and fills the Top-5; remove the duplicates and the missing evidence will rise.
- Experiment
- On the candidate pool of the first 60, three approaches, one run each (parameters fixed beforehand, not tuned): keep one of every identical chunk; also fold near-duplicates; MMR (λ=0.2).
- Result
- The duplication is real: in the first 60 there are 6 exact duplicates and 13 near-duplicates (including the former), and 4 of the Top-5 are redundant. But the Top-10 coverage of all three approaches is still 1/4. Once duplicates are removed, the freed places are filled by other secrets or local-development chunks. MMR even pushes C1 from position 52 to position 59 (ordering the whole candidate pool its way).
- Judgement
- Duplicates waste places, but they are not why the missing evidence ranks deep. Deduplication is useful; it solves a different problem.
New to MMR? Beyond relevance, it penalises chunks too similar to those already chosen, so the results spread out.
- Hypothesis
- The embedding is only a coarse ranking; if a stronger model judges “question–chunk” relevance one by one, the right ones will rise.
- Experiment
- Rerank the first 60 with
BAAI/bge-reranker-v2-m3, looking only at the reranker’s score, not mixed with the original cosine. - Result
- C2 and C3 move from rank 36 to rank 18, C1 from 52 to 49, and the Top-10 coverage is still 1/4. The piece that answers “how do I set a secret in production”, the
wrangler secret put chunk, actually drops from rank 43 to rank 55; an irrelevant chunk about remote bindings rises from rank 29 to rank 8. Half of the Top-10 is still duplicated content. - Judgement
- A stronger relevance scorer does not solve it by itself: the question has both “difference” and “local development” in it, the scorer is drawn to those two parts and rates the chunk that directly says how very low.
New to rerankers? Retrieve a batch of candidates first, then use a heavier model to judge the relevance of question and chunk one by one.
- Hypothesis
- Two questions are mixed together and the embedding is pulled off by one of them.
- Experiment
- Split mechanically at the question mark into two sentences (wording unchanged), retrieve for each, then interleave the results. The first sentence is about variables and secrets, the second about local versus production.
- Result
- Asking only the first sentence moves C1 from 52 to 39 and C2, C3 from 36 to 19: earlier, but still outside the Top-10. The second sentence finds none of the needed pieces, and even C4 only appears at rank 37. The merged Top-10 coverage is still 1/4.
- Another case
- A hand-written split was tried on Q8 as well: two sub-questions, “the Pages free allowance” and “the R2 free allowance”; the merged 10 candidates cover only 1 kinds, while the unsplit original already covers 2 kinds in its Top-10 (the missing piece is at rank 6). Splitting did not help and cost one more retrieval.
- Judgement
- A compound question does dilute the signal (it moved up by a dozen or so ranks), but splitting it does not get the evidence into the Top-10.
- Hypothesis
- If we ask in the document’s words, retrieval will be more accurate.
- Experiment
- Two rewrites written beforehand, one run each: one stays in Chinese and swaps in Cloudflare terms; the other is English written straight from the document’s wording, which is open to a suspicion of leakage and is used only to test the upper bound.
- Result
- Both versions score a higher cosine than the original question, yet C1/C2/C3 are still deep: the Chinese rewrite 85 / 107 / 72, the English rewrite 53 / 48 / 44 (the original question 52 / 36 / 36); the Top-10 coverage of both is 1/4.
- Judgement
- Sounding more like the document does not mean carrying more evidence. A higher cosine only says it is closer to that stretch of text.
See the originals of the two rewrites
Chinese-terms version: Cloudflare Workers 里的 environment variables、secrets 和普通配置(vars)应该分别怎么设置?本地开发(Wrangler)和线上部署(production)有什么区别?
(In English: In Cloudflare Workers, how should environment variables, secrets and plain configuration (vars) each be set? What is the difference between local development (Wrangler) and a production deployment (production)?)
Document-aligned English version: How do I configure plain text environment variables in the Wrangler configuration file, add encrypted secrets with wrangler secret put for production, and use .dev.vars or .env files for local development in Cloudflare Workers and Pages?
- Hypothesis
- The first-stage representation, bge-m3, is itself the cause.
- Experiment
- Change only the model, with document vectors regenerated for each: bge-m3 (the baseline), nomic-embed-text-v1.5, snowflake-arctic-embed-l-v2.0.
- Result
- On Q7, nomic is clearly better locally: C2 and C3 reach ranks 2 and 3, and the Top-10 coverage is 3/4. But the “plain variables” piece falls to rank 134; arctic, like bge-m3, has only 1/4 in the Top-10. Back to the whole benchmark (Chinese questions): on the 4 fully scorable questions both have Recall@5 of 0.56, but MRR is 0.88 against 0.54; on the 4 questions that can only be scored in part, Recall@5 is 0.88 against 0.38.
- Judgement
- Do not replace the global baseline to rescue one failure case. bge-m3 was chosen on the whole benchmark, and what nomic gains on Q7 is paid for with losses elsewhere. nomic leans toward English; why it ranks Chinese questions this way has not been verified.
Side by side
Below is where C1, C2 and C3 first appear after each attempt. Green means in the Top-10.
Six approaches, and not one gets C1–C3 all into the Top-10.Each mechanism answers a different failure; this question’s failure is not at the layer any of them works on.
06 · Recovery Loop · Putting it together
When search cannot find it, help the user carry on
By now Q7 could stop at “not found, so we do not know”. There is another way: the system admits the evidence is not enough and helps the user turn the question into one that is easier to answer.
- Original questionevidence incomplete
- Qwen proposes 3 narrower questions
- The user picks onetreated as a new ordinary query
- Retrieve again
- A grounded answer+ citation check
“Recovery Loop” is a name I gave this approach, not a general term. The model sees only the original question and the current Top-5 evidence, and the rules are general: do not answer the original question, do not add Cloudflare facts, give only 2–3 narrower questions. It was not told the categories, which chunks rank low, what experiments came before, or “the right way to split”. Parameters are as before: qwen3.8-flash, temperature 0, thinking off, called once; once generated the suggestions are frozen, with no retry and no polishing, and only then are they searched.
Its reason: “原问题同时询问了环境变量、secret和普通配置的设置方法,以及本地与线上部署的区别,涉及多个独立主题且当前证据主要聚焦于本地开发中的 secret 处理。Translation: The original question asks at once how to set environment variables, secrets and plain configuration, and how local and production deployments differ: several separate topics, while the current evidence is mainly about handling secrets in local development.” The questions it suggested:
- S1在 Cloudflare Workers 中,如何区分设置普通环境变量(vars)和敏感信息(secrets)?Translation: In Cloudflare Workers, how do I set plain environment variables (vars) and sensitive values (secrets) separately?
- S2线上部署时,如何通过 Wrangler CLI 或 Dashboard 为 Worker 添加 secrets?Translation: For a production deployment, how do I add secrets to a Worker through the Wrangler CLI or the Dashboard?
- S3本地开发和线上部署在加载 .dev.vars/.env 文件与生产环境配置时有何具体区别?Translation: When loading .dev.vars/.env files and production configuration, what exactly differs between local development and a production deployment?
Say the user taps S2. That is one new ordinary query, on the same bge-m3, with nothing changed. The result for each of the three suggestions:
Its top three are Via the dashboard, Via Wrangler, Upload secrets alongside code: the dashboard path, wrangler secret put, and --secrets-file for uploading secrets with the code.
S2: C3 goes from rank 36 to rank 2.The piece that ranked 36 for the original question gets into the top 3 once the question asks only “how do I set a secret in production”.
A final answer for S2 only (with S2’s own Top-5 and the same citation rules):
After deployment there are two ways to add a secret to a Worker [1][2]: in the dashboard, under Workers & Pages → your Worker → Settings → Variables and Secrets, choose the type Secret, then Deploy; or use wrangler secret put <KEY>, which creates and immediately deploys a new version. With gradual deployments, use wrangler versions secret put instead. You can also upload secrets together with the code using --secrets-file [3].
That is a paraphrase of the full answer; it has 3 citations and the check passed: the cited chunks all come from S2’s own Top-5, and the quoted text exists word for word.
See the model’s full answer
在线上部署时,可以通过以下两种方式添加 Worker secrets:
1. **通过 Dashboard**:
- 在 Cloudflare dashboard 中进入 **Workers & Pages** 页面 [1]。
- 选择你的 Worker > **Settings** [1]。
- 在 **Variables and Secrets** 下选择 **Add** [1]。
- 选择类型 **Secret**,输入 **Variable name** 和 **Value**(值将在 Wrangler 和 dashboard 中隐藏)[1]。
- 可选地,选择 **Add variable** 以添加更多 secrets [1]。
- 最后选择 **Deploy** 以实施更改 [1]。
2. **通过 Wrangler CLI**:
- 使用 `wrangler secret put <KEY>` 命令。这会创建一个新的 Worker 版本并立即部署 [2]。
- 如果使用 gradual deployments,则使用 `wrangler versions secret put <KEY>` 命令。这仅创建新版本,随后需使用 `wrangler versions deploy` 进行部署 [2]。
- 注意:Wrangler 3.73.0 之前的版本需要指定 `--x-versions` 标志 [2]。
- 此外,可以使用 `--secrets-file` 标志在部署代码的同时上传 secrets,例如 `npx wrangler deploy --secrets-file .env.production` 或 `npx wrangler versions upload --secrets-file secrets.json` [3]。
In English (translation):
For a production deployment, Worker secrets can be added in two ways:
1. **Through the Dashboard**:
- In the Cloudflare dashboard, go to the **Workers & Pages** page [1].
- Select your Worker > **Settings** [1].
- Under **Variables and Secrets**, select **Add** [1].
- Choose the type **Secret**, and enter the **Variable name** and **Value** (the value is hidden in Wrangler and the dashboard) [1].
- Optionally, select **Add variable** to add more secrets [1].
- Finally, select **Deploy** to apply the change [1].
2. **Through the Wrangler CLI**:
- Use the `wrangler secret put <KEY>` command. This creates a new Worker version and deploys it immediately [2].
- With gradual deployments, use the `wrangler versions secret put <KEY>` command instead. This only creates a new version, which you then deploy with `wrangler versions deploy` [2].
- Note: Wrangler versions before 3.73.0 need the `--x-versions` flag [2].
- You can also upload secrets together with the code using the `--secrets-file` flag, for example `npx wrangler deploy --secrets-file .env.production` or `npx wrangler versions upload --secrets-file secrets.json` [3].
The limits stay on record; recovery is not a cure-all:
- S1 (the difference between plain variables and secrets) is not solved: C1 moves up to rank 31 and C2 to rank 16, still outside the Top-10. The most basic piece of Q7 is still not recovered.
- S3 (local development versus production configuration) mostly just covers C4, which could already be found; no new progress.
- One case, one set of suggestions, one call; the final answer did not use the “evidence checklist” gate, because none was written for this question.
A good RAG need not answer every question in one go, but it should know when the evidence is not enough and help the user continue.The result here is partial: one sub-question was answered completely and with grounds, and the other still was not.
07 · Where it stands, and a look back
When you meet a RAG failure, think this way; what worked, what did not, what was not done
Relevant ≠ enoughQ7’s Top-10 has nothing off topic, yet covers only 1/4 of the four pieces needed.
A high cosine ≠ can answerWith no KV allowance in the corpus, the Top-5 scores are 0.574–0.593; Q5, which has an answer, scores 0.608 at rank 5. A score threshold cannot block it.
Not in the corpus ≠ retrieval missed itCheck the corpus before blaming retrieval; and the answer may only say “the evidence provided does not include it”, never “the official docs do not have it”.
A reranker, deduplication, rewriting and splitting each have conditions where they applyQ8’s problem was fixed by k=6; on Q7 none of these got the missing evidence into the Top-10. Decide which layer the failure is at, then choose the mechanism.
Do not overfit the baseline to one querynomic is better locally on Q7, but bge-m3 is still stronger on the whole benchmark; change the baseline on the whole benchmark, not on one case.
Recovery for compound questions can be a product capability; there is no need to keep stacking retrievalLet the system say the evidence is not enough and offer narrower questions, and S2 takes C3 from rank 36 to rank 2 and gets an answer with citations.
Looking back: what this round did
- Worked
- With evidence it answers with citations; with none in the corpus it declines instead of making something up (Q5, Q9a).
- Located a typical failure: Q7’s Top-10 has nothing off topic, yet covers only 1/4 of the pieces the answer needs.
- Top-k works for Q8: at k=6 the missing piece is filled in (1/2 → 2/2).
- Having the system admit the evidence is not enough and then offer narrower questions: the missing “how do I set a secret in production” piece goes from rank 36 to rank 2, with an answer that cites its sources.
- Did not work
- Top-k, deduplication, a reranker, splitting and rewriting: not one got Q7’s missing evidence into the Top-10.
- Changing the embedding: nomic reaches 3/4 on Q7, but the “plain variables” piece falls to rank 134, and it is not as good as bge-m3 on the whole benchmark either.
- The “plain variables” and “secrets are encrypted” pieces never reach the Top-10 on the bge-m3 baseline, even after recovery.
- Not done
- Not a product: no interface, no online deployment; no vector database, no BM25 / hybrid retrieval, no metadata filtering; no fine-tuning.
- The reranker, rewriting and deduplication are only experiments; none entered the baseline.
- The evaluation is small: 9 questions, Q7 only one of them; each experiment that calls a model ran once.
- Generation quality was not evaluated, and different generation models were not compared.
Still unresolved
- C1/C2 of Q7 (the difference between plain variables and secrets) still cannot be recovered, even with a question aimed at them.
- Whether chunks like “Compare secrets and environment variables” count as evidence is a Gold question for the author to review.
- Every conclusion comes from this snapshot, this benchmark and a few cases; each experiment ran once, which does not show that the model or the method is stable.
After reading, you should be able to answer this: facing a retrieval failure, what to judge first — is the evidence missing from the corpus, outside the top k, or there but ranked deep — and can the system, for now, admit it is not enough and help the person carry on.
08 · Reproduction notes
If you want to reproduce it, roughly what you need
This is not a complete setup guide. It tells an experienced engineer where the material came from, what was used and where the experiments are. The code and corpus are in the author’s repository and are not public at the moment; the page has no download.
- Data
- The official Cloudflare documentation, taking each page’s Markdown export (
index.md): 26 pages, snapshot batch v1-2026-10-04, fetched within a very short window on 2026-10-04, with a sha256 recorded per file. It is a span in which the data was taken and does not correspond to one commit of the official repository; a later update means a new batch, never overwriting this one. Three layers: raw (the responses as saved) → normalized (only the interface elements the documents record removed; for the long pages that mix several topics, only the relevant sections kept) → chunks. - Chunking
- Cut by structure: a heading, a paragraph, a list item, a table, a code block and a
<details> are each a block; a heading is bound to the explanation right below it; a chunk never crosses a heading; each chunk tries to stay within about 2400 characters, with zero overlap between neighbours; tables and code blocks are never split. Only one chunk in the whole corpus is over the limit (5759 characters, a compatibility matrix table), and it is kept whole. A chunk’s id carries a hash of its text. - Embedding
- Input = document title + heading path + chunk text. The baseline is
bge-m3 (1024 dimensions); nomic-embed-text-v1.5 and snowflake-arctic-embed-l-v2.0 were compared. An earlier bake-off also ran bge-small-en-v1.5 (512-token window, which truncates 57 chunks). The current baseline is still bge-m3. - Reranker
BAAI/bge-reranker-v2-m3, used only in the Q7 experiments and not in the baseline; run on CPU through a third-party ONNX export.- Generation
qwen3.8-flash (Alibaba Cloud Model Studio’s OpenAI-compatible interface), temperature 0, enable_thinking=false, JSON output; the key and address are read only from environment variables: DASHSCOPE_API_KEY, DASHSCOPE_BASE_URL.- Environment
- Python 3.11, ONNX Runtime 1.30.0, numpy 2.4.6, 4 CPU threads, no GPU; the model files are in a local cache, not in the repository.
Where the experiment material is (author’s repository)
experiments/rag/cloudflare-docs-practice-03/
snapshots/ corpus snapshots (raw, normalized, manifest)
ingest/ chunking rules and chunks
embedding/ embedding bake-off and benchmark
case-01/ Q5's full run, the Q9a boundary test
q8-decomposition/ hand-written splitting
topk-sensitivity/ Top-k
q7-diversity/ deduplication and MMR
q7-reranker/ reranker
q7-query-split/ splitting by sentence
q7-query-rewrite/ rewriting
q7-embedding-comparison/ embedding comparison
q7-recovery-loop/ suggestions, re-retrieval and the final answer
Each experiment folder has its own report.md and an offline check script; the numbers on this page come from the results.json there.