Qaf search · Luna cost study · 8 October 2026

Leaner Luna

Our latest search pipeline, with OpenAI's Luna making the search decisions, measured against production on 150 real questions. Then: how it works, and how we cut Luna's cost by 41%.

Against production: answers about as good, twice as fast, a quarter cheaper.

Answers vs production46%decided pairs, 150 questions; 50% is a tie
First word7.6 s vs 13.3 slean Luna vs production, median
Cost per turn$0.032 vs $0.042lean Luna vs production · jev $0.025
Passage ranking0.881 vs 0.78Luna's grades vs production's Cohere reranker, nDCG@20
1 · How it works

Small decisions made by Luna, one answer written by Gemini

Production hands the whole search to Gemini: it calls a search tool, reads everything that comes back, decides whether to search again, and repeats for up to 20 steps before writing. The new pipeline takes those decisions away from Gemini and gives each one to Luna as a small, typed question it answers in about a quarter of a second, many at once.

Production today Question Gemini, in a loopdecides what to search, up to 20 steps Search toolvector top 50 → Cohere → 20 Expand tool± 5 neighbouring passages calls every result back Gemini writes the answerfrom ~42 passages when Gemini stops searching New pipeline, lean Luna Question Luna triage1 call: type, parts, difficulty… Retrieve candidatesquestion + templates + Gemini-written queries templates Luna grades each~150 calls · 1 question ~150 Shortlist 24near-duplicates out Luna review24 calls · 8 questions Luna coverage1 call · parts answered? Pack ~23trimmed to need Gemini writes oncefrom the graded pack round 2, only on a measured gap: a part unanswered, a hadith without its source, a passage cut off
Luna callGemini callCohere rerankerPlain code
  1. Read the question. One Luna call answers about ten typed questions at once: what kind of question it is, its separate parts, whether it is a ruling and contested, how hard it is, and what form of reply the user wants.
  2. Retrieve wide. The user's own words, queries built from the triage answers, and a few queries Gemini writes are run against the library in parallel, giving about 150 candidate passages.
  3. Grade every candidate. Luna scores each passage 0–3 for how useful it is, one call per passage, all in parallel. This replaces Cohere and Gemini's own reading.
  4. Review the best 24. Code keeps the top passages and drops near-duplicates. Luna then marks which sentences of each are needed, whether a passage is cut off, and what position it takes; one more call checks whether every part of the question is answered.
  5. Fill gaps once. Only if that check finds a measured gap does a second, targeted round run.
  6. Write once. Gemini gets a pack of about 23 graded, trimmed passages and writes the answer in a single pass, instead of reading about 42 passages over many steps.
2 · Answers

About as good as production

Each setup answered the same real questions as production, and Claude Opus compared every pair blind, in both orders; a pair counts when both orders agree.

Share of decided pairs won against productionDot is the win rate, line its 95% interval.

All three land within noise of a tie. The one clear difference is completeness: judges preferred production's fuller answers (54 to 29 for lean Luna, 45 to 32 for tuned Luna), most likely because the pack holds about 24 passages against production's 42. Lean and tuned Luna tie on completeness head to head (37 to 37), so the cost cuts did not cause it; a larger pack is the lever to try.

3 · Speed

The answer starts 6 seconds sooner and finishes in half the time

Production waits for Gemini's search loop before it can write. The new pipeline's decisions run in parallel, and Gemini starts writing as soon as the pack is ready.

Median time per questionAll 150 questions. Production times are September replays of the same questions; lean Luna ran six to eight questions at a time.
4 · Cost

A quarter cheaper than production

Gemini reads far less in the new pipeline, so its share drops from $0.037 to $0.015 a turn; Luna's decisions add $0.017.

Mean cost per answered questionList prices, with Gemini at its introductory rate (it doubles on 1 January 2027); monthly at 22,000 questions a day. Luna and production on 150 questions, jev on 48 (lean Luna on those 48: $0.031).

jev makes the same decisions for less: Luna would need about $0.066 per million tokens, not $0.10, to match it.

5 · Grounding

Claims stay tied to their sources

On 18 audited questions, every factual claim in an answer was checked against the passages it cited. Lean Luna: 80% fully supported, 0.5% unsupported. Production: 78% and 5.1%. Serious errors were 7 answers against production's 6, too few to tell apart; they were mostly writing slips (a wrong source, an overstated consensus), and one unsafe framing of a question about a 12-year-old that several setups, production included, have failed before.

6 · How we cut Luna's cost

41% off Luna's bill, with no change in answers

Luna charges about 145 tokens for every question in a request, on top of the text, so the number of questions drove the bill. Before the cost work, grading about 150 passages per turn with three questions each was nearly two thirds of it, and two of those questions were never read by the pipeline.

Luna tokens per turn, by pipeline stepSame 12 questions through each setup.
ChangeSavingEffect on quality
Stop asking the two cut-off flags when gradinggrading call 1190 → 751 tokensnone: the pipeline never read them
Shorter wording for the 0–3 relevance grade→ 654 tokensranking 0.881 vs 0.888, a small dip within noise; answers tied on 150 questions
Review asks for the first and last needed sentence, not one question per sentencereview call 3.7k → 2.0k tokensnone measured
Duplicate check done in code, not by Lunacoverage call 13.5k → 9.0k tokensnone: only used once the pack is full
Fill in a neutral answer when Luna refuses one questionfailed requests 12 → 0 per 150 turnsno failed steps; a refused sentence span (about one turn in five) keeps the whole passage
Send up to 128 Luna calls at once instead of 48same tokenssearch stage 4.9 s vs 5.5 s; within the noise between timing runs

Against tuned Luna on all 150 questions, the lean setup won 52% of decided pairs (58 to 53), a tie. Each grading change was first tested on 205 searches graded by three AI judges (nDCG@20 0.888 before, 0.881 after). Everything else we tried either lowered quality clearly or saved too little:

ChangeExtra savingWhy not
Labels-only grade wording−8%ranking −0.016
Passages cut to 1,000 characters−2%ranking −0.013
Passages cut to 700 characters−13%ranking −0.044
Stripping each passage's context (the query that found it, empty fields, the preface)−13%ranking −0.048
Eight passages per call*−6%ranking −0.076 on the 179 searches that completed; 27 calls refused
Minified JSON input*−1%the same small dip as the kept wording change, for too little saving
Retrieving fewer candidates per search−7.5% at depth 30loses 8% of the passages answers use
Cheaper processing tier or cachingnot offeredthe endpoint rejects service_tier; repeats are billed in full

Savings are per grading call, on top of the kept 654-token call; ranking changes are against 0.888 and include the kept wording's −0.007. * Tested on the older, longer grading call, so its saving does not stack.

7 · Recommendation

Replace production's search with this pipeline; run it on jev, keep lean Luna as the fallback

Either decision model gives production-level answers in half the time at lower cost. jev is cheaper by about half a cent a turn. The next quality step is a larger evidence pack to close the completeness gap; the next cost step is Gemini, whose price doubles in January.

Method and limitsHow the numbers were produced.

Setups. Production: today's search, Gemini 3.7 Flash with search and expand tools and Cohere reranking, replayed on 17 September. Lean Luna: the new pipeline with OpenAI Decisions (gpt-6-luna) making every decision, probabilities mapped onto jev's scale, plus the six cost changes; run 7 October. Gemini 3.7 Flash writes every answer.

Answers. 150 real user questions (48 for comparisons with jev), judged blind by Claude Opus in both orders; a pair counts as decided only when both orders agree; 95% Wilson intervals. Grounding audit on 18 questions.

Speed and cost. Medians and means over the same 150 questions; production's from September replays on different load, so its times are indicative. List prices; Gemini cached input at its cached rate.

Grading tests. 205 real searches of 50 candidate passages, graded 0–3 by three AI judges (median grade).

Limits. 48–150 questions resolve a win rate to about ±8–15 points. Load at full production volume was not tested. The Luna API is in beta.