Qaf search · Luna cost study · 7 October 2026

Leaner Luna

How far can we cut what Luna costs, and how fast it answers, without losing answer quality? We traced where its tokens go, tried fourteen changes, and tested the survivors on all 150 real questions.

The lean setup cuts Luna's search-decision cost by 41%, with no loss in answer quality our tests could detect.

Luna cost per turn$0.017 vs $0.028lean vs tuned, 150 questions
Whole turn$0.032 vs $0.042lean Luna vs production, 150 questions · jev $0.025 on 48
Answers vs tuned52%decided pairs, 150 questions; 50% is a tie
First word7.3 s vs 8.1 slean vs tuned, 16 questions one at a time
1 · Where the tokens went

Most of the bill was questions, not text

Luna charges about 145 tokens for every question in a request, on top of the text. Grading about 150 candidate passages per turn, each with three questions, made up nearly two thirds of the bill, and two of those three questions were never read by the pipeline. The shortlist review asked one question per sentence.

Luna tokens per turn, by pipeline stepSame 12 questions through each setup.
2 · What we changed

Six changes made it in

Each grading change was first tested on 205 searches whose passages three AI judges had graded, scoring how well Luna ranks them (nDCG@20: 0.888 for tuned Luna's grading; production's Cohere reranker scores 0.78 on the same searches). Only changes that kept that score went on to full answers.

ChangeSavingEffect on quality
Stop asking the two cut-off flags when gradinggrading call 1190 → 751 tokensnone: the pipeline never read them
Shorter wording for the 0–3 relevance grade→ 654 tokensranking 0.881 vs 0.888, a small dip within noise; answers tied on 150 questions
Review asks for the first and last needed sentence, not one question per sentencereview call 3.7k → 2.0k tokensnone measured
Duplicate check done in code, not by Lunacoverage call 13.5k → 9.0k tokensnone: only used once the pack is full
Fill in a neutral answer when Luna refuses one questionfailed requests 12 → 0 per 150 turnsno failed steps; a refused sentence span (about one turn in five) keeps the whole passage
Send up to 128 Luna calls at once instead of 48same tokenssearch stage 4.9 s vs 5.5 s; within the noise between timing runs

Rejected. The other cuts either lowered ranking quality clearly or saved too little to be worth it. Two lessons stand out: the context around each passage, such as the search query that found it, matters a lot to Luna's grading; and retrieval rank barely predicts which passages end up in answers, so grading fewer candidates loses good ones.

ChangeExtra savingWhy not
Labels-only grade wording−8%ranking −0.016
Passages cut to 1,000 characters−2%ranking −0.013
Passages cut to 700 characters−13%ranking −0.044
Stripping each passage's context (the query that found it, empty fields, the preface)−13%ranking −0.048
Eight passages per call*−6%ranking −0.076 on the 179 searches that completed; 27 calls refused
Minified JSON input*−1%the same small dip as the kept wording change, for too little saving
Retrieving fewer candidates per search−7.5% at depth 30loses 8% of the passages answers use
Cheaper processing tier or cachingnot offeredthe endpoint rejects service_tier; repeats are billed in full

Savings are per grading call, on top of the kept 654-token call; ranking changes are against 0.888 and include the kept wording's −0.007. * Tested on the older, longer grading call, so its saving does not stack.

3 · Quality

No measurable change in answer quality

The lean setup answered all 150 real questions, and Claude Opus compared each answer blind with tuned Luna's and with production's, in both orders.

Lean against tuned LunaShare of decided pairs won.

One earlier, smaller run did lose: the shorter grade wording on its own lost to tuned Luna on 48 questions (9 won, 23 lost). The full lean setup with the same wording then tied on all 150, including those 48 (17 to 19), and so did the lean setup without it, so we read that run as noise.

Against productionOverall, lean Luna lands where tuned Luna and jev do.

One axis stands out: judges found production's answers more complete than lean Luna's (54 to 29). Tuned Luna leaned the same way (45 to 32), and lean and tuned tied on completeness head to head (37 to 37), so the gap likely comes from the new pipeline's smaller evidence pack (about 24 passages against production's 42), not from the cost changes.

Grounding held: 80% of claims fully supported and 0.5% unsupported, the same as tuned Luna. In the 18-question audit, lean answers had more serious errors (7 against tuned Luna's 3; production 6, jev 5), a gap that is not significant at this size. They were mixed: missing or wrong sources, a made-up page range and a madhhab mix-up (both also made by tuned Luna), a misapplied rule, a vowelling error, and an unsafe framing of a question about a 12-year-old that several setups have failed before. None traces to a specific change, but it is worth rechecking at larger scale.

4 · Cost and speed

Cheaper than production, still dearer than jev

Luna's decisions are still slightly more than half of a lean turn ($0.017 of $0.032). On 16 questions run one at a time, the median first word came after 7.3 s with lean Luna, 8.1 s with tuned Luna and 6.8 s with jev. Allowing 128 Luna calls in flight instead of 48 gave 6.7 s, but jev itself moved by half a second between the two passes, so that gain is not established. With several questions at once (the 150-question run), lean and tuned Luna were within 0.3 s.

Mean cost per answered questionList prices, with Gemini at its introductory rate (it doubles on 1 January 2027); monthly at 22,000 questions a day. Luna and production on 150 questions, jev on 48 (lean Luna on those 48: $0.031).
5 · Recommendation

Use the lean setup for any Luna deployment

The changes live in the bench adapter and need no change to the pipeline's logic; moving them into the engine is a small job. jev remains the cheaper choice for the same quality, about half a cent less per turn. Every further Luna cut we tried cost quality, so the next big savings are a lower Luna price or Gemini, whose share doubles when its introductory price ends.

Method and limitsHow the numbers were produced.

Setups. The new search pipeline with OpenAI Decisions (gpt-6-luna) making every search decision. Tuned: Luna's probabilities mapped onto jev's scale. Lean: tuned plus the six changes above. Production: today's search, replayed on 17 September. Gemini 3.7 Flash writes every answer.

Tokens. Measured from Luna's own usage numbers, per request type, on the same questions. Price $0.10 per million input tokens; nothing else is billed.

Grading tests. 205 real searches of 50 candidate passages, graded 0–3 by three AI judges (median grade); each change was run on all 10,250 passages (the eight-per-call test completed 179 of the 205 searches).

Answers. 150 real user questions (48 for comparisons with jev), judged blind by Claude Opus in both orders; a pair counts as decided only when both orders agree; 95% Wilson intervals. Grounding audit on 18 questions.

Limits. Speed was measured from Europe, 16 questions one at a time. The higher concurrency setting was not load-tested at production volume. The API is in beta, so price and limits may change.