Leaner Luna
How far can we cut what Luna costs, and how fast it answers, without losing answer quality? We traced where its tokens go, tried fourteen changes, and tested the survivors on all 150 real questions.
The lean setup cuts Luna's search-decision cost by 41%, with no loss in answer quality our tests could detect.
- $0.017 instead of $0.028 per turn for Luna's decisions; a whole turn costs $0.032 against production's $0.042.
- Answers as good: on all 150 questions blind judges preferred the lean setup in 52% of decided pairs against tuned Luna (58 won, 53 lost), a tie; the 95% range, 43%–61%, does not rule out a small loss. Against production it lands where tuned Luna does.
- A little faster: 7.3 s to the first word against 8.1 s for tuned Luna (16 questions one at a time, same settings); jev took 6.8 s.
- Still dearer than jev: on the 48 questions both ran, $0.031 a turn against $0.025. Luna would need about $0.066 per million tokens, not $0.10, to match it.
Most of the bill was questions, not text
Luna charges about 145 tokens for every question in a request, on top of the text. Grading about 150 candidate passages per turn, each with three questions, made up nearly two thirds of the bill, and two of those three questions were never read by the pipeline. The shortlist review asked one question per sentence.
Six changes made it in
Each grading change was first tested on 205 searches whose passages three AI judges had graded, scoring how well Luna ranks them (nDCG@20: 0.888 for tuned Luna's grading; production's Cohere reranker scores 0.78 on the same searches). Only changes that kept that score went on to full answers.
| Change | Saving | Effect on quality |
|---|---|---|
| Stop asking the two cut-off flags when grading | grading call 1190 → 751 tokens | none: the pipeline never read them |
| Shorter wording for the 0–3 relevance grade | → 654 tokens | ranking 0.881 vs 0.888, a small dip within noise; answers tied on 150 questions |
| Review asks for the first and last needed sentence, not one question per sentence | review call 3.7k → 2.0k tokens | none measured |
| Duplicate check done in code, not by Luna | coverage call 13.5k → 9.0k tokens | none: only used once the pack is full |
| Fill in a neutral answer when Luna refuses one question | failed requests 12 → 0 per 150 turns | no failed steps; a refused sentence span (about one turn in five) keeps the whole passage |
| Send up to 128 Luna calls at once instead of 48 | same tokens | search stage 4.9 s vs 5.5 s; within the noise between timing runs |
Rejected. The other cuts either lowered ranking quality clearly or saved too little to be worth it. Two lessons stand out: the context around each passage, such as the search query that found it, matters a lot to Luna's grading; and retrieval rank barely predicts which passages end up in answers, so grading fewer candidates loses good ones.
| Change | Extra saving | Why not |
|---|---|---|
| Labels-only grade wording | −8% | ranking −0.016 |
| Passages cut to 1,000 characters | −2% | ranking −0.013 |
| Passages cut to 700 characters | −13% | ranking −0.044 |
| Stripping each passage's context (the query that found it, empty fields, the preface) | −13% | ranking −0.048 |
| Eight passages per call* | −6% | ranking −0.076 on the 179 searches that completed; 27 calls refused |
| Minified JSON input* | −1% | the same small dip as the kept wording change, for too little saving |
| Retrieving fewer candidates per search | −7.5% at depth 30 | loses 8% of the passages answers use |
| Cheaper processing tier or caching | not offered | the endpoint rejects service_tier; repeats are billed in full |
Savings are per grading call, on top of the kept 654-token call; ranking changes are against 0.888 and include the kept wording's −0.007. * Tested on the older, longer grading call, so its saving does not stack.
No measurable change in answer quality
The lean setup answered all 150 real questions, and Claude Opus compared each answer blind with tuned Luna's and with production's, in both orders.
One earlier, smaller run did lose: the shorter grade wording on its own lost to tuned Luna on 48 questions (9 won, 23 lost). The full lean setup with the same wording then tied on all 150, including those 48 (17 to 19), and so did the lean setup without it, so we read that run as noise.
One axis stands out: judges found production's answers more complete than lean Luna's (54 to 29). Tuned Luna leaned the same way (45 to 32), and lean and tuned tied on completeness head to head (37 to 37), so the gap likely comes from the new pipeline's smaller evidence pack (about 24 passages against production's 42), not from the cost changes.
Grounding held: 80% of claims fully supported and 0.5% unsupported, the same as tuned Luna. In the 18-question audit, lean answers had more serious errors (7 against tuned Luna's 3; production 6, jev 5), a gap that is not significant at this size. They were mixed: missing or wrong sources, a made-up page range and a madhhab mix-up (both also made by tuned Luna), a misapplied rule, a vowelling error, and an unsafe framing of a question about a 12-year-old that several setups have failed before. None traces to a specific change, but it is worth rechecking at larger scale.
Cheaper than production, still dearer than jev
Luna's decisions are still slightly more than half of a lean turn ($0.017 of $0.032). On 16 questions run one at a time, the median first word came after 7.3 s with lean Luna, 8.1 s with tuned Luna and 6.8 s with jev. Allowing 128 Luna calls in flight instead of 48 gave 6.7 s, but jev itself moved by half a second between the two passes, so that gain is not established. With several questions at once (the 150-question run), lean and tuned Luna were within 0.3 s.
Use the lean setup for any Luna deployment
The changes live in the bench adapter and need no change to the pipeline's logic; moving them into the engine is a small job. jev remains the cheaper choice for the same quality, about half a cent less per turn. Every further Luna cut we tried cost quality, so the next big savings are a lower Luna price or Gemini, whose share doubles when its introductory price ends.
Method and limitsHow the numbers were produced.
Setups. The new search pipeline with OpenAI Decisions (gpt-6-luna) making every search decision. Tuned: Luna's probabilities mapped onto jev's scale. Lean: tuned plus the six changes above. Production: today's search, replayed on 17 September. Gemini 3.7 Flash writes every answer.
Tokens. Measured from Luna's own usage numbers, per request type, on the same questions. Price $0.10 per million input tokens; nothing else is billed.
Grading tests. 205 real searches of 50 candidate passages, graded 0–3 by three AI judges (median grade); each change was run on all 10,250 passages (the eight-per-call test completed 179 of the 205 searches).
Answers. 150 real user questions (48 for comparisons with jev), judged blind by Claude Opus in both orders; a pair counts as decided only when both orders agree; 95% Wilson intervals. Grounding audit on 18 questions.
Limits. Speed was measured from Europe, 16 questions one at a time. The higher concurrency setting was not load-tested at production volume. The API is in beta, so price and limits may change.