Leaner Luna
Our latest search pipeline, with OpenAI's Luna making the search decisions, measured against production on 150 real questions. Then: how it works, and how we cut Luna's cost by 41%.
Against production: answers about as good, twice as fast, a quarter cheaper.
- Answers: lean Luna won 46% of decided pairs against production (52 won, 62 lost), within noise of a tie (95% range 37%–55%). Judges did find production more complete (54 to 29); correctness was even (30 to 31).
- Speed: the first word arrives after 7.6 s instead of 13.3 s, and the answer is done in 14.1 s instead of 28.9 s.
- Cost: $0.032 a turn instead of $0.042; about $21k a month instead of $28k.
- Grounding: 80% of claims fully supported against 78%, and 0.5% unsupported against 5.1%.
- jev does the same for less ($0.025 a turn on 48 questions), so lean Luna is the fallback, not the first choice.
Small decisions made by Luna, one answer written by Gemini
Production hands the whole search to Gemini: it calls a search tool, reads everything that comes back, decides whether to search again, and repeats for up to 20 steps before writing. The new pipeline takes those decisions away from Gemini and gives each one to Luna as a small, typed question it answers in about a quarter of a second, many at once.
- Read the question. One Luna call answers about ten typed questions at once: what kind of question it is, its separate parts, whether it is a ruling and contested, how hard it is, and what form of reply the user wants.
- Retrieve wide. The user's own words, queries built from the triage answers, and a few queries Gemini writes are run against the library in parallel, giving about 150 candidate passages.
- Grade every candidate. Luna scores each passage 0–3 for how useful it is, one call per passage, all in parallel. This replaces Cohere and Gemini's own reading.
- Review the best 24. Code keeps the top passages and drops near-duplicates. Luna then marks which sentences of each are needed, whether a passage is cut off, and what position it takes; one more call checks whether every part of the question is answered.
- Fill gaps once. Only if that check finds a measured gap does a second, targeted round run.
- Write once. Gemini gets a pack of about 23 graded, trimmed passages and writes the answer in a single pass, instead of reading about 42 passages over many steps.
About as good as production
Each setup answered the same real questions as production, and Claude Opus compared every pair blind, in both orders; a pair counts when both orders agree.
All three land within noise of a tie. The one clear difference is completeness: judges preferred production's fuller answers (54 to 29 for lean Luna, 45 to 32 for tuned Luna), most likely because the pack holds about 24 passages against production's 42. Lean and tuned Luna tie on completeness head to head (37 to 37), so the cost cuts did not cause it; a larger pack is the lever to try.
The answer starts 6 seconds sooner and finishes in half the time
Production waits for Gemini's search loop before it can write. The new pipeline's decisions run in parallel, and Gemini starts writing as soon as the pack is ready.
A quarter cheaper than production
Gemini reads far less in the new pipeline, so its share drops from $0.037 to $0.015 a turn; Luna's decisions add $0.017.
jev makes the same decisions for less: Luna would need about $0.066 per million tokens, not $0.10, to match it.
Claims stay tied to their sources
On 18 audited questions, every factual claim in an answer was checked against the passages it cited. Lean Luna: 80% fully supported, 0.5% unsupported. Production: 78% and 5.1%. Serious errors were 7 answers against production's 6, too few to tell apart; they were mostly writing slips (a wrong source, an overstated consensus), and one unsafe framing of a question about a 12-year-old that several setups, production included, have failed before.
41% off Luna's bill, with no change in answers
Luna charges about 145 tokens for every question in a request, on top of the text, so the number of questions drove the bill. Before the cost work, grading about 150 passages per turn with three questions each was nearly two thirds of it, and two of those questions were never read by the pipeline.
| Change | Saving | Effect on quality |
|---|---|---|
| Stop asking the two cut-off flags when grading | grading call 1190 → 751 tokens | none: the pipeline never read them |
| Shorter wording for the 0–3 relevance grade | → 654 tokens | ranking 0.881 vs 0.888, a small dip within noise; answers tied on 150 questions |
| Review asks for the first and last needed sentence, not one question per sentence | review call 3.7k → 2.0k tokens | none measured |
| Duplicate check done in code, not by Luna | coverage call 13.5k → 9.0k tokens | none: only used once the pack is full |
| Fill in a neutral answer when Luna refuses one question | failed requests 12 → 0 per 150 turns | no failed steps; a refused sentence span (about one turn in five) keeps the whole passage |
| Send up to 128 Luna calls at once instead of 48 | same tokens | search stage 4.9 s vs 5.5 s; within the noise between timing runs |
Against tuned Luna on all 150 questions, the lean setup won 52% of decided pairs (58 to 53), a tie. Each grading change was first tested on 205 searches graded by three AI judges (nDCG@20 0.888 before, 0.881 after). Everything else we tried either lowered quality clearly or saved too little:
| Change | Extra saving | Why not |
|---|---|---|
| Labels-only grade wording | −8% | ranking −0.016 |
| Passages cut to 1,000 characters | −2% | ranking −0.013 |
| Passages cut to 700 characters | −13% | ranking −0.044 |
| Stripping each passage's context (the query that found it, empty fields, the preface) | −13% | ranking −0.048 |
| Eight passages per call* | −6% | ranking −0.076 on the 179 searches that completed; 27 calls refused |
| Minified JSON input* | −1% | the same small dip as the kept wording change, for too little saving |
| Retrieving fewer candidates per search | −7.5% at depth 30 | loses 8% of the passages answers use |
| Cheaper processing tier or caching | not offered | the endpoint rejects service_tier; repeats are billed in full |
Savings are per grading call, on top of the kept 654-token call; ranking changes are against 0.888 and include the kept wording's −0.007. * Tested on the older, longer grading call, so its saving does not stack.
Replace production's search with this pipeline; run it on jev, keep lean Luna as the fallback
Either decision model gives production-level answers in half the time at lower cost. jev is cheaper by about half a cent a turn. The next quality step is a larger evidence pack to close the completeness gap; the next cost step is Gemini, whose price doubles in January.
Method and limitsHow the numbers were produced.
Setups. Production: today's search, Gemini 3.7 Flash with search and expand tools and Cohere reranking, replayed on 17 September. Lean Luna: the new pipeline with OpenAI Decisions (gpt-6-luna) making every decision, probabilities mapped onto jev's scale, plus the six cost changes; run 7 October. Gemini 3.7 Flash writes every answer.
Answers. 150 real user questions (48 for comparisons with jev), judged blind by Claude Opus in both orders; a pair counts as decided only when both orders agree; 95% Wilson intervals. Grounding audit on 18 questions.
Speed and cost. Medians and means over the same 150 questions; production's from September replays on different load, so its times are indicative. List prices; Gemini cached input at its cached rate.
Grading tests. 205 real searches of 50 candidate passages, graded 0–3 by three AI judges (median grade).
Limits. 48–150 questions resolve a win rate to about ±8–15 points. Load at full production volume was not tested. The Luna API is in beta.