LLM as a judge: GPT‑6 Astra vs. GPT‑6.1 Sol
GPT‑6.1 Sol cost 76% less than GPT‑6 Astra to grade the same 900 question pairs, and took about the same time: 62.7 s per review vs. 61.0 s. Both matched Claude Opus 5.5’s factual verdicts 98.1% of the time. On pass/fail judgment, Astra agreed with Opus 94.4% vs. Sol’s 93.6%.
Main results and findings
GPT‑6.1 Sol cost about a quarter as much as GPT‑6 Astra for the job: $5.30 vs. $22.52. At listed API prices, a review of one question set would have cost $0.029 with Sol vs. $0.125 with Astra (76% less).
Both LLMs took about the same time. An average review took 62.7 s with Sol, and 61.0 s with Astra.
Both matched Claude Opus 5.5 on facts equally often. Each gave the same factual verdict as Opus on 883 of the 900 question pairs (98.1%).
GPT‑6 Astra matched Opus on quality slightly more often. It gave the same pass or fail as Opus on 850/900 pairs (94.4%) vs. Sol’s 842/900 (93.6%).
Both LLMs were stricter than Opus. Opus rejected 58/900 question pairs, Astra 92/900, and Sol 102/900.
Why this study
DaycareFacts, one of our projects, analyzes and publishes inspection and licensing information on thousands of childcare providers. For each provider, we wanted an LLM to write five questions, each with a short reason, that a parent could ask the provider about its inspection findings (or general topics, where a provider has fewer than five findings). Every set has to be graded before a parent sees it. At that scale, the judge’s cost matters as much as its accuracy.
We used two AI judges, Claude Opus 5.5 and GPT‑6 Astra, to grade the question sets in our Gemini 3.1 Flash‑Lite vs. Gemma 4 E4B study. At listed API prices, Astra’s 896 reviews would have cost $110.47.
GPT‑6.1 Sol’s tokens cost 1/5 of Astra’s. We wanted to know whether it could do the same grading for much less, and whether the speed and results would differ.
We gave GPT‑6 Astra and GPT‑6.1 Sol the same 180 question sets from that study, at the same moment, and compared their verdicts with the ones Opus 5.5 had already given.
GPT‑6.1 Sol vs. GPT‑6 Astra
Method
- Two LLM judges
- GPT‑6 Astra and GPT‑6.1 Sol at high reasoning, each run through Codex on a ChatGPT subscription.
- Reference
- Claude Opus 5.5 - high reasoning verdicts given earlier on the same question pairs for the Flash‑Lite vs. Gemma 4 E4B study. No human grading.
- What they graded
- 180 total question sets from that study (60 each from Florida, California and Texas): 900 question pairs from 150 childcare providers. The sets were drawn reproducibly and balanced by length, without looking at grades or which model wrote them.
- How they ran
- Each set went to both judges at the same moment, 0.15 s apart on average, with the same instructions and evidence. Each judge worked in a fresh session with no tools or web access, without knowing which model wrote a set or what the other judges said.
- Pass definition
- A question pair (one question and its reason) passes if a parent could use it, it covers its assigned finding or topic, does not make up information, doesn’t repeat anything, and does not imply an old problem still exists. A factual verdict says whether the evidence the model was given supports the pair, contradicts it, or does not establish it.
- How time was measured
- Codex’s own record of each full review, from request to verdicts on all five pairs.
- When
- October 3, 2026.
- Cost basis
- List-price estimates from the tokens each review used, at OpenAI’s API prices on October 3, 2026. Both judges ran on a ChatGPT subscription.
Grading: GPT‑6 Astra vs. GPT‑6.1 Sol
Both judges gave each question pair two verdicts: pass or fail on quality, and whether the evidence supported it. We compared each verdict with Opus’s for the same pair.
| Judge | Reasoning | Pairs passed / 900 | Quality matches / 900 | Factual matches / 900 |
|---|---|---|---|---|
| Claude Opus 5.5 | High | Pairs passed / 900842 | Quality matches / 900Reference | Factual matches / 900Reference |
| GPT‑6 Astra | High | Pairs passed / 900808 | Quality matches / 900850 (94.4%) | Factual matches / 900883 (98.1%) |
| GPT‑6.1 Sol | High | Pairs passed / 900798 | Quality matches / 900842 (93.6%) | Factual matches / 900883 (98.1%) |
Time and cost
GPT‑6.1 Sol reasoned twice as much, yet still cost about a quarter of what Astra cost. A Sol review used 648 reasoning tokens on average vs. 301 for Astra, and the written verdicts were about the same length: 1,244 tokens vs. 1,212. Sol’s list prices are a fifth of Astra’s: $2 vs. $10 per million input tokens, and $10 vs. $50 per million output tokens.
| Judge | Reasoning | Average time | Median time | Cost per review | Total / 180 |
|---|---|---|---|---|---|
| GPT‑6 Astra | High | Average time61.0 s | Median time55.6 s | Cost per review$0.125 | Total / 180$22.52 |
| GPT‑6.1 Sol | High | Average time62.7 s | Median time57.5 s | Cost per review$0.029 | Total / 180$5.30 |
Practical takeaways
In light of OpenAI’s recent announcement that it will cut its Pro 200 plan’s usage limits in half, and its claim that GPT‑6.1 Sol is nearly as good as the flagship GPT‑6 Astra, allowing subscribers to do as much work as before (even with the plan cuts), this was a great opportunity to see if GPT‑6.1 Sol was a viable replacement.
For this specific task, GPT‑6.1 Sol delivered essentially the same speed and results as GPT‑6 Astra, for about a quarter of the cost.
GPT‑6.1 Sol is the clear winner of this test.