Experiment

LLM as a judge: GPT‑6 Astra vs. GPT‑6.1 Sol

GPT‑6.1 Sol cost 76% less than GPT‑6 Astra to grade the same 900 question pairs, and took about the same time: 62.7 s per review vs. 61.0 s. Both matched Claude Opus 5.5’s factual verdicts 98.1% of the time. On pass/fail judgment, Astra agreed with Opus 94.4% vs. Sol’s 93.6%.

Main results and findings

GPT‑6.1 Sol cost about a quarter as much as GPT‑6 Astra for the job: $5.30 vs. $22.52. At listed API prices, a review of one question set would have cost $0.029 with Sol vs. $0.125 with Astra (76% less).

Both LLMs took about the same time. An average review took 62.7 s with Sol, and 61.0 s with Astra.

Both matched Claude Opus 5.5 on facts equally often. Each gave the same factual verdict as Opus on 883 of the 900 question pairs (98.1%).

GPT‑6 Astra matched Opus on quality slightly more often. It gave the same pass or fail as Opus on 850/900 pairs (94.4%) vs. Sol’s 842/900 (93.6%).

Both LLMs were stricter than Opus. Opus rejected 58/900 question pairs, Astra 92/900, and Sol 102/900.

Why this study

DaycareFacts, one of our projects, analyzes and publishes inspection and licensing information on thousands of childcare providers. For each provider, we wanted an LLM to write five questions, each with a short reason, that a parent could ask the provider about its inspection findings (or general topics, where a provider has fewer than five findings). Every set has to be graded before a parent sees it. At that scale, the judge’s cost matters as much as its accuracy.

We used two AI judges, Claude Opus 5.5 and GPT‑6 Astra, to grade the question sets in our Gemini 3.1 Flash‑Lite vs. Gemma 4 E4B study. At listed API prices, Astra’s 896 reviews would have cost $110.47.

GPT‑6.1 Sol’s tokens cost 1/5 of Astra’s. We wanted to know whether it could do the same grading for much less, and whether the speed and results would differ.

We gave GPT‑6 Astra and GPT‑6.1 Sol the same 180 question sets from that study, at the same moment, and compared their verdicts with the ones Opus 5.5 had already given.

GPT‑6.1 Sol vs. GPT‑6 Astra

$0.029
LLM cost per review
76% less than Astra’s $0.125
62.7 s
Average time per review
Astra 61.0 s
Quality verdicts matching Opus · Astra: 94.4% 93.6%
Factual verdicts matching Opus · Astra: 98.1% 98.1%
Reasoning tokens per review · Astra: 301 648

Method

Two LLM judges
GPT‑6 Astra and GPT‑6.1 Sol at high reasoning, each run through Codex on a ChatGPT subscription.
Reference
Claude Opus 5.5 - high reasoning verdicts given earlier on the same question pairs for the Flash‑Lite vs. Gemma 4 E4B study. No human grading.
What they graded
180 total question sets from that study (60 each from Florida, California and Texas): 900 question pairs from 150 childcare providers. The sets were drawn reproducibly and balanced by length, without looking at grades or which model wrote them.
How they ran
Each set went to both judges at the same moment, 0.15 s apart on average, with the same instructions and evidence. Each judge worked in a fresh session with no tools or web access, without knowing which model wrote a set or what the other judges said.
Pass definition
A question pair (one question and its reason) passes if a parent could use it, it covers its assigned finding or topic, does not make up information, doesn’t repeat anything, and does not imply an old problem still exists. A factual verdict says whether the evidence the model was given supports the pair, contradicts it, or does not establish it.
How time was measured
Codex’s own record of each full review, from request to verdicts on all five pairs.
When
October 3, 2026.
Cost basis
List-price estimates from the tokens each review used, at OpenAI’s API prices on October 3, 2026. Both judges ran on a ChatGPT subscription.

Grading: GPT‑6 Astra vs. GPT‑6.1 Sol

Both judges gave each question pair two verdicts: pass or fail on quality, and whether the evidence supported it. We compared each verdict with Opus’s for the same pair.

Table 1. Verdicts compared with Claude Opus 5.5
Judge Reasoning Pairs passed / 900 Quality matches / 900 Factual matches / 900
Claude Opus 5.5 High Pairs passed / 900842 Quality matches / 900Reference Factual matches / 900Reference
GPT‑6 Astra High Pairs passed / 900808 Quality matches / 900850 (94.4%) Factual matches / 900883 (98.1%)
GPT‑6.1 Sol High Pairs passed / 900798 Quality matches / 900842 (93.6%) Factual matches / 900883 (98.1%)
900 question pairs from 180 sets. Opus’s verdicts are the reference, given earlier for the Flash‑Lite vs. Gemma 4 E4B study. Opus found 876 of the 900 pairs supported by the evidence, so most factual matches are pairs that both Opus and the judge found supported. Download CSV

Time and cost

GPT‑6.1 Sol reasoned twice as much, yet still cost about a quarter of what Astra cost. A Sol review used 648 reasoning tokens on average vs. 301 for Astra, and the written verdicts were about the same length: 1,244 tokens vs. 1,212. Sol’s list prices are a fifth of Astra’s: $2 vs. $10 per million input tokens, and $10 vs. $50 per million output tokens.

Table 2. Time and cost per review
Judge Reasoning Average time Median time Cost per review Total / 180
GPT‑6 Astra High Average time61.0 s Median time55.6 s Cost per review$0.125 Total / 180$22.52
GPT‑6.1 Sol High Average time62.7 s Median time57.5 s Cost per review$0.029 Total / 180$5.30
One review covers one set of five question pairs. Both judges got each set at the same moment. Cost is API-equivalent at October 3, 2026 list prices; both judges ran on a ChatGPT subscription. Download CSV

Practical takeaways

In light of OpenAI’s recent announcement that it will cut its Pro 200 plan’s usage limits in half, and its claim that GPT‑6.1 Sol is nearly as good as the flagship GPT‑6 Astra, allowing subscribers to do as much work as before (even with the plan cuts), this was a great opportunity to see if GPT‑6.1 Sol was a viable replacement.

For this specific task, GPT‑6.1 Sol delivered essentially the same speed and results as GPT‑6 Astra, for about a quarter of the cost.

GPT‑6.1 Sol is the clear winner of this test.