Case study

Grading RAG answers: GPT‑6 Astra vs. GPT‑5.6 Sol

GPT‑6 Astra finished first in all 26 matched grading batches, averaging 4.1 minutes per 30 answers to GPT‑5.6 Sol’s 7.3. They agreed on 95% of verdicts, but Sol graded the answers more strictly.

Main results and findings

Use two AI judges, not one. The result depends on which one grades. Graded by Astra alone, the RAG experiment’s current setup got 6 fewer of its 240 answers right than the original. Graded by Sol alone, it got 12 fewer. A second judge also shows where the grading is borderline (35 of 717 answers in this study).

GPT‑6 Astra took 44% less time than GPT‑5.6 Sol to grade the same answers, even though it ran at extra-high reasoning vs. Sol’s high.

Why

The RAG experiment had all of its 1,500 requests graded twice, by two AI judges working blind. GPT‑6 Astra kept returning its batches first, even though it ran at the higher reasoning setting. The session records already held how long every batch took, so this comparison needed no new model calls.

GPT‑6 Astra against GPT‑5.6 Sol

4.1 min
Astra’s average per 30 answers
Sol 7.3 min; 44% less time
26 of 26
Batches Astra finished first
Its slowest full batch beat Sol’s fastest
Astra per answer, all 26 batches · Sol 14.9 s 8.4 s
Astra median batch · Sol 6.8 min 4.0 min
Verdicts the two agreed on 95%
Planted tests passed by each judge 13 of 13

Method

Two AI judges
GPT‑6 Astra at extra-high reasoning and GPT‑5.6 Sol at high, each run as a Codex agent on an existing ChatGPT subscription.
What they graded
The RAG experiment’s 1,500 requests, 717 once identical answers were merged, in 26 batches: 23 of 30 answers, then 16, 10 and 1. Each batch went to both judges, usually about 15 seconds apart, with the same answers, rubric, reference answers and retrieved passages. Both graded blind.
How time was measured
Each agent’s own start-to-finish time on its batch: reading, checking the manual, grading and writing up.
When
September 23, 2026, during the RAG experiment.

Grading time and results

On the same batch, Sol took 1.2 to 2.6 times as long as Astra, and 3.1 times in the one batch where it got a mid-task clarification.

Table 1. Seconds per batch of 30 answers
Judge Reasoning Average Median Fastest Slowest
GPT‑6 Astra Extra high Average247.5 s Median239.8 s Fastest192.0 s Slowest329.1 s
GPT‑5.6 Sol High Average440.2 s Median408.6 s Fastest341.7 s Slowest648.4 s
23 full batches per judge; Astra also finished first on the last three, of 16, 10 and 1 answers. Times cover the whole job, not just the model’s reply. Sol’s slowest, 648.4 s, is the one batch where it got a mid-task clarification. In the 8 full batches that overlapped none of the other 25, Astra still took 39% less time and finished first in all 8. Download CSV
Table 2. Correct answers out of 240
Setup Answer model Astra alone Sol alone Final
A · original setup Gemini 3 Flash Preview Astra alone207 Sol alone204 Final206
B1 · reranker removed Gemini 3 Flash Preview Astra alone210 Sol alone209 Final211
B2 · smaller chunks Gemini 3 Flash Preview Astra alone198 Sol alone190 Final198
B3 · 4 context chunks, not 5 Gemini 3 Flash Preview Astra alone193 Sol alone189 Final194
C · Flash-Lite, current setup Gemini 3.1 Flash-Lite Astra alone201 Sol alone192 Final201
The RAG experiment’s five setups, each answering its 80 answerable questions 3 times. Final is the count the RAG study published, after a fresh Astra settled every disagreement; it sided with Astra in 30 of 35 splits, so Final is not independent of Astra. Tire-pressure answers are counted as re-graded. Download CSV

They agreed on the verdict for 95% of the 717 answers (Cohen’s kappa 0.89). They split on 35 answers, where Sol gave the harsher grade 24 times vs. Astra’s 11. Most splits, 23 of 35, were only about completeness (whether an answer carried every required fact and condition).

Other checks

Astra noticed a contradiction in Ford’s manual itself. Explaining how long to drive after inflating the tires before the low-pressure light goes off, the manual says both “up to two minutes” and “at least two minutes”. Astra flagged it on all 8 answers to that question vs. Sol on 1. A fresh Sol, asked to check the manual after Astra’s flag, confirmed the conflict.

Sol checked Astra’s questions. Astra drafted 150 test questions and 12 for a trial run. Sol reviewed all 162, rejected none and edited 44, mostly to drop facts the answer key required, but the question never asked for. The final 100 were chosen from the 150.

Both passed the planted tests. Before grading began, each judge graded 13 planted test answers with known verdicts, from correct answers to deliberately wrong facts, and passed all 13. On the planted wrong facts, Astra was the stricter of the two.