Grading RAG answers: GPT‑6 Astra vs. GPT‑5.6 Sol
GPT‑6 Astra finished first in all 26 matched grading batches, averaging 4.1 minutes per 30 answers to GPT‑5.6 Sol’s 7.3. They agreed on 95% of verdicts, but Sol graded the answers more strictly.
Main results and findings
Use two AI judges, not one. The result depends on which one grades. Graded by Astra alone, the RAG experiment’s current setup got 6 fewer of its 240 answers right than the original. Graded by Sol alone, it got 12 fewer. A second judge also shows where the grading is borderline (35 of 717 answers in this study).
GPT‑6 Astra took 44% less time than GPT‑5.6 Sol to grade the same answers, even though it ran at extra-high reasoning vs. Sol’s high.
Why
The RAG experiment had all of its 1,500 requests graded twice, by two AI judges working blind. GPT‑6 Astra kept returning its batches first, even though it ran at the higher reasoning setting. The session records already held how long every batch took, so this comparison needed no new model calls.
GPT‑6 Astra against GPT‑5.6 Sol
Method
- Two AI judges
- GPT‑6 Astra at extra-high reasoning and GPT‑5.6 Sol at high, each run as a Codex agent on an existing ChatGPT subscription.
- What they graded
- The RAG experiment’s 1,500 requests, 717 once identical answers were merged, in 26 batches: 23 of 30 answers, then 16, 10 and 1. Each batch went to both judges, usually about 15 seconds apart, with the same answers, rubric, reference answers and retrieved passages. Both graded blind.
- How time was measured
- Each agent’s own start-to-finish time on its batch: reading, checking the manual, grading and writing up.
- When
- September 23, 2026, during the RAG experiment.
Grading time and results
On the same batch, Sol took 1.2 to 2.6 times as long as Astra, and 3.1 times in the one batch where it got a mid-task clarification.
| Judge | Reasoning | Average | Median | Fastest | Slowest |
|---|---|---|---|---|---|
| GPT‑6 Astra | Extra high | Average247.5 s | Median239.8 s | Fastest192.0 s | Slowest329.1 s |
| GPT‑5.6 Sol | High | Average440.2 s | Median408.6 s | Fastest341.7 s | Slowest648.4 s |
| Setup | Answer model | Astra alone | Sol alone | Final |
|---|---|---|---|---|
| A · original setup | Gemini 3 Flash Preview | Astra alone207 | Sol alone204 | Final206 |
| B1 · reranker removed | Gemini 3 Flash Preview | Astra alone210 | Sol alone209 | Final211 |
| B2 · smaller chunks | Gemini 3 Flash Preview | Astra alone198 | Sol alone190 | Final198 |
| B3 · 4 context chunks, not 5 | Gemini 3 Flash Preview | Astra alone193 | Sol alone189 | Final194 |
| C · Flash-Lite, current setup | Gemini 3.1 Flash-Lite | Astra alone201 | Sol alone192 | Final201 |
They agreed on the verdict for 95% of the 717 answers (Cohen’s kappa 0.89). They split on 35 answers, where Sol gave the harsher grade 24 times vs. Astra’s 11. Most splits, 23 of 35, were only about completeness (whether an answer carried every required fact and condition).
Other checks
Astra noticed a contradiction in Ford’s manual itself. Explaining how long to drive after inflating the tires before the low-pressure light goes off, the manual says both “up to two minutes” and “at least two minutes”. Astra flagged it on all 8 answers to that question vs. Sol on 1. A fresh Sol, asked to check the manual after Astra’s flag, confirmed the conflict.
Sol checked Astra’s questions. Astra drafted 150 test questions and 12 for a trial run. Sol reviewed all 162, rejected none and edited 44, mostly to drop facts the answer key required, but the question never asked for. The final 100 were chosen from the 150.
Both passed the planted tests. Before grading began, each judge graded 13 planted test answers with known verdicts, from correct answers to deliberately wrong facts, and passed all 13. On the planted wrong facts, Astra was the stricter of the two.