GPT‑6.1 Sol cost 76% less than GPT‑6 Astra to grade the same 900 question pairs, and took about the same time: 62.7 s per review vs. 61.0 s. Both matched Claude Opus 5.5’s factual verdicts 98.1% of the time. On pass/fail judgment, Astra agreed with Opus 94.4% vs. Sol’s 93.6%.
Published October 5, 2026
Gemini 3.1 Flash‑Lite wrote complete, usable question sets for 90% of 300 childcare providers, grounded in official state inspection records processed by DaycareFacts.com. Gemma 4 E4B with thinking on, run locally on a Mac Mini (M4 Pro, 24 GB RAM): 73%, and with thinking off: 65%. Gemini was also substantially faster, with 1.81 s per set vs. 7.45 s and 4.64 s, respectively.
Published October 3, 2026
Average response time fell from 9.1 s to 1.6 s, and LLM cost from $4.03 to $0.46 per 1,000 questions. The difference in answer accuracy was too small to call either way.
Published September 25, 2026