Experiment

Hosted vs. local LLM: Gemini 3.1 Flash‑Lite vs. Gemma 4 E4B

Gemini 3.1 Flash‑Lite wrote complete, usable question sets for 90% of 300 childcare providers, grounded in official state inspection records processed by DaycareFacts.com. Gemma 4 E4B with thinking on, run locally on a Mac Mini (M4 Pro, 24 GB RAM): 73%, and with thinking off: 65%. Gemini was also substantially faster, with 1.81 s per set vs. 7.45 s and 4.64 s, respectively.

Main results and findings

Gemini 3.1 Flash‑Lite wrote the most usable questions. 270 of its 300 question sets passed, vs. 220 for Gemma 4 with thinking on, and 196 with thinking off, as graded by two AI judges at high reasoning, Claude Opus 5.5 and GPT‑6 Astra, with a third breaking ties.

Gemini 3.1 Flash‑Lite also made the fewest factual errors. The judges found 2 in Flash‑Lite’s 1,500 question pairs, vs. 36 and 54 for Gemma 4 E4B.

Gemini 3.1 Flash‑Lite was the fastest. A set of five question pairs took 1.81 s on average, vs. 7.45 s for Gemma 4 with thinking on, and 4.64 s with thinking off. Writing all 300 sets with Gemini cost an estimated $0.25. Gemini ran on Google’s servers, and Gemma 4 on a local Mac Mini.

Thinking improved Gemma 4’s results, but not enough to close the gap with Gemini. With thinking on, 24 more of its sets passed and it made fewer factual errors, 36 vs. 54. It also took 61% more time per set, and 4 of its sets came back broken.

Questions built on inspection findings were the challenging part. Gemini 3.1 Flash‑Lite passed 95% of them, vs. 86% and 82% for Gemma 4 E4B. On general topics, all three setups passed 97% or more.

Why this study

DaycareFacts, one of our projects, publishes inspection and licensing information on childcare providers. We wanted an LLM to write five questions, each with a short reason, a parent could ask a childcare provider based on their inspection findings. A question should name the actual finding, if any, without implying the problem still exists today. The other questions cover general topics.

We wanted to know whether a small LLM running on our own machine could do this as well as a hosted model. We ran Gemma 4 E4B locally, with thinking on and off, against Gemini 3.1 Flash‑Lite through Google’s API, using the same 300 childcare providers.

The question sets were written on September 10, 2026, and graded blind on October 2–3 by two AI judges, Claude Opus 5.5 and GPT‑6 Astra, whose comparison became a second part of this study.

Gemini 3.1 Flash‑Lite vs. Gemma 4 E4B

90%
Question sets passed
Gemma 4 73% with thinking on, 65% with it off
1.81 s
Average time per set
Gemma 4 7.45 s with thinking on, 4.64 s with it off
Question pairs passed · Gemma 4: 92% and 90% 98%
Factual errors found · Gemma 4: 36 and 54 2
Gemini LLM cost for all 300 sets $0.25

Method

Three setups
Gemini 3.1 Flash‑Lite through Google’s API, thinking off. Gemma 4 E4B on a local Mac Mini through Ollama, once with thinking on and once with it off. Same instructions for all three, temperature 0.
Test set
300 childcare providers, 100 each from Florida, California and Texas, drawn reproducibly from 54,645 across provider types and numbers of inspection findings. Each setup wrote one set of five question pairs per provider: 900 sets, 4,500 pairs.
What each model was given
The provider’s name, type and state, plus five question slots: up to five most recent distinct inspection findings, and general topics for the rest.
Pass definition
A question pair (one question and its reason) passes if a parent could use it, it covers its assigned finding or topic, does not make up information, doesn’t repeat anything, and does not imply an old problem still exists. A set passes only if all five pairs pass. A factual error is a claim not supported by, or contradicted by, the evidence the model was given.
How sets were graded
Blind, by Claude Opus 5.5 and GPT‑6 Astra at high reasoning, each in a fresh session with the original instructions and evidence, and no tools or web access. Neither was told which model wrote a set, or what the other judge said. Disagreements were settled by Gemini 3.8 Flash, which reviewed the full set on its own, and two of three votes decided. No human grading.
When
Sets written September 10, 2026, one request at a time. Graded October 2–3, 2026.
Cost basis
List-price estimates: Gemini API charges for writing the sets, and API-equivalent prices on October 3, 2026 for the judges, which ran on subscriptions. Gemma 4’s local running cost was not measured.

Three setups · 300 providers

Table 1. Question sets that passed, by setup
Setup Runs on Sets passed / 300 Pairs passed / 1,500 Factual errors Average time per set
Gemini 3.1 Flash‑Lite · thinking off Google API Sets passed / 300270 Pairs passed / 1,5001,466 Factual errors2 Average time per set1.81 s
Gemma 4 E4B · thinking on Local Mac Mini Sets passed / 300220 Pairs passed / 1,5001,381 Factual errors36 Average time per set7.45 s
Gemma 4 E4B · thinking off Local Mac Mini Sets passed / 300196 Pairs passed / 1,5001,357 Factual errors54 Average time per set4.64 s
A set passes only if all five of its question pairs pass. 4 of Gemma 4’s sets with thinking on came back broken and count as failed, so its factual errors are out of 1,480 pairs, not 1,500. Undecided grades, from a missing or split tie-break vote: 1 set and 7 pairs on quality, 8 pairs on factual errors. The ranking holds whichever way they fall. Time is from the request to the full set of five. Download CSV
Table 2. Question sets that passed, by state
Setup Runs on Florida / 100 California / 100 Texas / 100
Gemini 3.1 Flash‑Lite · thinking off Google API Florida / 10086 California / 10097 Texas / 10087
Gemma 4 E4B · thinking on Local Mac Mini Florida / 10070 California / 10076 Texas / 10074
Gemma 4 E4B · thinking off Local Mac Mini Florida / 10069 California / 10065 Texas / 10062
100 providers per state. Texas has 1 undecided Flash‑Lite set. Between Gemma 4 with and without thinking, the difference in each state is too small to call. Download CSV
Table 3. Question pairs that passed, by kind of question
Setup Runs on From inspection findings / 678 General topics / 822
Gemini 3.1 Flash‑Lite · thinking off Google API From inspection findings / 678644 General topics / 822822
Gemma 4 E4B · thinking on Local Mac Mini From inspection findings / 678580 General topics / 822801
Gemma 4 E4B · thinking off Local Mac Mini From inspection findings / 678554 General topics / 822803
Across all 300 providers, 678 of each setup’s 1,500 question slots were built from a provider’s inspection findings, and the rest from general topics. Undecided: 1 findings pair for Flash‑Lite, 4 and 2 for Gemma 4. Download CSV

Grading: Claude Opus 5.5 vs. GPT‑6 Astra

Claude Opus 5.5 and GPT‑6 Astra, both at high reasoning, graded every question set on their own, without knowing which model wrote it. Gemini 3.8 Flash was used as a tie-breaker on question sets where Opus and Astra disagreed.

The two judges agreed on most grades, but GPT‑6 Astra was stricter. They gave the same pass or fail to 93.6% of the 4,480 question pairs (Cohen’s kappa 0.59, an agreement score that allows for chance, where 1 is perfect). Astra rejected 444 of the 4,480 question pairs the models wrote (9.9%), and Opus 303 (6.8%). In the 285 disagreements, Astra gave the harsher grade 213 times, vs. Opus’s 72.

Opus flagged more factual errors. It marked 133 question pairs as unsupported or contradicted by the evidence, vs. Astra’s 116.

Different LLMs gave different grading results. Graded by Opus alone, Gemini 3.1 Flash‑Lite passed 271 of its 300 sets. Graded by Astra alone, it passed 243.

Table 4. Claude Opus 5.5 vs. GPT‑6 Astra as judges
Judge Reasoning Pairs rejected / 4,480 Factual errors Median time Cost per review Total
Claude Opus 5.5 High Pairs rejected / 4,480303 (6.8%) Factual errors133 Median time16.46 s Cost per review$0.093 Total$83.41
GPT‑6 Astra High Pairs rejected / 4,480444 (9.9%) Factual errors116 Median time55.42 s Cost per review$0.123 Total$110.47
Both judges reviewed the same 896 valid sets, 4,480 question pairs. One review covers one set of five. Times come from each provider’s own clock and are not a controlled race: the judges ran at different times, and part of Astra’s work ran four jobs at once. Astra’s median leaves out 1 review that included a pause. Cost is API-equivalent at October 3, 2026 list prices; both judges ran on subscriptions. Download CSV

A third tie-breaking LLM sided with Opus on most quality disagreements. When Opus and Astra disagreed, Gemini 3.8 Flash reviewed the whole set on its own, without seeing their grades. On quality, it agreed with Opus 215 times out of 278 (77.3%), and with Astra 63 times (22.7%). On factual errors it was close: 50 (49.5%) vs. 45 (44.6%), out of 101.

Opus 5.5 took less time and cost less than GPT‑6 Astra. Its median review took 16.46 s, vs. 55.42 s for Astra (70% less time). At list prices, the 896 reviews would have cost $83.41 with Opus vs. $110.47 with Astra (32% more). The judges ran at different times and under different loads, so this is what we observed, not a controlled speed test.

Table 5. Question sets that passed, by judge
Setup Runs on Opus alone Astra alone Final
Gemini 3.1 Flash‑Lite · thinking off Google API Opus alone271 Astra alone243 Final270
Gemma 4 E4B · thinking on Local Mac Mini Opus alone213 Astra alone197 Final220
Gemma 4 E4B · thinking off Local Mac Mini Opus alone187 Astra alone177 Final196
Out of 300 sets per setup. Final is the count in Table 1: where Opus and Astra disagreed, Gemini 3.8 Flash’s vote decided. It is decided pair by pair, so Final can be higher than either judge alone, as it is for Gemma 4 with thinking on. Download CSV

Practical takeaways

Gemini 3.1 Flash‑Lite was a clear winner on both fronts - speed and quality. Both subjectively and by the data, Gemma 4 E4B is no match for this affordable Gemini model.

Opus 5.5 outperformed GPT‑6 Astra on both speed (a lot) and cost (moderately) for the task of judging other LLMs’ outputs.