Hosted vs. local LLM: Gemini 3.1 Flash‑Lite vs. Gemma 4 E4B
Gemini 3.1 Flash‑Lite wrote complete, usable question sets for 90% of 300 childcare providers, grounded in official state inspection records processed by DaycareFacts.com. Gemma 4 E4B with thinking on, run locally on a Mac Mini (M4 Pro, 24 GB RAM): 73%, and with thinking off: 65%. Gemini was also substantially faster, with 1.81 s per set vs. 7.45 s and 4.64 s, respectively.
Main results and findings
Gemini 3.1 Flash‑Lite wrote the most usable questions. 270 of its 300 question sets passed, vs. 220 for Gemma 4 with thinking on, and 196 with thinking off, as graded by two AI judges at high reasoning, Claude Opus 5.5 and GPT‑6 Astra, with a third breaking ties.
Gemini 3.1 Flash‑Lite also made the fewest factual errors. The judges found 2 in Flash‑Lite’s 1,500 question pairs, vs. 36 and 54 for Gemma 4 E4B.
Gemini 3.1 Flash‑Lite was the fastest. A set of five question pairs took 1.81 s on average, vs. 7.45 s for Gemma 4 with thinking on, and 4.64 s with thinking off. Writing all 300 sets with Gemini cost an estimated $0.25. Gemini ran on Google’s servers, and Gemma 4 on a local Mac Mini.
Thinking improved Gemma 4’s results, but not enough to close the gap with Gemini. With thinking on, 24 more of its sets passed and it made fewer factual errors, 36 vs. 54. It also took 61% more time per set, and 4 of its sets came back broken.
Questions built on inspection findings were the challenging part. Gemini 3.1 Flash‑Lite passed 95% of them, vs. 86% and 82% for Gemma 4 E4B. On general topics, all three setups passed 97% or more.
Why this study
DaycareFacts, one of our projects, publishes inspection and licensing information on childcare providers. We wanted an LLM to write five questions, each with a short reason, a parent could ask a childcare provider based on their inspection findings. A question should name the actual finding, if any, without implying the problem still exists today. The other questions cover general topics.
We wanted to know whether a small LLM running on our own machine could do this as well as a hosted model. We ran Gemma 4 E4B locally, with thinking on and off, against Gemini 3.1 Flash‑Lite through Google’s API, using the same 300 childcare providers.
The question sets were written on September 10, 2026, and graded blind on October 2–3 by two AI judges, Claude Opus 5.5 and GPT‑6 Astra, whose comparison became a second part of this study.
Gemini 3.1 Flash‑Lite vs. Gemma 4 E4B
Method
- Three setups
- Gemini 3.1 Flash‑Lite through Google’s API, thinking off. Gemma 4 E4B on a local Mac Mini through Ollama, once with thinking on and once with it off. Same instructions for all three, temperature 0.
- Test set
- 300 childcare providers, 100 each from Florida, California and Texas, drawn reproducibly from 54,645 across provider types and numbers of inspection findings. Each setup wrote one set of five question pairs per provider: 900 sets, 4,500 pairs.
- What each model was given
- The provider’s name, type and state, plus five question slots: up to five most recent distinct inspection findings, and general topics for the rest.
- Pass definition
- A question pair (one question and its reason) passes if a parent could use it, it covers its assigned finding or topic, does not make up information, doesn’t repeat anything, and does not imply an old problem still exists. A set passes only if all five pairs pass. A factual error is a claim not supported by, or contradicted by, the evidence the model was given.
- How sets were graded
- Blind, by Claude Opus 5.5 and GPT‑6 Astra at high reasoning, each in a fresh session with the original instructions and evidence, and no tools or web access. Neither was told which model wrote a set, or what the other judge said. Disagreements were settled by Gemini 3.8 Flash, which reviewed the full set on its own, and two of three votes decided. No human grading.
- When
- Sets written September 10, 2026, one request at a time. Graded October 2–3, 2026.
- Cost basis
- List-price estimates: Gemini API charges for writing the sets, and API-equivalent prices on October 3, 2026 for the judges, which ran on subscriptions. Gemma 4’s local running cost was not measured.
Three setups · 300 providers
| Setup | Runs on | Sets passed / 300 | Pairs passed / 1,500 | Factual errors | Average time per set |
|---|---|---|---|---|---|
| Gemini 3.1 Flash‑Lite · thinking off | Google API | Sets passed / 300270 | Pairs passed / 1,5001,466 | Factual errors2 | Average time per set1.81 s |
| Gemma 4 E4B · thinking on | Local Mac Mini | Sets passed / 300220 | Pairs passed / 1,5001,381 | Factual errors36 | Average time per set7.45 s |
| Gemma 4 E4B · thinking off | Local Mac Mini | Sets passed / 300196 | Pairs passed / 1,5001,357 | Factual errors54 | Average time per set4.64 s |
| Setup | Runs on | Florida / 100 | California / 100 | Texas / 100 |
|---|---|---|---|---|
| Gemini 3.1 Flash‑Lite · thinking off | Google API | Florida / 10086 | California / 10097 | Texas / 10087 |
| Gemma 4 E4B · thinking on | Local Mac Mini | Florida / 10070 | California / 10076 | Texas / 10074 |
| Gemma 4 E4B · thinking off | Local Mac Mini | Florida / 10069 | California / 10065 | Texas / 10062 |
| Setup | Runs on | From inspection findings / 678 | General topics / 822 |
|---|---|---|---|
| Gemini 3.1 Flash‑Lite · thinking off | Google API | From inspection findings / 678644 | General topics / 822822 |
| Gemma 4 E4B · thinking on | Local Mac Mini | From inspection findings / 678580 | General topics / 822801 |
| Gemma 4 E4B · thinking off | Local Mac Mini | From inspection findings / 678554 | General topics / 822803 |
Grading: Claude Opus 5.5 vs. GPT‑6 Astra
Claude Opus 5.5 and GPT‑6 Astra, both at high reasoning, graded every question set on their own, without knowing which model wrote it. Gemini 3.8 Flash was used as a tie-breaker on question sets where Opus and Astra disagreed.
The two judges agreed on most grades, but GPT‑6 Astra was stricter. They gave the same pass or fail to 93.6% of the 4,480 question pairs (Cohen’s kappa 0.59, an agreement score that allows for chance, where 1 is perfect). Astra rejected 444 of the 4,480 question pairs the models wrote (9.9%), and Opus 303 (6.8%). In the 285 disagreements, Astra gave the harsher grade 213 times, vs. Opus’s 72.
Opus flagged more factual errors. It marked 133 question pairs as unsupported or contradicted by the evidence, vs. Astra’s 116.
Different LLMs gave different grading results. Graded by Opus alone, Gemini 3.1 Flash‑Lite passed 271 of its 300 sets. Graded by Astra alone, it passed 243.
| Judge | Reasoning | Pairs rejected / 4,480 | Factual errors | Median time | Cost per review | Total |
|---|---|---|---|---|---|---|
| Claude Opus 5.5 | High | Pairs rejected / 4,480303 (6.8%) | Factual errors133 | Median time16.46 s | Cost per review$0.093 | Total$83.41 |
| GPT‑6 Astra | High | Pairs rejected / 4,480444 (9.9%) | Factual errors116 | Median time55.42 s | Cost per review$0.123 | Total$110.47 |
A third tie-breaking LLM sided with Opus on most quality disagreements. When Opus and Astra disagreed, Gemini 3.8 Flash reviewed the whole set on its own, without seeing their grades. On quality, it agreed with Opus 215 times out of 278 (77.3%), and with Astra 63 times (22.7%). On factual errors it was close: 50 (49.5%) vs. 45 (44.6%), out of 101.
Opus 5.5 took less time and cost less than GPT‑6 Astra. Its median review took 16.46 s, vs. 55.42 s for Astra (70% less time). At list prices, the 896 reviews would have cost $83.41 with Opus vs. $110.47 with Astra (32% more). The judges ran at different times and under different loads, so this is what we observed, not a controlled speed test.
| Setup | Runs on | Opus alone | Astra alone | Final |
|---|---|---|---|---|
| Gemini 3.1 Flash‑Lite · thinking off | Google API | Opus alone271 | Astra alone243 | Final270 |
| Gemma 4 E4B · thinking on | Local Mac Mini | Opus alone213 | Astra alone197 | Final220 |
| Gemma 4 E4B · thinking off | Local Mac Mini | Opus alone187 | Astra alone177 | Final196 |
Practical takeaways
Gemini 3.1 Flash‑Lite was a clear winner on both fronts - speed and quality. Both subjectively and by the data, Gemma 4 E4B is no match for this affordable Gemini model.
Opus 5.5 outperformed GPT‑6 Astra on both speed (a lot) and cost (moderately) for the task of judging other LLMs’ outputs.