01
The judge is compromised
An LLM cannot be your CI gate.
The research is blunt. On identical tasks, deterministic checks reach precision 1.00; LLM judges manage 0.50–0.67 — a coin flip. Change the sampling temperature, and the verdict changes with it. That’s not a gate. That’s a slot machine.
Worse: the misses are silent. LLM judges under-count the most dangerous bugs — failures dressed up as success — by 44 to 1.