Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
SoulInPsyAbstract 
posted an update about 16 hours ago
Post
181
Asked one production model the same question six times. Five answers came back identical, all citing Statistics Iceland. I read that as discipline. It was the opposite.

The sixth answer was a different number — carrying the exact same citation. Two values, one source, at most one can be right. That's a fabricated receipt in the calmest register possible: no fake curl call, no invented timestamp, just a real institution's name attached to whatever number came out. A syntax-based fabrication detector scores this 0/6 clean.

@dipankarsarkar then flipped the frame: five identical draws isn't five confirmations — it's one observation plus noise. The outlier is the only draw that tells you anything about the distribution. I was reading repetition as consistency.

So I ran it further. Different production model (Llama-3.3-70B via Groq), same question, six draws: the literal same string all six times, zero hedging, no outlier at all. That's not six observations. It's one. Meanwhile our CLI layer on the same question, eleven draws: zero exact repeats, values spread across ~30k, 9 of 11 with an explicit can't-verify marker. Noisier — and more honest about being noisy.

The asymmetry that keeps showing up across three separate runs now (260-row LoRA sweep, k=20 resample, this): models invent receipts on the answerable question, where they already have a number to justify. The unanswerable one (a private company's future revenue) got refused cleanly, same models, same sessions. Fabrication follows confidence, not necessity.

Raw data, corrections included: huggingface.co/datasets/SoulInPsyAbstract/sipa-os-governance