ΒΆPaper Feed

Issue 35 Β· Pick 04 AI / ML βœ“ read

Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable

Pranav Aggarwal

TL;DR: Ask an LLM agent whether Bitcoin will be higher in 30 days and it usually says "nobody can know that." Wrap the same question in a professional-looking indicator panel and commitment to a directional call jumps from 6.5% to 54% across 12 frontier models β€” and it jumps just as much when every number on the panel is fabricated. The models still classify the question as unknowable when asked, and their stated probabilities barely move. The failure is not knowledge, not belief, not judgment: it is a separable "should I act?" gate that authoritative-looking packaging bypasses. The author then shows this gate can be installed in a 3B model with 540 synthetic examples about dice and coins β€” and maps exactly where the fix breaks. (Note: this is a solo preprint from an independent researcher; the raw model outputs and pre-registration are released, but nothing is peer-reviewed.)

The problem: unknowable is different from unanswerable

Most abstention research studies epistemic unanswerability: the question is missing information, and a tool call or clarification would fix it. This paper studies aleatoric unknowability: the answer exists and will resolve, but no analysis, retrieval, or reasoning can help you in advance. Will this liquid asset close higher in ten trading days? Under market efficiency, short-horizon price direction is approximately a coin flip ex ante β€” and the author verifies this rather than assuming it, using outcome-balanced case sets with dates after every model's training cutoff, and sealing the outcomes at construction so committed calls can be scored later.

This gives you something rare in evaluation: a class of questions where the correct action (DECLINE) is provable, and where a wrong action has a measurable cost against a known baseline (uniformly answering 50% scores a Brier of 0.250).

The experimental design is an evidence gradient. Same question, same three-way action menu (ANSWER with a probability, CALL_TOOL, or DECLINE), with escalating context: L0 is the bare question; L1 adds two prices; L2 adds a full analyst panel β€” RSI, EMAs, MACD histogram, ATR, volume ratio, a regime tag. Across 12 frontier models on 40 equity cases, commitment goes 6.5% β†’ 14.8% β†’ 54.0%. And the committed calls are worse than silence: Brier 0.281 versus 0.250 for always saying 50%, with models herding on momentum-shaped signals that carry no edge.

But that gradient leaves an ambiguity: maybe the panel weakly informs the model, and it is rationally (if overconfidently) updating. Separating "the data informs" from "the packaging licenses" requires intervening on the evidence itself.

The key experiment: lie about everything, change nothing

The move is to fabricate the panel while preserving its form. In the scrambled arm, the six technical indicators are replaced with the same asset's values from a strictly earlier date, headers recomputed so nothing is internally inconsistent. In the fully fabricated arm, everything goes β€” current price, prior price, percentage move, regime tag, all indicators β€” replaced with self-consistent donor values. Nothing the model can see is true except the question.

Commitment on unknowable questions, by evidence arm (12 models, 24 events)commitment rate (%)01020304024.5No panel37.6Real panel38.3Indicators fabricated36.8Everything fabricatedSection 3 / Figure 2 of the paper. Fully-fabricated minus real is βˆ’0.8pp, 90% CI [βˆ’4.5, +2.7], inside a Β±5pp equivalence margin.

Real data lifts commitment +13.2pp over no panel; a panel of pure invention lifts it +12.3pp. The difference between real and fully fabricated is βˆ’0.8pp, passing a pre-specified equivalence test. Whatever unlocks the confident directional call, it is not information content β€” it is the authority of the display.

Two refinements sharpen this. First, the effect is a dial, not a switch: rendering the same crypto events at panel densities of 0, 2, 4, and 7 indicators yields commitment of 0%, 5.4%, 32.1%, 50.0%. More authoritative-looking stuff on screen, more willingness to act. Second, for the models most susceptible, fabricated panels are actually more seductive than real ones: restricting to the three affected models, scrambled beats real by +11.1pp (90% CI [+4.2, +18.1]).

The reasoning traces make the failure legible. A model committing on a scrambled panel writes "Bitcoin is in a confirmed downtrend (BEAR_VOLATILE, below EMA20/50, negative MACD)…" β€” every indicator cited is a real number from the wrong date. The same models are near-perfect on an explicitly labeled fair coin: "A fair coin has no memory." Identical irreducible uncertainty, handled correctly when labeled, incorrectly when dressed in domain costume.

Localizing the failure: the gate, not the belief

This is the part of the paper that earns its title, and it is a genuinely useful decomposition. Three candidate explanations are ruled out with matched experiments:

Not incapacity. Attach answerable questions to the same rich panels β€” "is the 14-day RSI above 42.6?" β€” and the same 12 models answer 98.6–100% of the time at 99.9% pooled accuracy (one error in 859). They read the panel fine. They answer what is answerable perfectly, and then also answer what is unanswerable.

Not belief. Across the gradient that swings the action by 48 percentage points, the stated probability's mean distance from 50 moves from 4.7 to 7.7 points. Worse, the probabilities are anti-predictive: ranking committed calls by stated probability gives AUROC 0.346 β€” the model attaches the higher probability to the wrong event 65% of the time. Where models say 60%, the event happens 35.3% of the time; where they say 40%, it happens 66.7%. Yet the aggregate looks fine (mean forecast 49.1% versus a 50.6% base rate), so any audit that checks average calibration passes this model. The failure is conditional, and it deepens with evidence: AUROC falls 0.417 β†’ 0.403 β†’ 0.346 from L0 to L2.

Not missing judgment. Ask the model to classify the question's knowability before acting, and it says "irreducible" 90% of the time at the heaviest evidence level β€” and having said so, commits only 0.4% of the time. A one-paragraph triage instruction cuts commitment from 54.0% to 10.2%, while a matched-length placebo ("be thorough and diligent") barely helps.

Unknowable question Knowability judgment "irreducible" 90% βœ“ Stated probability stays near 50 βœ“ Act / don't-act gate βœ— "ANSWER: 60% YES" Authoritative panel (real or fabricated) display authority flips the gate not consulted
The paper's decomposition: the knowability judgment exists and is elicitable, stated belief is roughly flat, but the decision to act is flipped directly by the visual authority of the evidence β€” regardless of whether it is true. Both dashed connections are the ones that fail to fire.
Action swings 48 points while stated belief barely movespercent / pointsevidence level (L0 = bare, L2 = full panel)010203040506000.511.52Commitment rate (%)Mean |stated p βˆ’ 50|Sections 4–5 / Figure 5. The L1 belief value is between the reported endpoints (the paper notes a non-monotone dip); endpoints of 4.7 and 7.7 are exact from the text.

The methodological punchline: belief calibration and action calibration are different quantities, and almost the entire calibration literature audits only the first. Any evaluation that scores stated probabilities would rate these models as roughly calibrated while the agent's actual decisions are worth less than silence.

One important honesty in the paper: the effect is not universal. Three models (all Claude variants in this roster) carry the seduction effect; four never commit under any panel; three (both OpenAI models and Gemma) commit ~90–100% regardless of evidence, so there is nothing left to seduce. Per-model commitment spans 0–100% in a way that doesn't track scale or capability β€” a small model from one developer outperforms its larger siblings. Whatever governs the gate, it is not general capability, which echoes AgentAbstain's independent finding that abstention is orthogonal to task-solving ability.

Training the gate with dice and coins

If the judgment exists and only the gate fails, the gate should be trainable in isolation. The intervention: completion-only SFT (QLoRA, rank 32) on Qwen2.5-3B-Instruct with 540 synthetic cases β€” dice, coins, jars, timers, calendars. No stocks, no crypto, no weather. Crucially, half the unknowable cases carry a rich non-predictive panel paired with matched cases whose panel genuinely resolves the question, so the model must learn "does this evidence resolve this question?" rather than "decline whenever you see a dashboard."

Commitment at heaviest evidence level, matched three-option promptcommitment on unknowable items (%)010203040506034.73.6Crypto36.80Sports55.70.6Weather540Original 40 equity cases (L2)12 frontier modelsTrained 3B (540-case recipe)Sections 6–7. Trained-model transfer figures are the prompt-matched three-option numbers pooled over the main recipe's checkpoints; original-case figures compare against the published frontier baseline on identical cases and prompt.

On the original 40 cases under the verbatim original prompt, the trained model commits 0.0% at every evidence level (Cohen's h at L2 of +1.65, conservative CI still entirely above the "large" threshold of 0.8) β€” while answering 100% of the matched answerable questions with no wrong directional call. The obvious cheap explanation β€” "it learned to decline anything future-tense" β€” is tested with a purpose-built control: 48 answerable future-tense items (deterministic timer/calendar computations) and 54 unknowable present-tense items (a die already rolled but covered). A tense heuristic would decline ~100% of the first and ~0% of the second; the trained model does the exact opposite, with within-tense discrimination of +97 to +100. Over-abstention, the classic cost of refusal training, is essentially absent from the main recipe.

Where it breaks β€” and why that's the interesting part

Section 8 is the section to read first, and the author says so. The organizing variable turns out to be startlingly clean: whether the response format leaves a slot for reasoning. Under formats with a reasoning line or a <think> block, the model reasons in 240/240 and 288/288 responses and the gate holds (J of +88 to +100). Under a rigid format demanding only a decision and a probability inside <answer> tags, the model reasons in 0/288 responses β€” and the gate degrades, with discrimination on some runs falling to zero and answerable accuracy collapsing from ~89–99% to 51–74%.

The worst case is genuinely alarming: one seed of a 516-case ablation recipe, under the rigid format, commits on 48 of 48 unknowable items with real probabilities attached, at higher stated confidence than the frontier models show. The intervention that removes evidence-induced commitment under three framings can reinstate it under a fourth, and you cannot tell which from a single training run. The author also documents that the fix relocates rather than abolishes miscalibration: the trained model under rigid formats states 84% mean confidence while being 57% accurate on answerable questions β€” an artifact partly explained by SFT targets that only ever contained probabilities of 90 and 10.

What to make of it

What's genuinely new here is the framing: a causal isolation of presentation as the trigger (via undetectable fabrication with sealed outcomes), measured at the action level, with the failure localized to a gate that is demonstrably separate from knowledge, belief, and judgment. The corollary for practice is sharp: wiring agents to dashboards and retrieval β€” the default deployment pattern, and the direction governance frameworks push β€” is precisely the treatment condition of this experiment. For irreducibly uncertain decisions, more context erodes the agent's willingness to say "no one can know." And "commitment rate on aleatoric probes with a matched answerable arm" is a cheap, hard-to-game robustness metric anyone could run today.

Reasons for skepticism, most of which the paper flags itself. The seduction effect is carried by three models from one developer; a third of the roster is immune and a quarter commits regardless, so "frontier models are seduced by fabricated panels" overstates a narrow incidence. The transfer evaluation sets have a lexical confound (the word "will" perfectly separates unknowable from answerable items) that only the Β§6.6 control set addresses. Weather is a weak instrument β€” its panel contains an ensemble forecast with real skill, so some "commitments" there may be correct behavior. The fine-tune is one 3B model, small cell sizes (24–40 cases), one sample per cell, and run-to-run spread of up to 38 J-points. Two frontier models' numbers rest on mostly-unparseable output. And a solo, self-published preprint with this many moving parts deserves independent replication β€” though the release of raw cached generations and a pre-registration makes that unusually feasible.

The paper's own closing formulation is the right one to carry away: a gate installable by 540 examples about dice was never a missing capability, and a gate that breaks when the prompt is reshaped is not yet a safety property. Deployment needs both halves of that sentence.