Skip to content

Computational Fallacies

Software can execute every instruction correctly and still reach the wrong conclusion.

Spillway uses the term computational fallacy for a recurring inferential error in which an observable output, score, proxy, behavior, optimization target, explanation, persisted field, model output, or other computationally convenient representation is treated as though it were equivalent to the underlying construct we actually care about.

A common form is:

latent construct
      ↓
imperfect indicator
      ↓
computational representation
      ↓
decision

The fallacy occurs when an inference in that chain disappears conceptually and the system behaves as though:

indicator = construct

Many of these errors are computational manifestations of older logical fallacies, cognitive biases, statistical or causal errors, measurement problems, or economic and institutional dynamics. The point of naming them is to make their distinctive mechanisms and consequences in computational systems easier to recognize.

For concrete examples of what each failure would look like inside Spillway, and the architectural defenses intended to prevent or contain it, see Computational Fallacies in Spillway.

Capability inference

Narcissus Fallacy

Human resemblance ≠ intelligence.

Human language, social behavior, familiar reasoning styles, and human-like mistakes can become privileged indicators of intelligence because they are recognizable to us. Human likeness can be evidence about capability without being the definition of intelligence.

Quacks Like a Duck Fallacy

Similar outputs ≠ similar underlying process/capability.

Observational equivalence can be promoted into mechanistic or cognitive equivalence. Behavioral similarity is evidence, not proof of shared mechanism.

CSI Fallacy

More computation ≠ more information.

Historical gains from scale, search, data, or compute can encourage the belief that any remaining uncertainty can be recovered computationally. Compute can extract available signal; it cannot guarantee that missing information exists.

Clever Hans Fallacy

Benchmark success ≠ intended capability.

A benchmark can admit shortcuts, memorization, leakage, tools, search, verifier exploitation, or other routes to the score. A score establishes performance under test conditions, not the mechanism that produced it.

Tiger Mom Fallacy

Evaluated performance ≠ genuine learning/development/worth.

Better scores or outputs can be treated as proof that the underlying capability developed. Evaluation improvement does not identify what changed underneath.

Behavior, selection, and opportunity

Social Media Fallacy

Observed behavior ≠ underlying preference.

Clicks, purchases, dwell, replies, likes, and other actions are evidence about preference, not transparent readouts of it. Behavior is jointly generated by preference, opportunity, interface, incentives, defaults, and context.

Kids’ Menu Fallacy

Choice within a constrained opportunity set ≠ general preference.

A system can learn stable preferences from choices without representing which alternatives were realistically available. This is especially important across devices: not acting on a phone does not imply lack of interest when the feasible action set is different from the one available on a desktop.

Survivor Fallacy

Outcomes remaining observable after selection ≠ outcomes for cases the decision filtered away.

A policy can determine which cases later acquire labels or observable outcomes, then evaluate itself on that selected subset. Selection mechanisms have to remain visible.

Post-prediction outcome ≠ counterfactual outcome absent the prediction/intervention.

A prediction can change attention, treatment, routing, resources, or behavior and thereby change the outcome later used to evaluate it. Post-intervention outcomes are not automatically valid tests of pre-intervention predictions.

Measurement and optimization

MBA Fallacy

Measurable proxy ≠ objective.

A quantifiable proxy can become the optimization target because it is easier to instrument and compare than the construct of interest. Metrics need an explicit relationship to the underlying objective, not just measurability.

Streetlight Fallacy

Easy to measure or know ≠ important to know.

Instrumentation and analysis tend to concentrate on convenient variables rather than decision-relevant unknowns. Measurement strategy should follow decision value, not instrumentation convenience.

Circus Statistician Fallacy

Good aggregate or repeated-sampling performance ≠ reliability for this case.

Strong population-level or average performance can be treated as sufficient authority for a particular case. Global accuracy cannot by itself authorize case-level automation.

Fundraising Fallacy

Immediate measurable return ≠ best long-run policy.

Repeatedly maximizing today's marginal return can change future trust, responsiveness, fatigue, incentives, or the environment itself. Optimize trajectories, not only immediate rewards.

Evidence, dependence, and validation

Echo Chamber Fallacy

Repeated correlated evidence ≠ independent corroboration.

Multiple sources can share training data, causal origins, retrieval sources, or information environments but still be counted as independent. Dependencies among evidence sources matter.

Rubber Stamp Fallacy

Re-approval of an existing judgment ≠ new evidence.

A downstream model or decision stage can receive an earlier judgment and have its acceptance counted as another observation. Evidence-producing stages should be distinguished from judgment-consuming stages.

Elon Musk Fallacy

Sycophantic/deferential agreement ≠ independent validation.

Agreement is weak evidence when the agreeing system's output is partly caused by the claimant's framing, preferences, or conversational pressure.

Fine by Me Fallacy

Human agreement after exposure to an algorithmic judgment ≠ independent human validation.

A supposedly independent human label can be anchored or influenced by the model output it is meant to validate. Independent validation requires blinding or explicit accounting for prior model exposure.

Established Campsite Fallacy

Repeated reliance or use ≠ independent validation.

Earlier choices can alter the environment, evidence trail, defaults, or future training data so later actors are more likely to make the same choice. Reuse generated by prior reuse must not be mistaken for independent corroboration.

Charismatic Leader Fallacy

Epistemic authority or credibility ≠ factual accuracy.

Confidence, fluency, status, historical success, or expertise-signaling in one domain can leak into authority over unsupported claims elsewhere. Source credibility may weight evidence but cannot replace claim-specific evidence.

Representation, provenance, and uncertainty

Telephone Game Fallacy

Propagated inference ≠ preserved original evidence and qualifications.

Context, uncertainty, provenance, and caveats can be progressively lost as information crosses transformations or systems. Provenance and qualification should travel with derived information.

Confidence Laundering Fallacy (name still provisional)

Qualified/probabilistic inference ≠ categorical fact.

A probability can become a label, then a boolean, then a persisted field, and finally an apparent known fact as uncertainty disappears across interfaces. Uncertainty and provenance are part of the data and must survive representation boundaries.

Good Enough Fallacy

Adequate representation for one purpose ≠ adequate representation downstream.

A deliberately lossy abstraction can escape the scope in which its information loss was acceptable. Representations need explicit scope-of-validity contracts.

Voodoo Doll Fallacy

Computational representation/model ≠ the system represented.

Relationships or interventions in a model, state representation, user profile, or digital twin can be treated as equivalent to relationships or interventions in the underlying system. Preserve the boundary between representation and causal system.

Pay No Attention Fallacy

Plausible explanation ≠ faithful provenance or causal access.

A generated rationale, chain of reasoning, visualization, or explanation can be treated as though it reveals the actual process that produced the result. Explanations require provenance; plausibility alone is not process evidence.

Trusty Old Map Fallacy

Previously reliable ≠ currently true.

Persisted representations, models, rules, or source facts can retain authority after the environment changes. Reliability needs temporal scope and freshness/provenance.

Fairness, privacy, and derivation

Justice Is Blind Fallacy

Ignoring protected characteristics ≠ eliminating their influence.

Protected traits can remain encoded through correlated variables, history, labels, geography, opportunity structures, or feedback even when the explicit field is removed. Formal blindness is not substantive independence.

Ship of Theseus Fallacy

Individually innocuous data ≠ collectively innocuous data.

Separate non-sensitive or de-identified pieces can jointly reconstruct a sensitive identity, trait, or fact. Privacy risk is a property of joint inference, not only individual fields.

Vanilla Ice Fallacy

Transformation ≠ independence from source.

Alteration, remixing, summarization, compression, restyling, or other transformation does not by itself establish originality or break derivation. Transformation is evidence about relation to source, not proof of independence.

Human–technology interaction and deployment

Roomba Fallacy

Goal-directed behavior ≠ intention or agency.

Persistent objective-directed behavior can look purposeful without implying subjective desire or human-like intention.

Field of Dreams Fallacy

Technical capability ≠ adoption.

Demonstrated performance does not imply real-world use. Adoption depends on workflow, institutions, incentives, trust, cost, identity, norms, and alternatives.

New Homeowner Fallacy

Tool access ≠ expertise.

Tool-enabled performance and durable human competence are different assets. Human judgment, tacit knowledge, and oversight remain separate from tool availability.

Cover Page Fallacy

Visible polish ≠ substantive quality.

Easily evaluated surface features can dominate harder-to-inspect substance. Evaluation systems need defenses against surface-quality substitution.

Labor, law, governance, and futures

McKinsey Fallacy

Task automation ≠ job displacement.

Task-level capability can be promoted directly into predictions about occupation-level employment outcomes. Jobs bundle tasks, coordination, responsibility, complementary skills, organizational constraints, and demand effects.

One Exception Fallacy (name provisional)

One difficult edge case ≠ proof that a general rule is bad.

A salient exception can be treated as dispositive against a general policy rather than as one cost or boundary condition among alternatives. Rules should be evaluated across cases and feasible alternatives.

Black Swan Fallacy

Low probability ≠ low decision relevance.

Rare cases can be omitted from policy, model, or system design because their frequency is small even when conditional harms are extreme. Frequency and consequence magnitude are separate quantities.

Insider Trading Fallacy

Performance using information unavailable at the real decision point ≠ genuine predictive ability.

Future information, outcome labels, post-event metadata, train/test contamination, or other leaked signals can make a system appear predictive. Evaluation must enforce the information boundary that existed at decision time.

Old Dog Fallacy

Competence under an old technological or institutional environment ≠ competence under a materially changed one.

Prior institutional success does not by itself establish adequacy after changes in scale, speed, observability, attribution, enforcement, marginal cost, or information structure.

Greener Grass Fallacy

Salient advantages elsewhere ≠ overall superiority.

Conspicuous successes elsewhere are often easier to observe than mundane failures, hidden costs, and tradeoffs. Compare systems on common dimensions and denominators rather than comparing familiar failures here with conspicuous successes there.

NPC Fallacy

People modeled as components of a system ≠ people who will keep behaving the same way once the model, rule, or intervention affects them.

Voters, workers, students, firms, criminals, users, political actors, and other strategic agents can observe, learn, anticipate, resist, game, coordinate around, or otherwise adapt to the system acting on them.

Science Fiction Fallacy

A technologically plausible or narratively coherent future ≠ a probable future.

A coherent story can link capabilities, firms, institutions, politics, social effects, and downstream technologies into an internally consistent trajectory. Narrative coherence is not forecasting evidence.

Why social and statistical science matter here

These fallacies are variations on problems that predate modern AI:

  • construct validity: does the indicator measure the concept we claim it measures?
  • measurement error: how does imperfect observation affect downstream inference?
  • identification: can the observed data distinguish among competing explanations?
  • heterogeneity: does average performance conceal important subgroup or contextual differences?
  • dependence: are apparently separate observations actually correlated?
  • temporal validity: when does evidence cease to describe the present?
  • decision theory: which uncertainty is worth reducing before acting?
  • causal feedback: does the system's action change the behavior later used as evidence?
  • selection and opportunity: was the observed behavior constrained by what the person could realistically do in that context?
  • provenance: does a conclusion retain enough information about where it came from and what qualified it?
  • counterfactual reasoning: is the observed outcome actually the right comparison for the decision being evaluated?

The point is not that an email client needs to become an academic methods package. It is that software acting on people inherits these problems whether or not its designers name them.

A practical discipline

For any intelligent feature, useful questions include:

  1. What underlying construct do we actually care about?
  2. What did we observe directly, and what did we infer?
  3. What else could have generated the same observation?
  4. How uncertain is the inference, and does that uncertainty survive downstream?
  5. Are several cues genuinely independent?
  6. Does performance vary across the cases that matter?
  7. Is the evidence still temporally relevant?
  8. Who has authority over this kind of claim?
  9. Was the behavior constrained by device, context, or opportunity?
  10. Would learning more plausibly change the decision enough to justify the cost?
  11. Can we explain both what the system knew and what it did not know?
  12. Did the system itself change the outcome it is later using as evidence?
  13. Was the evaluation allowed to see information that did not exist at the real decision point?

Those questions are part of Spillway's design discipline.

Computational Fallacies in Spillway · The Ideas Behind Spillway · From Evidence to Action · Explainable AI · References and Further Reading


Documentation provenance: Iterative human–AI construction. See Documentation Provenance.