What Does “Proven” Actually Mean?
Proof, indications, probability, statistical significance, causation, testimony, and degrees of certainty
We use the word “proven” very broadly.
Science proved it.
The court proved it.
The photograph proves it.
Statistics prove it.
The document proves it.
The witness proved it.
But these statements do not mean the same thing. In mathematics, a proof can derive a proposition from axioms and rules of inference. In empirical science, we almost always work with:
- measurements;
- uncertainty;
- probability;
- models;
- incomplete data.
In law, “proven” is tied to a particular standard of proof. In history, we assess:
- documents;
- testimony;
- provenance;
- consistency among multiple sources.
So the first question should be:
Proven in what sense?
Evidence is not one single universal thing
Philosophy of science treats evidence as a relationship between:
evidence
and
hypothesis.
The Stanford Encyclopedia of Philosophy, in its discussion of confirmation, emphasizes that true data can increase the credibility of a hypothesis without necessarily proving it with absolute logical certainty.[1] In the empirical world, several hypotheses can often be logically compatible with the same data. So evidence frequently does not mean:
“the alternative is logically impossible.”
It means:
“this information increases or decreases support for a particular explanation.”
Mathematical proof and empirical proof are not the same thing. In mathematics, we can derive a proposition from:
- definitions;
- axioms;
- rules of inference.
If the formal proof is correct, the result follows from the starting assumptions. In empirical science, however, we must also ask:
- Is the measurement correct?
- Is the sample representative?
- Is the model appropriate?
- Is there an alternative explanation?
- Does the result replicate?
That is why empirical science more often speaks in terms of:
- support;
- probability;
- certainty of evidence;
- consistency of data with a model;
rather than absolute proof in the mathematical sense.
A fact is not the same as an explanation of the fact
Imagine a security camera. The recording shows:
person A entered the building at 20:14.
That may strongly support the claim:
A was in the building.
It does not by itself prove:
why A entered.
Nor:
what A did there.
Nor:
whether A later caused a particular event.
This is one of the most important distinctions:
evidence has a limited evidentiary scope.
The question is not only:
“Do I have evidence?”
It is:
“Which specific claim does this evidence actually support?”
An indication is not the same as more direct evidence. In everyday language, we often distinguish between:
indication. a circumstance that increases the reason to suspect or consider a hypothesis;
more direct evidence
information that more directly connects an actor to an act or a claim. Example:
the person had a motive.
That is an indication.
their verified digital signature appears on the order.
That is a different kind of evidentiary connection. Motive matters. But by itself it does not prove the act.
Multiple indications can together become very strong evidence
There does not have to be one “smoking gun.” A strong case can be built from:
- a timeline;
- communications;
- financial flows;
- physical evidence;
- several independent witnesses;
- technical data.
Each element by itself may permit several explanations. Together, they can create powerful convergence of evidence. That is why:
circumstantial evidence
is not automatically “weak evidence.” What matters is the whole evidentiary structure.
Repetition of the same source is not convergence
If ten articles summarize one anonymous statement, we have:
ten publications.
Not necessarily:
ten independent pieces of evidence.
The same applies in science. Ten papers may rely on:
- the same dataset;
- the same cohort;
- the same measurement device;
- the same underlying assumption.
So THY-REALITY should always ask:
How many genuinely independent paths lead to the same conclusion?
Statistical significance is not proof of importance
The p-value is one of the most misunderstood concepts in statistics. The American Statistical Association issued a dedicated statement in 2016 precisely because misinterpretations were so widespread.[2] A p-value is not:
- the probability that the hypothesis is true;
- the probability that the result is “just chance”;
- the size of the effect;
- the practical importance of the result.
The ASA specifically warns:
scientific and business decisions should not be based only on whether a p-value crosses a particular threshold.[2]
What does a p-value roughly tell us?
Simplified: if we assume a particular statistical model and the null hypothesis, the p-value tells us how unusual the observed result—or a more extreme one—would be under those assumptions. That is very different from:
“There is only a 3% chance that the result is wrong.”
That second statement is generally incorrect. A p-value concerns:
the data under an assumed hypothesis,
not directly:
the probability of the hypothesis given the data.
p < 0.05 is not a magical boundary of truth. If one study finds:
p = 0.049
and another:
p = 0.051,
it is not rational to declare the first proven and the second worthless. The two results are statistically very similar. The 0.05 threshold is a convention. It is not a law of nature. The ASA therefore warns against bright-line thinking in which one decimal threshold determines whether a claim is “true.”[2]
A large sample can make a tiny effect statistically significant
With a very large sample, we may detect a very small difference. Example: a medicine reduces a symptom by:
0.2 points on a 100-point scale.
That can be statistically significant. But the question remains:
Is the effect practically or clinically meaningful?
Statistical significance and practical importance are different things.
A small sample can miss an important effect
Conversely: a study may observe a fairly large effect but have:
- few participants;
- high variability.
The result may therefore fail to cross a traditional significance threshold. That does not prove:
there is no effect.
It may mean:
the study is too imprecise to estimate it reliably.
That is why we need uncertainty intervals.
A confidence interval is more informative than a point estimate alone. Cochrane emphasizes that an effect estimate should be presented together with a confidence interval, because the interval conveys statistical uncertainty around the estimate.[3] A narrow interval means:
greater precision.
A wide interval means:
greater uncertainty.
If the estimate says:
20% reduction
but the interval allows:
anything from a large benefit to almost no effect,
we should not communicate only the number:
20%.
A confidence interval does not capture every kind of uncertainty
Even a narrow interval does not necessarily mean a study is highly reliable. There may still be:
- bias;
- an unrepresentative sample;
- a wrong model;
- publication bias;
- a measurement problem.
Cochrane therefore emphasizes that statistical imprecision is only one domain of overall certainty in the evidence.[3] This is crucial:
precisely measured bias is still bias.
Association ≠ causation. If two variables occur together, we have:
an association.
If they change together in a linear way, we may talk about:
correlation.
That does not yet imply:
one causes the other.
Modern methodological literature clearly distinguishes association, correlation, and causation.[4]
Why is correlation not enough? Three classic possibilities:
A causes B
B causes A
C causes both
Example: ice-cream sales may be associated with more drownings. Ice cream does not cause drowning. Summer increases:
- ice-cream sales;
- swimming.
That is confounding.
Even a strong correlation does not prove causation
A correlation can be very high and still be:
- indirect;
- caused by a third factor;
- the result of selection.
So when someone makes a causal claim, we should ask:
What is the causal identification strategy?
How does the study distinguish:
“X occurs together with Y”
from
“X changes the probability of Y”?
A randomized experiment is powerful because it tries to break confounding
In a randomized trial, people are assigned randomly to groups. If randomization works, known and unknown factors are, on average, better balanced. The difference between groups is therefore more likely to result from the intervention.
The U.S. National Library of Medicine describes randomized controlled trials as one of the designs especially powerful for estimating causal effects.[5] But even an RCT is not perfect.
Randomized trials have limits. They may have:
- too small a sample;
- poor adherence;
- substantial participant loss;
- inadequate measurement;
- short follow-up;
- an unrepresentative population.
And randomization is not always:
- ethical;
- practical;
- possible.
We cannot randomly assign people so that:
half smoke for 30 years and half do not.
That is why causal inference needs other approaches too.
Bradford Hill did not create nine automatic “tests” for causation
In 1965, Bradford Hill described several considerations that can help us judge whether an observed association is likely to be causal. They are often called:
“Hill criteria.”
But historical review warns that Hill did not intend them as rigid criteria.[6] His considerations include:
- temporality;
- consistency;
- strength;
- biological gradient;
- plausibility.
The most necessary is:
temporality.
The cause must occur before the effect.
Causal inference has no single universal philosophy. Epidemiology uses several ways of reasoning about causation:
- counterfactual;
- probabilistic;
- sufficient-component;
- other models.
Systematic philosophical and methodological reviews emphasize that there is no single definition of causation that completely resolves every question.[7] So the phrase:
“causality has been proven”
should always be tied to:
- type of evidence;
- research design;
- alternative explanations.
Bayes: evidence should change the degree of confidence
A Bayesian view is especially useful for THY-REALITY. The basic idea is:
before new evidence, we have some initial degree of confidence in a hypothesis.
When new information appears, we ask:
how likely would this information be if the hypothesis were true, and how likely would it be under a competing hypothesis?
We then update our confidence. The Stanford Encyclopedia describes this as moving from prior to posterior credence, with evidence confirming a hypothesis when the hypothesis becomes more credible after taking the evidence into account.[8]
Evidence can support a hypothesis without making it certain
Imagine a disease that affects 1 person in 10,000. The test is very good. A positive result greatly increases the probability of the disease. But because the disease is very rare, a positive result does not necessarily mean:
almost 100% probability.
This is the base-rate problem. Evidence must be interpreted together with prior probability.
“How likely is this evidence if the suspect is guilty?” is not the same as “How likely is the suspect to be guilty given this evidence?”
This is a classic logical reversal. A forensic result may say:
this sample would be much more likely if it came from person A than if it came from an unrelated person.
That is not the same as:
there is a 99.99% probability that A is the perpetrator.
For the latter, we also need:
- other evidence;
- context;
- base rates.
NIST emphasizes that likelihood ratios and other statistical expressions are tools for evaluating evidentiary weight, with their own limitations and uncertainties.[9]
Forensic evidence is not error-free merely because it is “scientific”. Forensic methods have:
- validation;
- limitations;
- error rates;
- human factors;
- measurement uncertainty.
The National Academies warned in 2009 that different forensic disciplines did not all have equally strong scientific foundations and that more validation research was needed.[10] NIST now develops statistical approaches precisely to improve the evaluation and communication of uncertainty and error rates in forensic science.[9][11]
“Match” should not mean “identity with no alternative” when the method cannot support that claim
In some pattern-evidence fields, historically there has been a tendency toward very absolute language. But if a method has:
- a known error rate;
- a subjective component;
- incomplete validation,
that belongs in the interpretation. A stronger form of communication is:
what the result supports, and with what uncertainty.
Not:
“science proved identity”
when the method does not justify that level of certainty.
Testimony is evidence—but human memory is not a video recording
A witness may sincerely say:
“I am certain I saw that person.”
That is evidence. It is not automatically reliable evidence. In historical DNA-exoneration cases, the Innocence Project found that mistaken eyewitness identification appeared in a large share of wrongful convictions.[12] That does not mean:
witnesses should not be believed.
It means:
testimony should be evaluated with the same methodological care as other forms of evidence.
High witness confidence is not universally proof of accuracy. Memory can be affected by:
- time;
- stress;
- suggestion;
- identification procedures;
- feedback.
Experimental literature has shown that confirming feedback after an identification can alter a witness's later confidence and memory of the original experience.[13] That is why it matters especially:
how the identification procedure was conducted,
not only:
how confident the witness later appears in court.
A document is strong evidence only if we know its provenance. For a document, check:
- who created it;
- when;
- whether it is complete;
- whether it is authentic;
- in what context;
- for whom it was intended.
An anonymous screenshot of one sentence has much less evidentiary weight than:
a verifiable document from a known archive with full context.
That is why THY-REALITY pays so much attention to a provenance ledger.
A primary document primarily proves its own existence and content. If a memorandum says:
“We propose operation X,”
it strongly supports:
proposal X existed.
It does not prove:
X was approved.
Nor:
X was carried out.
This again confirms our locked rule:
PROPOSAL ≠ APPROVAL ≠ EXECUTION ≠ PROOF OF AUTHORSHIP.
Photographs and video have limited evidentiary scope. A video may document:
- a scene;
- movement;
- sound;
- time, if properly verified.
It does not necessarily establish:
- prior context;
- motivation;
- a person outside the frame;
- authorship of an attack.
In an era of manipulation and generative tools, we also need to verify:
- origin;
- geolocation;
- time;
- file continuity.
Absence of evidence is sometimes informative
The phrase:
“absence of evidence is not evidence of absence”
is often useful. But it is not universal. If a hypothesis predicts:
if X is true, we should very likely observe Y,
and a good search fails to find Y, then the absence of Y reduces support for X. In Bayesian terms:
if the missing evidence would be highly expected under the hypothesis, not observing it counts against the hypothesis.
“There is no evidence” has at least three different meanings
A. We did not look. The evidentiary value of the absence is small.
B. We looked, but the method is not sensitive enough. There is still substantial uncertainty.
C. We looked carefully where the evidence should have been
The absence becomes important evidence against the hypothesis. A good article should therefore explain:
what kind of absence of evidence we are dealing with.
Legal truth and scientific truth use different thresholds
Law uses different burden-of-proof standards. The Cornell Legal Information Institute lists, in the U.S. system for example:
- preponderance of the evidence in many civil matters;
- beyond a reasonable doubt in criminal cases;
- clear and convincing evidence in some other proceedings.[14]
This is only an illustration from a particular legal system. Countries and branches of law use different standards. But the broader idea is:
law must make decisions at some level of uncertainty.
A verdict of “not guilty” does not necessarily mean “proven innocent”. In a system where the prosecution must prove guilt beyond a specified threshold, acquittal can mean:
the burden of proof was not met.
That is a different claim from:
it has been proven that the person did not commit the act.
Likewise:
the allegation was not proven
is not the same as:
the allegation was proven false.
That distinction is extremely important outside law too.
“There is not enough evidence” is not the same as “it did not happen”
THY-REALITY will encounter this position in many historical disputes. We may have:
- motives;
- anomalies;
- partial documents.
But not enough for a stronger conclusion. The correct formulation is then not:
“the event has been refuted.”
It is:
“the currently available evidence is insufficient to confirm this claim reliably.”
And “the official explanation has gaps” is not proof of the alternative
If model A does not explain 100% of the data, it does not follow that:
model B is correct.
B needs positive evidence of its own. This is one of the most important rules for:
- false-flag claims;
- conspiracy theories;
- historical disputes.
Failure of A ≠ proof of B.
An anecdote is evidence, but of a very specific kind. If someone says:
“I felt better after taking the medicine,”
that is real information about their experience. It does not automatically prove:
the medicine caused the improvement.
Possible alternatives include:
- natural course of illness;
- placebo;
- another therapy;
- regression to the mean.
An anecdote can be:
a signal for further investigation.
It is not the same as a controlled estimate of effect.
A case report is especially valuable for detecting signals. In medicine, one unusual case can alert us to:
- a rare adverse effect;
- a new disease;
- an unexpected phenomenon.
That is valuable. But to estimate:
how often the phenomenon occurs
we need a different type of data. Again:
the type of evidence must match the type of claim.
Expert opinion is an evidentiary layer, not a substitute for data. An expert has:
- contextual knowledge;
- experience;
- the ability to integrate complex information.
That is very important. But the opinion is stronger when the expert can show:
which data and methods support it.
“Professor X said so” is weaker than:
“Professor X explained an evidence chain that other qualified people can inspect.”
Extraordinary claims and evidentiary weight. A popular maxim says:
“extraordinary claims require extraordinary evidence.”
A more precise formulation would be:
the more a claim conflicts with well-supported existing knowledge, the more evidentiary weight we need in order to rationally revise our high prior confidence in the existing model.
That is Bayesian logic. It does not mean:
reject a new idea merely because it is unusual.
It means:
a strong existing model already explains a great deal of evidence that the new model must also account for.
Negative evidence can be very strong
If a hypothesis predicts:
phenomenon X should produce Y,
and well-designed experiments repeatedly show:
Y does not occur,
that is important evidence against the hypothesis. Science therefore advances not only through:
confirmations.
It also advances through:
good failed predictions.
Replication changes evidentiary status. One result:
interesting.
Several independent replications:
much stronger.
Different methods with the same basic conclusion:
stronger still.
A systematic review:
shows the wider pattern.
Evidence is therefore not only an object. It can also be a process of accumulation and checking.
Degrees of certainty are more honest than the word “proven”. For THY-REALITY, I would use the following descriptive categories:
DOCUMENTED. Primary or direct sources clearly support the claim.
STRONGLY SUPPORTED. Several independent kinds of evidence point to the same conclusion.
SUPPORTED / PROBABLE. The evidence supports the conclusion, but important limitations remain.
PLAUSIBLE. The hypothesis is possible and consistent with part of the evidence, but the evidentiary base is not strong enough.
SPECULATIVE. There is an idea or some indications, but not enough connecting evidence.
UNSUPPORTED. There is currently no relevant support for the claim.
CONTRADICTED
The available high-quality evidence actively conflicts with the claim. This is not a mathematical scale. It is editorial language designed to prevent false certainty.
EVIDENCE CHAIN. For future articles, we introduce another standard:
CLAIM → EVIDENCE TYPE → PROVENANCE → VALIDITY → INDEPENDENCE → ALTERNATIVE EXPLANATIONS → UNCERTAINTY → CONCLUSION STRENGTH
For every important claim, ask:
- What exactly are we claiming?
- What type of evidence supports it?
- Where does the evidence come from?
- Is the method validated?
- Are the sources independent?
- Which alternatives remain?
- What is the uncertainty?
- How strong a conclusion are we justified in drawing?
Evidence-to-claim matching
This is probably the most important editorial principle. The conclusion must not be stronger than the evidence. Examples:
motive → supports suspicion, not guilt;
correlation → supports association, not necessarily causation;
a document proposing something → supports the existence of the proposal, not its execution;
a photograph of a consequence → supports the consequence, not necessarily attribution;
a failed investigation → supports “not proven,” not necessarily “did not happen”;
statistical significance → supports inconsistency with a specified null model, not practical importance.
The conclusion must be as precise as the evidence
A major public-communication problem often develops like this:
Data. “We observed an association in the sample.”
Abstract. “The results suggest a possible relationship.”
Press release. “Scientists discover the cause.”
Headline
“X causes Y.” Each step increases certainty. By the end, the public receives a much stronger claim than the study actually supports. THY-REALITY should reverse that process:
reduce the conclusion to the strongest wording the evidence genuinely supports.
Evidence is not only “for” or “against”. Evidence can:
- strongly support;
- slightly support;
- change little;
- slightly weaken;
- strongly weaken.
Bayesian logic shows why we do not have to label every new piece of information as either:
proof
or
debunking.
A piece of information may only:
move our confidence by a few percentage points.
“I don't know” is a valid evidentiary result. After reviewing all available sources, we may conclude:
the data do not allow us to reliably distinguish between the competing hypotheses.
That is a result. It is not failure. False certainty is epistemically worse than an honest:
“we do not currently know.”
The next article in the series is therefore devoted to how to live and make decisions under that kind of uncertainty.
THY-REALITY Evidence Status. For every major disputed article, I would therefore use three separate labels:
FACT STATUS. What happened?
ATTRIBUTION STATUS. Who is responsible?
INTERPRETATION STATUS
What does the event mean in the wider context? For example:
the event may be DOCUMENTED;
attribution only PROBABLE;
the broader geopolitical interpretation CONTESTED.
This prevents one strong layer of evidence from “contaminating” other layers with false certainty.
Conclusion: “proven” is often the beginning, not the end, of the question. When someone says:
“This has been proven,”
the most useful response is:
What exactly?
Then:
With what type of evidence?
Under which standard?
With how much uncertainty?
Does the evidence support the event, attribution, intent, or only an association?
Are there independent confirmations?
What would weaken the conclusion?
This does not mean we cannot know anything. On the contrary. We can place very high confidence in some claims precisely because they have survived:
- multiple measurements;
- different methods;
- independent researchers;
- attempts at refutation;
- time.
But even for very strong empirical knowledge, the language:
high confidence
is often more precise than:
absolute certainty.
The greatest intellectual mistake is not only believing without evidence. It is also demanding from evidence more than any empirical method can realistically provide.
Methodological note
This article deliberately distinguishes:
- mathematical proof from empirical evidence;
- an indication from a conclusion;
- statistical significance from practical importance;
- correlation from causation;
- the probability of evidence under a hypothesis from the probability of a hypothesis given the evidence;
- a legal standard of proof from scientific certainty;
- absence of sufficient evidence from evidence that an event did not occur.
U.S. legal standards are used only to illustrate different heights of evidentiary thresholds, not as a universal description of all legal systems.
Sources and further reading
- Stanford Encyclopedia of Philosophy, Confirmation, substantive revision 2025, and Evidence. Used for the relationship between evidence and hypothesis and the fact that true evidence often increases support without producing absolute logical certainty. Source 1 Source 2
- American Statistical Association, Statement on Statistical Significance and P-Values, 2016. Used for the limitations of p-values and the warning that decisions should not rest solely on crossing an arbitrary threshold. Source 1 Source 2
- Cochrane Handbook, Chapter 15, Interpreting results and drawing conclusions. Used for point estimates, confidence intervals, precision, and the distinction between statistical imprecision and other sources of uncertainty in evidence. Source
- Core concepts in statistics and research methods. Part 2: clinical research principles and observational studies, 2025. Used to distinguish association, correlation, and causation. Source
- U.S. National Library of Medicine, Finding and Using Health Statistics – Causation. Used for the role of randomized controlled trials in causal inference. Source
- Association and causation in epidemiology – half a century since the publication of Bradford Hill’s interpretational guidance, 2015. Used to explain that Hill's nine considerations should not be treated as rigid criteria and for discussion of strength, consistency, gradient, and temporality. Source
- Parascandola M, Weed DL., Causation in epidemiology, Journal of Epidemiology and Community Health 55, 2001, and Vandenbroucke et al., Causality and causal inference in epidemiology: the need for a pluralistic approach, 2016. Source 1 Source 2
- Stanford Encyclopedia of Philosophy, Bayesian Epistemology. Used for prior/posterior credence, likelihood, and updating degrees of confidence with evidence. Source
- National Institute of Standards and Technology, Evidential Statistics. Used for likelihood ratios, statistical evaluation of forensic evidence, and the need to communicate uncertainty. Source 1 Source 2
- National Research Council, Strengthening Forensic Science in the United States: A Path Forward, 2009. Used for differences in scientific validation among forensic methods and the need for stronger research foundations. Source
- NIST, Inconclusive Decisions and Error Rates in Forensic Science, 2024/updated 2026, and Human Factors in Forensic Science. Source 1 Source 2
- Innocence Project, DNA Exonerations in the United States (1989–2020) and current exoneration data. Used as a historical sample of documented mistaken identifications in DNA-exoneration cases, not as an estimate of the error rate of all eyewitness testimony. Source 1 Source 2
- Wells GL, Bradfield AL., Distortions in eyewitnesses’ recollections: Can the postidentification-feedback effect be moderated? and APA materials on eyewitness reports. Used for the effect of confirming post-identification feedback on witness confidence and memory. Source
- Cornell Legal Information Institute, Wex, burden of proof, reviewed 2024. Used only as an illustration of different legal thresholds in the U.S. system: preponderance, clear and convincing evidence, beyond a reasonable doubt. Source