How Does Expert Consensus Form?
Peer review, replication, meta-analysis, paradigms, minority views, and why consensus is not the same as truth
When we do not know enough about a complex field, we often ask:
What do the experts think?
That is reasonable. No individual can personally:
- repeat every laboratory measurement;
- read every clinical study;
- verify every geological sample;
- reanalyze all satellite data;
- master the entire statistical literature of a field.
Expert consensus is therefore an important epistemic signal. But a problem appears immediately. If we say:
“Most experts agree, therefore it is true,”
we can turn consensus into an argument from authority. If, on the other hand, we say:
“Scientists have been wrong many times before, therefore consensus means nothing,”
we discard one of the most important mechanisms for collectively checking knowledge. The better question is:
How did the consensus form, what evidence supports it, how robust is it, and does the system allow better evidence to change it?
Consensus is not a vote on truth
Scientific consensus is not a referendum. It is not:
51% of scientists against 49%.
At its best, it means something more complex:
- different research groups;
- using different methods;
- working with different data;
- over time;
arrive at a sufficiently similar conclusion that a claim begins to function as a working foundation of the field. That distinction matters. Consensus does not gain strength because many people think the same thing. It gains strength when agreement emerges from converging evidence.
Why is one expert not enough?
An expert can:
- make an error;
- misinterpret a result;
- choose a poor model;
- become attached to their own theory;
- be influenced by professional incentives;
- use data others cannot verify.
Science is therefore not designed around the idea:
“Find a genius and trust them.”
It is designed around:
a community in which other people can inspect, criticize, and correct an individual's work.
The Stanford Encyclopedia of Philosophy, in its discussion of scientific objectivity, emphasizes the importance of transformative criticism: a scientific community needs channels of criticism, shared standards, genuine uptake of criticism, and sufficient equality of intellectual authority among qualified participants.[1] Objectivity is therefore not only a personal trait:
“this scientist is unbiased.”
It can also be a property of a well-organized process of criticism.
Peer review: a first filter, not a certificate of truth
Before a scientific paper is published, it is often evaluated by other specialists. That is peer review. Reviewers may examine:
- whether the question makes sense;
- whether the method fits the question;
- whether the conclusions are supported by the results;
- whether important limitations are acknowledged;
- whether the author engages with the relevant literature.
The Royal Society describes peer review as a fundamental part of maintaining the quality and progress of the scientific literature.[2] But peer review does not mean:
“The paper has been proven true.”
It means roughly:
“The paper has passed a particular form of expert pre-publication scrutiny.”
A reviewer usually does not repeat the experiment
A peer reviewer normally does not have the time, money, or access required to:
- rerun the experiment;
- recruit a new sample;
- verify every laboratory step;
- forensically inspect all raw data.
Much of the review is based on what is written in the manuscript and accompanying data. Peer review can therefore miss:
- a statistical error;
- an accidental coding mistake;
- a problem that becomes visible only during replication;
- sometimes even fraud.
In its public discussion of peer review, the Royal Society has itself highlighted both sides: the system helps improve papers and provides an important filter, but it is also vulnerable to bias, disagreement among reviewers, and limited ability to detect fraud or error.[3]
Published ≠ settled forever
This is one of the most important rules for reading science. A paper in a prestigious journal may mean:
the result is sufficiently interesting and methodologically acceptable to be shown to the scientific community for further scrutiny.
It does not necessarily mean:
the result has become a permanent scientific fact.
The real test often begins after publication. Other researchers:
- use the finding;
- try to replicate it;
- discover limitations;
- test another population;
- improve the measurement;
- compare it with other studies.
Reproducibility and replicability are not exactly the same thing
The National Academies of Sciences, Engineering, and Medicine noted in its 2019 report that the two terms are often used inconsistently, and therefore drew a clear distinction.[4]
Reproducibility
Can another researcher, using the same data, code, and analytical procedures, obtain the same computational result?
Replicability
Can a new study that tests the same scientific question with new data obtain a result consistent with the earlier finding? These are different tests. A paper can be perfectly reproducible computationally and still produce a finding that does not replicate in new data.
A failed replication does not always mean “the original study was a lie”
A replication can fail for many reasons:
- the original result was a chance finding;
- the original effect estimate was exaggerated;
- the original sample was too small;
- the new study failed to reproduce a crucial condition;
- the effect depends on context;
- the populations differ;
- an important moderator is missing from the theory.
The National Academies therefore emphasizes that replicability requires careful interpretation and that inconsistency can sometimes lead to new scientific discoveries.[4] So:
failed replication ≠ automatic refutation.
But multiple high-quality failures to replicate do, of course, reduce confidence in the original claim.
The replication crisis in psychology
In 2015, the Open Science Collaboration attempted to replicate 100 experimental and correlational findings from three psychology journals.[5] In the original studies, about:
97% of the findings were statistically significant.
In the replications:
36%.
The average effect size in the replications was about half the size of the original effects.[5] That was an important warning signal. It was not proof that:
“psychology does not work.”
The project examined a specific sample of studies from three journals and used several different criteria of replication success.
By 2026, we also have a larger newer result
A large study published in Nature in April 2026 attempted to replicate 274 positive claims from 164 papers in the social and behavioral sciences.[6] By one of the principal criteria, about:
55% of claims
produced a statistically significant result in the same direction. That is higher than the famous 2015 project, but still far from complete replication.[6] The important lesson is not one magical percentage. It is:
replicability is an empirical property of a claim that has to be tested, not assumed from the prestige of the journal.
We should not generalize the “replication crisis” to all of science without distinction
Different fields work with very different kinds of data. Replicating:
- a laboratory psychology experiment;
- an astronomical event;
- an epidemic;
- a paleontological discovery;
- a clinical trial;
is not the same task. NASEM therefore warns against using one global non-replication rate as a universal measure of the health of science.[4] Some researchers also argue that “crisis” rhetoric is too broad in certain fields and that reforms should be adapted to the type of research being conducted.[7]
Science responded to the problem
The replication debate is not only a story about weakness. It is also a story about self-correction. There has been greater use of:
- preregistration;
- registered reports;
- open data;
- open code;
- more transparent peer review;
- multi-site replication projects;
- improved statistical reporting;
- metascience.
A 2025 Nature Communications review examined how different scientific fields are adopting reforms to improve reproducibility and transparency.[8] That matters for the entire series:
a system that detects its own weakness and changes its procedures is epistemically different from a system that interprets weakness as evidence of its own infallibility.
A systematic review is stronger than a single study—but not automatically
If we have ten studies on the same question, we can conduct a systematic review. Its strength lies in specifying in advance:
- a search strategy;
- inclusion criteria;
- the method for assessing bias;
- the procedure for synthesizing results.
That reduces the danger of:
“I selected only the studies that support my view.”
But a systematic review is only as good as its protocol and the literature that enters it.
Meta-analysis is not “voting among studies”
Meta-analysis statistically combines results from multiple studies. That can be extremely powerful. It allows:
- greater statistical precision;
- estimation of an average effect;
- comparison among studies;
- investigation of heterogeneity.
But the final number is not magical truth. Cochrane emphasizes that meta-analysis should always assess heterogeneity between studies, and that statistical pooling can sometimes be misleading.[9] If studies measure things that are too different, an average can hide more than it reveals.
Heterogeneity is information
Imagine five studies. Three show benefit. Two show harm. The average result is:
close to zero.
That does not necessarily mean:
there is no effect.
Perhaps there are two different populations. Perhaps the effect depends on:
- dose;
- age;
- context;
- method.
Cochrane therefore requires heterogeneity to be examined and pooled results not to be interpreted without understanding variation among studies.[9]
GRADE: “How certain are we?”
Cochrane and many health guidelines use the GRADE approach. It classifies the certainty of evidence as:
- high;
- moderate;
- low;
- very low.[10]
The assessment considers:
- risk of bias;
- inconsistency;
- indirectness;
- imprecision;
- publication bias.[10]
This is a very useful epistemic model. Instead of:
“proven / unproven”
we get:
how much confidence we have in the estimate, and why.
The point of consensus is not the point of certainty
Experts can agree while still acknowledging:
- an interval of uncertainty;
- open questions;
- exceptions;
- varying quality among different pieces of evidence.
A good consensus is not:
“we know everything.”
It is:
“given the current evidence, this is the most robust shared conclusion.”
National Academies 2026: consensus has to be properly formed
The fourth edition of the National Academies guide On Being a Scientist, published in September 2026, includes an important warning.[11] Scientific debates and important disagreements should be visible to the public. When communicating consensus, it should be clear:
- what kind of agreement exists;
- what important disagreements remain;
- what the consensus rests on.
The guide emphasizes that consensus is trustworthy only when it is properly formed:
through robust debate, consideration of multiple perspectives, and rigorous evaluation of evidence.[11]
That is probably the best short definition of healthy expert consensus for this project.
Artificial unanimity can reduce trust
If experts tell the public:
“everyone agrees”
when in fact an important methodological debate exists, communication may seem simpler in the short term. In the long term, it is risky. If the public later discovers:
- disagreement;
- a change in guidelines;
- new evidence;
people may conclude:
“They lied before.”
That is why the National Academies warns that artificial unanimity is dishonest and can reduce trust when uncertainty later becomes visible.[11]
But “there is disagreement” does not mean “the field is split 50:50”
This is the opposite communication error. If 95 research groups support model A and five support model B, it is not fair to say:
“experts cannot agree.”
A minority view may matter. But the size and quality of support should be visible. Dissent should be represented in proportion to the evidence, not simply to the fact that it exists.
The minority can be right
The history of science includes cases in which a minority position was later accepted. That is one reason science must allow:
- disagreement;
- new models;
- anomalies;
- criticism of the prevailing paradigm.
But this does not imply:
“Because some historical dissenters were right, today's dissenter is probably more right than the consensus.”
For every historically successful dissenter, there have been many mistaken minority ideas that were correctly rejected. Minority status is not evidence. Evidence is.
Galileo is not a universal argument
“People did not believe Galileo either” is rhetorically powerful. It is not evidence for a specific modern claim. The correct lesson from historical dissenters is:
institutions must allow better evidence to defeat authority.
It is not:
everyone who opposes the experts is the new Galileo.
Kuhn: science needs paradigms
Thomas Kuhn profoundly influenced how scientific development is understood. In his view, mature science often operates within a paradigm, or shared disciplinary framework:
- foundational theories;
- instruments;
- methods;
- exemplars of good problem-solving.[12]
Such a shared framework is useful. Researchers do not need to debate every day:
whether the field's basic equations work at all.
They can focus on more specific problems.
A paradigm is not the same as dogma
Kuhn is often simplified as saying:
“Science is just social fashion; one paradigm replaces another.”
That is not a good summary. The Stanford Encyclopedia emphasizes that Kuhn himself rejected interpretations in which scientific change is simply determined by political or social factors.[12] In theory choice he emphasized values such as:
- accuracy;
- consistency;
- scope;
- simplicity;
- fruitfulness.[12]
Scientific communities are social. That does not mean:
evidence is irrelevant.
Normal science is conservative—and that is not necessarily a flaw
If scientists abandoned a foundational theory after every anomaly, research could not progress. Every experiment contains:
- measurement error;
- chance;
- incomplete data.
A degree of conservatism is therefore rational. A new theory generally has to explain:
- what the old model already explained well;
- plus the anomalies the old model could not resolve.
This creates a balance:
stability without complete rigidity.
When does conservatism become a problem?
When anomalies:
- accumulate;
- recur;
- appear across different methods;
- require ever more ad hoc explanations.
At that point, willingness to change should increase. The system should not use:
“this is our consensus”
as a way to prevent investigation of the anomaly. Consensus should be the result of inquiry, not a ban on inquiry.
Expertise and social status are not the same thing
In science, some people have more:
- citations;
- prestige;
- laboratory funding;
- institutional power.
That can affect debate. But the ideal of a scientific community is:
an argument is not correct because the most famous professor made it.
The Stanford discussion of Longino's approach therefore also highlights equality of intellectual authority among qualified practitioners as part of transformative criticism.[1] This does not mean all opinions are equally valid. It means a qualified argument should be judged by the standards of the field.
Peer review can protect quality and still inhibit novelty
This is a tension that cannot be eliminated completely. A reviewer should reject:
- poor methods;
- unsupported conclusions.
But a genuinely novel model may use:
- unusual assumptions;
- different terminology;
- new methods.
The Royal Society has also discussed the concern that traditional peer review can sometimes be too conservative toward original work.[3] Some systems therefore separate:
scientific soundness
from
novelty / perceived importance.
Consensus is not measured only by polling scientists
Surveys of experts can be useful. But stronger indicators may include:
- textbooks;
- guidelines;
- systematic reviews;
- consensus reports;
- repeated findings;
- practice across independent laboratories.
An expert can give an opinion in a survey. The literature shows:
what that opinion ought to rest on.
A consensus report is different from one expert's paper
National Academies Consensus Study Reports are produced through:
- a panel of experts;
- review of the evidence;
- deliberation;
- an independent peer-review process.[13]
That does not mean:
the report is infallible.
It means it has a different epistemic structure from:
a column written by one professor.
For THY-REALITY, we therefore always need to know:
What kind of expert document are we looking at?
A systematic review, guideline, and consensus statement are not the same thing
Systematic review
Systematically synthesizes research according to a prespecified protocol.
Meta-analysis
Statistically combines results of compatible studies.
Clinical/policy guideline
In addition to evidence, it often considers:
- benefits;
- harms;
- feasibility;
- costs;
- values.
Consensus statement
Represents a coordinated expert conclusion of a particular group. So a recommendation:
“do X”
is not necessarily a direct scientific claim:
“X is biologically true.”
It may also contain a value judgment.
Consensus can contain values
When moving from evidence to policy, the question often becomes:
How much risk are we willing to accept?
Science can estimate:
the risk is 1 in 10,000.
It cannot by itself determine:
whether that risk is socially acceptable.
So in expert recommendations it is useful to separate:
- the empirical component;
- the normative component.
This prevents a political or ethical choice from acquiring the false appearance of:
“science requires only one option.”
Two legitimate expert groups can derive different recommendations from the same evidence
Even if both groups accept the same data, they can differ on:
- risk thresholds;
- costs;
- priorities;
- the precautionary principle.
That does not necessarily mean:
“science is incompetent.”
It may mean:
the data are similar, while the value weights differ.
THY-REALITY should therefore always ask of expert recommendations:
Where does evidence assessment end and policy judgment begin?
“Follow the consensus” is a good default, not an absolute rule
For a non-specialist, it is usually rational to give more weight to:
- robust expert consensus
than to:
- one isolated online post.
But that is a default, not dogma. Confidence in a consensus should be greater when:
- it is supported by multiple methods;
- it has independent replications;
- dissent is treated transparently;
- data are accessible;
- conflicting incentives are dispersed.
Confidence should be lower when:
- the field rests on very little data;
- everyone uses the same dataset;
- key findings have not been replicated;
- missing results are likely;
- access to evidence is monopolized.
A thousand papers are not a thousand independent pieces of evidence
This is extremely important. A large literature may rely on:
- the same dataset;
- the same cohort;
- the same measure;
- the same foundational assumption.
The apparent number of publications is therefore not the same as:
epistemic independence.
When evaluating consensus, ask:
How many genuinely independent paths lead to the same conclusion?
Convergence beats repetition
The strongest case for consensus is not:
50 studies using the same method.
A better case is:
- a laboratory experiment;
- field data;
- a longitudinal study;
- an independent theory;
- results from different countries;
all pointing in a similar direction. That reduces the chance that all findings arise from:
the same hidden methodological error.
Meta-analysis cannot fix a monoculture of evidence
If all ten studies were conducted:
- with the same faulty measurement device;
a meta-analysis may estimate the wrong effect with great precision. So the strength of consensus is not only:
quantity.
It is:
diversity of independent methods and sources.
Consensus is stronger if it includes former skeptics
If different groups:
- started with different hypotheses;
- gradually converged as new evidence accumulated;
that is often a stronger signal than a field where agreement existed before most of the new evidence. This is not a formal rule. But it is a useful indicator:
Do the data actually change minds?
Consensus is weaker if disagreement is unsafe
If a scientist can lose:
- a career;
- access to data;
- the ability to publish;
simply for methodologically challenging the dominant model, agreement becomes harder to interpret as independent convergence. That does not mean every professional conflict is evidence of censorship. It means healthy consensus requires a real possibility of expert dissent.
But dissent also has to survive criticism
Protecting minority opinion does not mean:
all opinions deserve equal treatment.
A dissenter must show:
- data;
- a method;
- an answer to criticism;
- a prediction that can be tested.
If a minority model responds to every failure by adding:
“the evidence is being hidden,”
we move from scientific debate toward a self-sealing explanation.
When can we say a consensus is truly robust?
For THY-REALITY, I would use several criteria.
Evidence diversity
Multiple independent methods.
Replication
The effect appears in new data.
Transparency
Visible protocols, data, or at least enough information for verification.
Synthesis
High-quality systematic reviews/meta-analyses exist.
Heterogeneity understood
Differences among results are not simply hidden inside an average.
Dissent visible
Important opposing arguments are publicly addressed.
Conflict diversity
The result is not dependent on one funder or institution.
Correction capacity
The field can change direction.
CONSENSUS CHAIN
For future articles, I would use the standard:
CLAIM → PRIMARY STUDIES → REPLICATIONS → SYSTEMATIC REVIEW → META-ANALYSIS / EVIDENCE GRADE → EXPERT DELIBERATION → CONSENSUS STATEMENT / GUIDELINE → LATER REVISION
At each link, ask:
- how many independent sources there are;
- what their quality is;
- where the uncertainties are;
- whether dissent exists;
- what could change the consensus.
Five degrees of expert agreement
For practical use, we can describe consensus as:
Emerging
A few promising studies, many open questions.
Converging
Several independent studies point in the same direction.
Broad
A large majority of relevant evidence and experts supports the same basic model.
Mature
The model has been tested over time, is supported by multiple methods, and is used as the ordinary foundation of the field.
Foundational
The claim is so well integrated across independent systems of evidence that abandoning it would require explaining a very large body of successful evidence. This is not an official scientific classification. It is a THY-REALITY editorial heuristic so we do not rely only on the phrase:
“the experts say.”
How should we treat a legitimate minority view?
If the minority position is:
- published;
- methodologically serious;
- directly connected to relevant data;
we present it. But we state clearly:
how much support it has.
For example:
“This is a minority interpretation defended by X; the majority of the current evidence supports Y.”
Not:
“There are two sides.”
when the evidentiary support is not symmetrical.
How should we treat a marginal claim?
If a claim:
- has no primary data;
- does not appear in the specialist literature;
- relies mainly on videos or blogs;
there is no need to artificially place it next to a robust consensus as an equal alternative. We may mention it if it is socially important. But its epistemic status should remain clear.
What if there is no consensus yet?
Then we do not invent a winner. We use:
CONTESTED / EVIDENCE INCOMPLETE
and explain:
- what is agreed on;
- what remains disputed;
- what data would help resolve the disagreement.
That is much more useful than:
“science knows nothing.”
What happens when consensus changes?
A change in consensus is not automatically proof that:
“scientists knew nothing before.”
It can reflect:
- new data;
- a better method;
- a larger sample;
- a corrected measurement.
The scientific system is valuable partly because it must be revision-capable. The bigger problem would be if consensus could never change even in the face of strong contrary evidence.
Revising consensus should preserve its history
A field should be able to show:
- what it previously believed;
- what evidence supported it;
- which new data changed the conclusion.
This is the same changelog logic introduced in the article on institutional error. A history of change does not inherently reduce credibility. When transparent, it can strengthen it.
Consensus and THY-REALITY
In future articles we will not simply write:
“scientific consensus says X.”
We will check:
- who measured or formulated the consensus;
- which disciplines are included;
- how much evidence there is;
- what kinds of evidence;
- which relevant dissent exists;
- where the limits of the claim are.
Consensus will therefore be an evidentiary layer, not a rhetorical endpoint.
What should a non-expert do?
If someone does not have time to review an entire specialist field:
- find a high-quality recent systematic review;
- check whether multiple relevant professional bodies have a consensus statement or guideline;
- examine certainty and limitations;
- ask whether dissent is methodologically serious or mainly media-driven;
- do not build a conclusion on one study;
- watch whether the finding replicates.
This does not make the person an expert. It does allow them to use expert systems more intelligently.
Conclusion: good consensus is not the opposite of doubt
The healthiest expert consensus is not a system that says:
“the question is closed.”
It is a system that says:
“this is currently the most robust conclusion, but we know what kind of new evidence could change it.”
Its strength is not perfect unanimity. Its strength is that it has survived:
- criticism;
- replication;
- different methods;
- competing explanations.
So expert consensus should neither be:
worshipped blindly
nor
dismissed automatically.
We need a third option:
examine how it was formed.
If a consensus was built:
- through open debate;
- from multiple independent kinds of evidence;
- with room for dissent;
- with transparent methods;
- and with a real capacity for correction,
it is rational to give it substantial weight. Not because the majority is infallible. But because a well-designed collective process can be more reliable than individual intuition. And it must still remain something that future reality can correct.
Methodological note
This article does not claim that:
- peer review proves a paper true;
- a failed replication automatically refutes the original claim;
- the replication crisis describes every scientific field equally;
- meta-analysis automatically represents the highest truth;
- expert consensus means complete certainty;
- minority views are automatically wrong;
- historical changes in consensus prove modern consensus should be ignored;
- Kuhn's concept of paradigms means evidence is only a social construct.
Consensus is treated as an epistemic signal whose strength depends on the process by which it was formed.
Sources and further reading
- Stanford Encyclopedia of Philosophy, Scientific Objectivity. Used for transformative criticism, channels of criticism, shared standards, uptake, and equality of intellectual authority among qualified participants. Source
- Royal Society, Peer review, editorial standards and processes and Reviewing with the Royal Society. Used for the purpose and organization of expert review. Source 1 Source 2
- Royal Society, Peering at Review: is peer review fit for purpose?, 2015. Used as a balanced institutional discussion of the value and limitations of peer review. Source
- National Academies of Sciences, Engineering, and Medicine, Reproducibility and Replicability in Science, 2019. Used to distinguish reproducibility from replicability, discuss causes of failed replication, and caution against universalizing the idea of a scientific “crisis.” Source 1 Source 2
- Open Science Collaboration, Estimating the reproducibility of psychological science, Science 349(6251), 2015. Source 1 Source 2
- Investigating the replicability of the social and behavioural sciences, Nature 652, 143–150, 2026. Used for the newer large replication project: 274 claims from 164 papers and about 55% of claims with a statistically significant result in the original direction. Source
- Lash TL et al., The replication crisis in epidemiology: snowball, snow job, or winter solstice? Used as a caution that the term replication crisis is not equally suitable for all fields and that reforms themselves require evaluation. Source
- Nature Communications, Reproducibility and transparency: what’s going on and how can we help, 2025. Used for contemporary open-science reforms and differences among fields. Source
- Cochrane Handbook, Chapter 10, Analysing data and undertaking meta-analyses. Used for heterogeneity, limits of statistical pooling, and caution in interpreting average effects. Source
- Cochrane Handbook, Chapter 14, Completing ‘Summary of findings’ tables and grading the certainty of the evidence. Used for GRADE and the certainty domains: risk of bias, inconsistency, indirectness, imprecision, and publication bias. Source
- National Academies of Sciences, Engineering, and Medicine, On Being a Scientist: A Guide to the Responsible and Ethical Conduct of Research, 4th edition, 2026, module Earning and Maintaining Trust. Used for the requirement that scientific dissent remain visible, the warning against artificial unanimity, and the criterion that consensus is trustworthy only when properly formed through robust debate and rigorous evaluation of evidence. Source
- Stanford Encyclopedia of Philosophy, Thomas Kuhn, substantive revision 2025. Used for paradigms, normal science, scientific revolutions, values in theory choice, and the warning against a simplified sociological interpretation of Kuhn. Source
- National Academies, Process and description of Consensus Study Reports. Used to distinguish a consensus report from an individual expert opinion and to describe the independent review process. Source 1 Source 2