Evaluation Rubric

Every paper is scored on whichever dimensions apply to it, each rated 1 to 5. Not all dimensions apply to every paper — a philosophy paper won't have study design or statistical analysis scores, while a data-driven paper won't have an argument soundness score. The pattern of which dimensions apply is itself informative: it tells you what kind of evidence the paper offers.

The overall verdict is non-compensatory: a high score in one area cannot offset a fatal flaw in another. A paper with brilliant statistics but a fundamental construct validity problem is still dubious.

Verdicts

SolidEvidence or argument supports the conclusions.
CredibleLikely right, with caveats.
DubiousProbably doesn't hold up.
NonsenseThe argument doesn't work. Even if the conclusion is true, this paper doesn't show it.

How verdicts are computed

Dimensions are grouped by how much weight they carry. The verdict is driven by the worst score in the most important tier, not by averaging. Dimensions that don't apply to a given paper are excluded from the computation.

TierDimensionsEffect
CriticalStudy design, Argument soundness, Construct validity, Reasoning qualityAny score of 1 → nonsense. Any score of 2 → dubious.
ImportantStatistical analysis, Transparency, Citation integrity, IndependenceTwo or more scores of 2 → dubious.

If no critical or important dimension triggers a downgrade, the paper issolid — unless two or more dimensions across all tiers score 3 or below, in which case it's credible.

Dimensions

The first three dimensions below apply only to certain paper types. The remaining five apply to all papers, though some use different evaluation criteria depending on whether the paper is empirical or argumentative.

Study design empirical papers only

Was the study designed and executed in a way that can support its conclusions?

  • Study type and fit — Is the design (RCT, observational, meta-analysis, etc.) appropriate for the questions being asked?
  • Measurement — Are the measures validated? Could the measurement process itself influence results?
  • Confounds and bias — Are obvious confounding variables acknowledged and addressed? Is there blinding where appropriate?
  • Sample and generalizability — How generalizable are the findings? Does the paper overclaim?

Statistical analysis empirical papers only

Are the statistical methods and reported results sound?

  • Sample and power — Is the sample size adequate? Is there a power analysis? Any selection bias?
  • Statistical tests — Are the tests appropriate? Are assumptions met? Are multiple comparisons corrected?
  • Effect sizes and precision — Are effect sizes reported (not just p-values)? Confidence intervals? Are effects plausible?
  • Red flags — p-values clustering below 0.05, post-hoc outcome selection, selective reporting, numerical inconsistencies.

Argument soundness argumentative papers only

Are the paper's arguments logically sound?

  • Logical validity — Does each step follow from the previous one? Are there unstated premises doing load-bearing work?
  • Completeness — Is the reasoning chain complete? Does the paper establish its premises or assume them?
  • Objection handling — Does the paper anticipate the strongest objections? Are counterexamples handled honestly?
  • Scope match — Do the conclusions stay within what the argument actually establishes?

Construct validity

Is the bridge between what was studied (or argued) and what is claimed structurally sound?

  • For empirical papers: Does the thing being measured correspond to the thing being claimed about? Can the design distinguish between the claimed explanation and plausible alternatives?
  • For argumentative papers: Are key concepts defined precisely enough to apply to new cases? Do definitions do real work — excluding things, generating predictions — or do they accommodate everything?

Reasoning quality

Do the conclusions follow from what the paper offers, or is the framing doing the work?

  • Framing vs. substance — Could you swap in different (plausible) results or alternative conclusions and reach the same framing?
  • Smuggled assumptions — Are normative claims presented as empirical findings? Are conclusions predetermined by definitions?
  • Falsifiability — Could the core claims be shown wrong? Does the paper say what that would look like?
  • Rhetorical moves — Emotional language, appeals to authority, hedged results with unhedged conclusions.
  • Alternative explanations — Does the paper seriously engage with alternative interpretations?

Transparency

How open and reproducible is the research or reasoning?

  • For empirical papers: Pre-registration, open data/materials/code, sample size justification, replication status, limitations.
  • For argumentative papers: Are premises explicit? Can a reader trace every step and locate exactly where they disagree? Are scope limitations acknowledged?
  • Conflicts and funding — Are conflicts of interest and funding sources disclosed?

Citation integrity

Do the cited papers actually say what this paper implies they say?

  • Accuracy — Is each cited work represented fairly, or selectively?
  • Load-bearing analysis — Would the claim collapse if this citation were removed? Each citation is tagged as load-bearing, supportive, decorative, or misrepresented.
  • Citation quality — Are preprints, pilot studies, or retracted papers cited as primary evidence?
  • Missing citations — Are obvious contradictory studies omitted? Does the literature review construct a one-sided narrative?

When a cited paper's full text is available, it's checked directly. Otherwise the abstract is used. Citations that can't be verified against source material are flagged as such.

Independence

Do structural conflicts of interest visibly shape the findings?

  • Disclosed conflicts — Funding, affiliations, consulting, patents, industry positions.
  • Design choices — Does the study include features that could find an unfavorable result (active comparators, pre-registration, independent data analysis)?
  • Alignment of conclusions — Do findings consistently favor the party with the structural interest?
  • Context — Is this an area where industry funding has historically distorted the evidence base?

The question is not whether conflicts exist (they often do and are fine), but whether they create pressure that is visible in the work.

Rating scale

All dimensions use the same 1–5 scale:

ScoreLabelMeaning
5ExemplaryCould serve as a model. No meaningful issues.
4StrongMinor issues that don't affect conclusions.
3AdequateSome issues worth noting but not disqualifying.
2ConcerningMajor issues that substantially reduce confidence.
1UnsoundFundamental problems. This alone is reason to distrust the conclusions.

Dimensions marked N/A on a paper's scorecard were excluded because they don't apply to that type of paper. The pattern of which dimensions apply tells you what kind of knowledge the paper is offering.