← Back to Hub
📜 Scharp Scale Experiment

Philosophy Paper Quality:
Noûs 6.57, PPR 5.82

76 philosophy papers scored on the Scharp Scale (0-10, where 7 = average Nature paper). Real data from a March 2026 scoring experiment across three leading philosophy journals.

76Philosophy papers scored
6.57Noûs mean score (Scharp Scale)
5.82PPR mean score
6.27Philosophers' Imprint mean score

Choose your depth. The data doesn't change — just the explanation.

A philosopher created a scoring system to rate how good philosophy papers are, from 0 to 10. A 7 on this scale means the paper is as good as a typical paper in Nature, a famous science journal. Then 76 recent philosophy papers from top journals were scored. The average Noûs paper scored 6.57. Philosophy & Phenomenological Research (PPR) averaged 5.82. Philosophers' Imprint averaged 6.27. Some papers scored an 8 or higher — genuinely excellent work.
The Scharp Scale (0-10) calibrates philosophy paper quality against scientific norms: 7 = average Nature paper, 8 = high-impact Nature paper, 9 = landmark paper (Nobel-adjacent), 10 = Gödel's incompleteness theorems. A March 2026 experiment scored 76 papers from three journals: Noûs (15 papers), PPR (13 papers), and Philosophers' Imprint (48 papers). Key finding: Noûs outperforms PPR by 0.75 points on average. Several papers — Blame's Topography (8.0), Symmetries of Value (7.8), Zermelian Extensibility (7.8) — score at genuinely excellent levels.
Scharp Scale calibration: 7 = average Nature paper (rigorous, original, significant); 8 = high-impact (top 10% of Nature); 9 = landmark (redefines a field); 10 = generational. Philosophy calibration challenge: analytic philosophy lacks the standardized empirical feedback loops of science, making quality assessment inevitably more subjective. The three journals selected are among the top 5 in analytic philosophy by citation impact. PPR's lower mean may reflect broader topic coverage (empirical, applied) vs. Noûs's technical analytic focus. Philosophers' Imprint variance is highest due to largest sample. Top scorers: Blame's Topography (Reis-Dennis, PI, 8.0), Symmetries of Value (Goodsell, Noûs, 7.80), Zermelian Extensibility (Bacon, PPR, 7.80). Notable low score: Consciousness Makes Things Matter (Lee, PI, 4.0) — striking title, apparently not commensurate execution.
Raw data: /home/node/.openclaw/workspace/experiments/philosophy-scored-papers-2026-03-29.csv. Journals: Philosophers' Imprint (n=48), PPR (n=13), Noûs (n=15). Score range: 4.0-8.0. Scale anchor: 7 = average Nature paper. Scharp Scale documentation: Kevin Scharp's personal calibration, March 2026. The experiment was designed to test whether AI-assisted quality scoring of philosophy papers could produce reliable, interpretable signal across venues. Raw CSV available at path above.

Score Distributions by Journal

Score Distribution — All 76 Papers

Frequency of papers at each score level (0.5 bins)

Mean Score by Journal with Range

Journal-level comparison — higher is better

Top 15 Papers by Score — All Journals

Highest-scoring papers across the entire dataset

All Papers — Full Ranking Table (Sorted by Score)

Complete dataset — 76 papers, 3 journals, Scharp Scale

ScoreVenueTitle (abbreviated)Author(s)

📊 What the Scores Mean

On the Scharp Scale, 7 is not "excellent" — it's the average for Nature, arguably the most selective scientific journal. Most philosophy papers in this dataset score between 5.5 and 7.5. A score below 6.0 suggests solid professional work that makes a genuine contribution but without the originality or significance that would define a landmark paper. A score of 8.0 (Blame's Topography, Symmetries of Value) indicates a paper that would be transformative in its subfield — the kind that gets assigned in graduate seminars a decade later. The gap between Noûs (6.57) and PPR (5.82) — 0.75 points — is substantial when the entire active range is roughly 4 points wide.