pdfdrill

formula reports

What went wrong, and
what was done about it.

Re-measured, all twenty. A defect in the Scan column once pushed every scan crop onto a second line, so the residual was measured against an empty cell and every row scored as a difference. Seventeen reports were re-measured on 6 September. The remaining three — 0902.0431, gilmore-lie-groups and kohlhase-omdoc — were re-measured on 11 September, and every residual column on this page now comes from a measurement taken after that fix. A fourth document, penev_A, was retired from this index on 11 September. Residuals are measured on display equations; inline-formula rows carry a position mark instead — a rectangle showing where in its line an expression was found, drawn on 8,318 of 36,041 rows. The rest carry none deliberately: a mark in the wrong place sends a reader to the wrong glyphs and looks authoritative doing it, so the measurement declines rather than guesses.

Per-equation QC of PDF→LaTeX OCR. residuals.pdf carries the findings; the four evidence files carry every row (equations, inline formulas with their host line's picture, tables, image regions — six columns each, for lookup): a row appears only when there is something to say about it — corrected (the accepted change, both readings against one scan), unresolved (does not render, or two records of a refinement disagree), flagged (an independent ink measurement differs from the reading), and doubted but correct (MathPix was unsure and the ink agrees anyway). A document with nothing in any section is one page saying so. Across 20 documents that is 83 pages where the full listings were 1,230.

Every equation is still published, in formula-report.html and tables.html, which is where the catalogue belongs. The flag column reads shown/total: the flagged section lists the rows whose component delta is 20 or more and states the rest as a count, because 543 rows is a wall and 37 is a morning's work. A high-confidence flag is evidence of a rendering difference, not of a defect — twenty sampled from that band held eighteen purely typographic differences and no content errors.

DocumentSourcePagesEquationsFindingsResidualMeasuredSizeFiles
0902.0431
Exceptional Lie groups · Ichiro Yokota
arXiv 57 of 1,219
not rebuilt
not measured
lattice under-detects — withdrawn
3.5 MB
OpenDownload
1605.05775
Supervised Learning With Quantum-Inspired Tensor Networks · Stoudenmire & Schwab
arXiv 2 26
cor0 unr0 flag0/2 dbt0
W3 N18 K3
2026-09-01 0.04 MB
OpenDownload
1510.06699
arXiv 1510.06699
arXiv 3 279
cor0 unr0 flag1/68 dbt1
C29 W28 S2 N142 K47
2026-09-01 0.12 MB
OpenDownload
2106.07890
An Enriched Category Theory of Language: From Syntax to Semantics · Bradley, Terilla & Vlassopoulos
arXiv 2 51
cor0 unr0 flag0/14 dbt0
C3 W10 N31 K2
2026-09-01 0.04 MB
OpenDownload
2501.06662
The Magnitude of Categories of Texts Enriched by Language Models · Bradley & Vigneaux
arXiv 2 60
cor0 unr0 flag1/14 dbt0
C5 W5 N39 K4
2026-09-01 0.09 MB
OpenDownload
BradleyGastaldiTerilla2023
The Structure of Meaning in Language: Parallel Narratives in Linear Algebra and Category Theory · Bradley, Gastaldi & Terilla
Notices AMS 71(2), Feb 2024 2 8
cor0 unr0 flag0/3 dbt0
W3 N4 K1
2026-09-01 0.04 MB
OpenDownload
bradley_spring22
A New Perspective of Entropy · Tai-Danae Bradley
J. Math3ma Inst. 2 16
cor0 unr0 flag0/3 dbt0
W3 N10 K2
2026-09-01 0.04 MB
OpenDownload
penev_B
Local Feature Analysis (cont.): §4.4.4 – ch. 5 · Penio S. Penev · scanned inkjet print
PhD thesis, Rockefeller Univ., May 1998 2 62
cor0 unr0 flag0/24 dbt0
C7 W15 S5 N31 K1
2026-09-01 0.04 MB
OpenDownload
rph8-g15q
Toward structure-preserving quantum encodings · Parzygnat, Bradley, Vlasic & Pham
Phys. Rev. Research 7, 041001 (2025) 3 81
cor0 unr1 flag0/11 dbt0
C1 W12 S3 N39 K21
2026-09-01 0.07 MB
OpenDownload
mielke-geometrodynamics
Geometrodynamics of Gauge Fields: Yang–Mills and Gravitational Gauge Theories · Eckehard W. Mielke
book 4 1149
cor3 unr2 flag0/207 dbt0
C56 W121 S21 U2 N698 K158
2026-09-01 0.09 MB
OpenDownload
voloshin-hypergraph
Introduction to Graph and Hypergraph Theory · Vitaly I. Voloshin
book 3 266
cor3 unr0 flag0/24 dbt0
C8 W18 S3 N120 K101
2026-09-01 0.06 MB
OpenDownload
johnston-linear-matrix-algebra
Introduction to Linear and Matrix Algebra · Nathaniel Johnston
book 19 1670
cor8 unr7 flag37/543 dbt0
C216 W233 S58 U3 N683 K267
2026-09-01 1.7 MB
OpenDownload
lyche-numerical-linear-algebra
Numerical Linear Algebra and Matrix Factorizations · Tom Lyche
book 8 1056
cor8 unr5 flag1/239 dbt0
C38 W166 S27 U4 N542 K156
2026-09-01 0.22 MB
OpenDownload
gilmore-lie-groups
Lie Groups, Physics and Geometry · Robert Gilmore
book 8 1037
cor0 unr10 flag5/350 dbt1
C53 W225 S40 U1 N491 K100
2026-09-01 0.72 MB
OpenDownload
2010.14265
arXiv 2010.14265
arXiv 2 16
cor0 unr0 flag0/1 dbt0
W5 N6 K4
2026-09-01 0.04 MB
OpenDownload
fong-spivak-invitation
An Invitation to Applied Category Theory · Fong & Spivak · published edition
book 4 262
cor0 unr2 flag7/81 dbt0
C50 W29 S9 U3 N94 K56
2026-09-01 0.30 MB
OpenDownload
kohlhase-omdoc
OMDoc: An Open Markup Format for Mathematical Documents, v1.2 · Michael Kohlhase
book 3 6
cor0 unr4 flag0/1 dbt0
C3 U2 N1
2026-09-01 0.10 MB
OpenDownload
fong-spivak-seven-sketches
Seven Sketches in Compositionality · Fong & Spivak · preprint edition
preprint edition 4 264
cor0 unr1 flag3/78 dbt3
C40 W32 S18 U3 N119 K30
2026-09-01 0.19 MB
OpenDownload
0707.4470
arXiv 0707.4470
arXiv 4 62
cor1 unr0 flag2/13 dbt0
C1 W9 N26 K19
2026-09-01 0.28 MB
OpenDownload
cardona-qft-methods
Geometric, Algebraic and Topological Methods for Quantum Field Theory · Alexander Cardona (ed.)
book 3 1012
cor0 unr1 flag1/110 dbt0
C20 W70 S15 U2 N562 K266
2026-09-01 0.10 MB
OpenDownload

Every report carries the same six columns. Identifier · Page · Conf. · LaTeX source · Rendered · Scan image — the equation's id, the page it sits on, MathPix's confidence in its own reading, that reading as LaTeX, the LaTeX rendered, and the crop the reading came from. The coloured square is the confidence:  ≥0.9,  0.5–0.9,  <0.5, and — for no value.

Seventeen carry a second reading. An inkdrill residual measured between the rendered LaTeX and the scan, knowing nothing about the confidence value, printed as a coloured bullet with a selectable code:  C component,  W weak,  S stable,  U unrendered,  N noise,  K clean. The Residual column above is that distribution per report, and report.ink.json carries the measurements machine-readably. The two instruments disagree usefully — neither predicts the other.

Four do not, and the reason is the point of the column. A residual is only worth printing if each one is attached to the right equation, and rows pair to identifiers by position. Where that pairing cannot be shown to hold, no residual is published rather than a plausible one.

  • BradleyGastaldiTerilla2023not measured. The lattice finds 16 rows for 8 equations on a single page. A count taken after pairing would truncate to 8 and look correct while every row was 2:1 mispaired; the check that catches it compares raw detected rows against identifiers, before truncation.
  • 0902.0431withdrawn, having previously been published here. The lattice detects 55 rows for 57 identifiers: two equations, on report pages 13 and 19, are not detected at all. The cause is enclosure, not size — the lattice builds rows from background enclosed inside the table region, and content spanning from rule to rule leaves none, so the row does not exist as far as it is concerned. Page 13 yields 8 holes, all in the header band, against page 12’s 27; row coverage 0.015 against a corpus median of 0.986. Height only correlates, because a tall row is likelier to span. A row missing mid-sequence does not truncate, it shifts — so from page 13 on, each residual was attached to the following equation. Confirmed by recovering identifiers page by page from the published PDF: exact through page 12, then exactly one displacement per page, always forward, ten of them with no exception. That report has since been rebuilt without the residual apparatus, so the PDF above carries no bullets at all rather than wrong ones; its Conf. column, LaTeX, renders and crops were never affected. If you downloaded its report.ink.json before 2026-08-24, discard it.
  • penev_Bno ink pass yet. It was drilled after the residual pipeline stalled on two open lattice defects, so nothing has been measured against its scans. Not a refusal and not a withdrawal: simply not done. The Conf. column is complete.

One carries a second reading of a different kind. The Penev scan was OCR’d once in 2023 and again in 2026. The 2023 pass was merged into a single LaTeX document and hand-corrected until it compiled; that reconstruction is attached to the 2026 model as a competing tex provenance and scored against it, which is what the compare link opens: one row per equation, both readings rendered side by side, with the scan crop between them.

  • penev_B — 59 of 62 have both. 25 agree exactly, 15 score 0.90–0.99, 5 score 0.70–0.89, and 14 fall below 0.70. Mean 0.874.

This is not a gold column. The author’s LaTeX for this thesis does not exist here — the source is a 1998 print, scanned four pages to a sheet. What the tex column holds is one reader’s corrected transcription, so a disagreement means the two passes differ, not that the OCR is wrong. That is the useful shape: the rows below 0.70 are where a human once decided the machine had misread something, and they are the first place to look.

Equations counts what is in the report, not what is in the document. 0902.0431 is filtered to --max-conf 0.5 --types equation, so its 57 rows are the 57 display equations MathPix doubted out of 1,219. The other seven are unfiltered: every display equation in the document, at whatever confidence.

Sources are linked where they resolve. Crops are excerpts of the papers named above, reproduced for verification of the extraction.

why mathpix

Three things per equation,
from one call.

MathPix returns three things per equation, and this page is built on all three.

latex
rendered
confidence
selects 57 of 1,219
crop
what the render is measured against

The recognised LaTeX, a confidence in that reading, and the crop the reading came from. Each one carries a column of this report: the LaTeX is what gets rendered, the confidence is what selects 57 rows out of 1,219, and the crop is what the render is measured against. lines.json is the only format we have found that ships all three per equation, in one response, each addressed to a rectangle on a page.

The third one is the one that matters here. A confidence score you cannot check is a number you have to trust. Because the crop comes back with the reading, there is something to check it against — which is what the second column of every row below actually is. → the ink residual

The confidence is conservative, and that is the safe direction to be wrong in.

Of 1,219 equations, 57 came back below 0.5 — under 5% flagged as uncertain. inkdrill then measured 55 of those against the scan, knowing nothing about the confidence value. Eleven came back clean or noise: pixel distance 0–7, component delta 0–1. Indistinguishable from the page. The two it could not measure are the two widest in the set — blocks that span the full column, leaving the lattice no enclosed background to read a row from.

So roughly one in five equations at the very bottom of MathPix's own scale renders correctly anyway. It under-rates itself, and that costs review time and nothing else. The opposite error — high confidence on a wrong reading — is the one that quietly corrupts a corpus, and it has not been the problem here. A recogniser that under-rates itself is one you can build a quality-control pipeline on top of. That is why the two readings sit side by side on every row and are never averaged into one score.

lines.json became our interchange format, not just our input.

This was not planned. The MathPix line format turned out to be a good enough description of a page — typed rectangles, per-element geometry, child rectangles inside text lines — that everything else in the toolkit now speaks it. pdfdrill ocr renders pages through tesseract and emits lines.json. inkdrill's emit.py writes its findings as lines.json. The pdfminer.six route was extended to produce it as well, and that extension was offered upstream.

So nothing here is locked to one vendor: the format is the neutral layer, and a keyed MathPix call is one producer of it among several. The document model downstream merges rectangles from every producer into a single containment forest, which is what lets a formula number descend all the way to the ink it was read from. → the model

What it does not claim to do — and where the rest of the toolkit starts.

MathPix reads pixels. Two things sit outside that by construction, and both are ours to solve, not theirs:

The annotation layer

A page-one sentence reading “our code is available here”, where here is a hyperlink with no anchor text, has no pixels at all — the URL lives in the PDF object graph. No recognition route sees it, and none should be expected to. pdfdrill links reads it directly in ~50 ms, keyless. → the invisible link

Content inside images

A table returned as a Picture is a reading decision, not a defect. inkdrill goes back to the ink in that rectangle — blobs, holes, ruled lines — and reports what is in there. The residual is the product. → inkdrill

Every accepted correction in the corpus, on one page — MathPix's reading above, the accepted one below, both against the same scan. → corrections

Neither is an argument against the first instrument. Both are the reason there is a second one.

Support turns out to be a measurable property, and nobody benchmarks it.

Every question sent to the MathPix Admin over the life of this project has been answered inside 24 hours — and answered by reproducing the test case, not by quoting the documentation. For a one-person project that is not a courtesy, it is the difference between a bug fixed on Monday and a bug worked around forever.

The other tools here are volunteer and research projects that owe nobody a reply, and this is not a complaint about them. It is a note that a key buys something a star count does not show: a counterparty.