mirror of
https://github.com/wassname/ml_debug.git
synced 2026-08-25 11:21:20 +08:00
These three held hand-transcribed excerpts (381 / 914 / 573 words) with a header claiming extraction was too hard. It was not: jina reads the sculley and lones PDFs, and plain pdftotext (no -layout) reads all 3 columns of the domingos CACM PDF in order. Now 5736 / 15285 / 8180 words.