PDF copy-paste text fixer
Join up the line breaks, put back words split by a hyphen, and drop the page numbers and running headers that come with text copied out of a PDF. Every rule is a switch you can turn off.
Paste the text from the PDF
Your text stays in this tab. The work is done by JavaScript in your browser. None of it is uploaded, logged or saved, and the tool keeps working with the network off.
What to repair
A PDF does not record where its paragraphs are, so these three work it out from the shape of the lines. They are right on ordinary prose and wrong on poetry, addresses and tables, where a short line was deliberate. Compare the two versions in the panel before you use the result.
Repaired text
The repaired text appears here.
A PDF has no paragraphs in it
This is the fact everything else follows from, and it surprises most people. A PDF does not contain paragraphs, or sentences, or even words. It contains instructions to draw runs of letters at particular coordinates on a page.
The paragraph you can see is an arrangement of separate lines that happen to sit under one another at even spacing. Nothing in the file records that they belong together, because nothing needed to: the file’s job was to put ink in the right places.
So when you select text and copy it, your viewer has to reconstruct something that was never written down, and it does the safe thing: one line of text per drawn line. That is why you get a column of fragments instead of prose.
Three other things come out of the same fact.
The hyphens are real. When a word will not fit at the end of a line it is broken with a hyphen, and that hyphen is a character in the file like any other. There is nothing marking it as different from the one in “self-esteem”.
The page furniture is in the flow. A running header and a page number are drawn the same way as the body text, so they come out interleaved with it, once per page, in the middle of your sentences.
Some letters are joined. A typesetter sets “fi” as a single shape so the f’s hook does not collide with the i’s dot, and in the file it is one character rather than two. It looks right and it defeats a search: looking for “field” will not find it.
Which repairs know, and which repairs guess
The switches are in two groups on purpose, because the difference tells you where to look when something comes out wrong.
The ones that know. Expanding a ligature has one right answer. A line containing nothing but a number is a page number. A run of spaces left by justified text is a run of spaces. These cannot really be wrong.
The ones that guess. Whether a paragraph ended, whether a repeated line is a header, and whether a hyphen was the typesetter’s are all inferences from what is left after the layout was thrown away. They are worked out from line lengths: a line that ran to the width of the column was wrapped and carries on, and a line that stopped short and ended a sentence finished its paragraph.
That reasoning is sound for ordinary prose and wrong wherever a short line was deliberate. Poetry, addresses, tables, lists of names and anything in two columns will come out worse rather than better. Turn the paragraph rebuilding off for those and the repairs that know what they are doing still work.
The hyphen rule has one case it reliably gets wrong, and it is worth knowing: a genuine compound that happens to break at its own hyphen. “Anglo-” followed by “Saxon” is safe, because the capital gives it away. “Self-” followed by “esteem” is not, and comes out as one word.
What this deliberately is not
It is not a PDF reader. Reading the file would be the better way to do all of this: the coordinates of every line are in there, and the columns, the headers and the paragraph breaks could be worked out properly from them rather than inferred. But by the time text has been through the clipboard those coordinates are gone, and no amount of cleverness gets them back.
It is also not an OCR tool. If your PDF is a scan and selecting the text gives you nothing at all, there is no text in the file to repair: what you are looking at is a photograph of a page, and it needs recognising rather than tidying.
Everything here runs in this tab, on text you pasted. There is no request in the page that could send it anywhere, nothing is stored, and it keeps working with the network off. The panel has a toggle between the repaired version and what you pasted, which is how to check any of the above rather than take it on trust.
Questions
Why does text copied from a PDF break into lines?
Why are words split with hyphens in the middle?
It joined a word that should have kept its hyphen. Why?
What counts as a page number?
How does it know which lines are headers?
It ran two paragraphs together. Can that be fixed?
What are the joined-up letters?
Does it read the PDF itself?
Does my text leave the browser?
More tools
Word counter
Words, sentences, paragraphs and how long it takes to read
Character counter
With spaces, without, and what a character even is
Case converter
Seven cases at once, no switching back and forth