Skip to content
ToolShelf

PDF copy-paste text fixer

Join up the line breaks, put back words split by a hyphen, and drop the page numbers and running headers that come with text copied out of a PDF. Every rule is a switch you can turn off.

Paste the text from the PDF

Your text stays in this tab. The work is done by JavaScript in your browser. None of it is uploaded, logged or saved, and the tool keeps working with the network off.

What to repair

Repairs with one right answer
Repairs that have to guess

A PDF does not record where its paragraphs are, so these three work it out from the shape of the lines. They are right on ordinary prose and wrong on poetry, addresses and tables, where a short line was deliberate. Compare the two versions in the panel before you use the result.

Repaired text

The repaired text appears here.

A PDF has no paragraphs in it

This is the fact everything else follows from, and it surprises most people. A PDF does not contain paragraphs, or sentences, or even words. It contains instructions to draw runs of letters at particular coordinates on a page.

The paragraph you can see is an arrangement of separate lines that happen to sit under one another at even spacing. Nothing in the file records that they belong together, because nothing needed to: the file’s job was to put ink in the right places.

So when you select text and copy it, your viewer has to reconstruct something that was never written down, and it does the safe thing: one line of text per drawn line. That is why you get a column of fragments instead of prose.

Three other things come out of the same fact.

The hyphens are real. When a word will not fit at the end of a line it is broken with a hyphen, and that hyphen is a character in the file like any other. There is nothing marking it as different from the one in “self-esteem”.

The page furniture is in the flow. A running header and a page number are drawn the same way as the body text, so they come out interleaved with it, once per page, in the middle of your sentences.

Some letters are joined. A typesetter sets “fi” as a single shape so the f’s hook does not collide with the i’s dot, and in the file it is one character rather than two. It looks right and it defeats a search: looking for “field” will not find it.

Which repairs know, and which repairs guess

The switches are in two groups on purpose, because the difference tells you where to look when something comes out wrong.

The ones that know. Expanding a ligature has one right answer. A line containing nothing but a number is a page number. A run of spaces left by justified text is a run of spaces. These cannot really be wrong.

The ones that guess. Whether a paragraph ended, whether a repeated line is a header, and whether a hyphen was the typesetter’s are all inferences from what is left after the layout was thrown away. They are worked out from line lengths: a line that ran to the width of the column was wrapped and carries on, and a line that stopped short and ended a sentence finished its paragraph.

That reasoning is sound for ordinary prose and wrong wherever a short line was deliberate. Poetry, addresses, tables, lists of names and anything in two columns will come out worse rather than better. Turn the paragraph rebuilding off for those and the repairs that know what they are doing still work.

The hyphen rule has one case it reliably gets wrong, and it is worth knowing: a genuine compound that happens to break at its own hyphen. “Anglo-” followed by “Saxon” is safe, because the capital gives it away. “Self-” followed by “esteem” is not, and comes out as one word.

What this deliberately is not

It is not a PDF reader. Reading the file would be the better way to do all of this: the coordinates of every line are in there, and the columns, the headers and the paragraph breaks could be worked out properly from them rather than inferred. But by the time text has been through the clipboard those coordinates are gone, and no amount of cleverness gets them back.

It is also not an OCR tool. If your PDF is a scan and selecting the text gives you nothing at all, there is no text in the file to repair: what you are looking at is a photograph of a page, and it needs recognising rather than tidying.

Everything here runs in this tab, on text you pasted. There is no request in the page that could send it anywhere, nothing is stored, and it keeps working with the network off. The panel has a toggle between the repaired version and what you pasted, which is how to check any of the above rather than take it on trust.

Questions

Why does text copied from a PDF break into lines?
Because a PDF does not contain paragraphs. It contains instructions to draw runs of letters at particular positions on a page, and that is genuinely all it records. The paragraph you can see is an arrangement of separate lines that happen to sit under one another, and nothing in the file says they belong together. So when your viewer copies the text out, it plays safe and gives you one line of text per drawn line.
Why are words split with hyphens in the middle?
Because the typesetter put them there. When a word will not fit at the end of a line, it is broken with a hyphen, and that hyphen is a real character in the file exactly like every other. The file has no way of marking it as different from the hyphen in a compound word, which is why rejoining them is a guess rather than a certainty.
It joined a word that should have kept its hyphen. Why?
Because the two cases look identical in the file. A word broken at a line end and a compound word that happens to break at its own hyphen are the same characters in the same order. The only signal available is the letter that follows: a lower-case one is treated as a continuation and joined, a capital is left alone, so Anglo-Saxon survives and end-ing does not. Compounds like self-esteem broken at exactly that point are the case it gets wrong, and they are worth a look in the comparison view.
What counts as a page number?
A line that is nothing else: a bare number, a number in brackets or between dashes, 3 of 18, 3/18, or a roman numeral, which is how front matter is numbered. A line with a number in a sentence is not touched. The one genuine trap is a short roman numeral, because mix, did and mild are all made of numeral letters, so a document with single words on their own lines is worth checking.
How does it know which lines are headers?
By how often they repeat, because their position is exactly what the copy lost. A short line that appears three times or more is treated as a running header or footer. Long lines are left alone at any frequency: a clause repeated in every schedule of a contract is somebody's content, not page furniture, and removing it would be the tool deciding their document has a mistake in it.
It ran two paragraphs together. Can that be fixed?
Sometimes only by hand. Whether a paragraph ended is worked out from the shape of the lines: one that runs to the width of the column was wrapped, and one that stops short and ends a sentence finished the paragraph. That is right on ordinary prose and wrong where a short line was deliberate, which is poetry, addresses, tables and lists of names. Turn off the paragraph rebuilding for those and the rest of the repairs still work.
What are the joined-up letters?
Ligatures. A typesetter sets fi, fl, ff and ffi as single shapes so the letters do not collide, and in the file they are single characters rather than two. Copied out, they look right and defeat a search: looking for field will not find it, because the fi in it is not an f followed by an i. Expanding those has exactly one right answer, which is why it is in the group of repairs that do not guess.
Does it read the PDF itself?
No, it works on text you have already copied. Reading the file would be the better way to do this, because the coordinates of every line are in there and the paragraphs could be worked out properly from them, but by the time text has been through the clipboard those coordinates are gone. What is left is line breaks and line lengths, and this makes the best of them.
Does my text leave the browser?
No. The repairs run as JavaScript in this tab. There is no request in this page that could carry your text anywhere, nothing is stored, and the page works with your network off. Text copied out of a statement, a contract or a medical letter stays where you pasted it.

More tools