Skip to content
ToolShelf

Text cleaner

Collapse repeated spaces, trim the ends of lines, drop blank lines and remove duplicates. Each one is a switch you turn on, and they run in a stated order.

Your text

Your text stays in this tab. The work is done by JavaScript in your browser. None of it is uploaded, logged or saved, and the tool keeps working with the network off.

What to clean

Cleaning steps

Applied in the order listed: collapse, then trim, then blank lines, then duplicates. That order matters, a line of spaces only counts as blank once it has been trimmed.

Cleaned

The cleaned text appears here.

Nothing happens until you ask

Every switch starts off, and with none of them on the output is character-for-character what you pasted. A cleaner that quietly rewrites text the moment it arrives is one you cannot hand anything you care about, so the untouched state is genuinely untouched.

The order is fixed, and it matters

Collapse repeated spaces, then trim the ends of lines, then remove blank lines, then remove duplicates. Applied in the order the switches were flicked, the results would differ, and not in a way anyone could predict.

The clearest case: a line containing three spaces is not blank. It only becomes blank once it has been trimmed, so removing blank lines before trimming would leave it standing. The same goes for duplicates, two lines that differ only in trailing whitespace are different strings until they have been trimmed, which is why a deduplication that “does not work” almost always starts working when trimming is turned on too.

What each step does

  • Collapse repeated spaces turns runs of spaces and tabs into one space, within each line. Line breaks are untouched, so paragraphs survive.
  • Trim the ends of lines removes leading and trailing whitespace from every line, leaving the middle alone.
  • Remove blank lines drops empty lines and whitespace-only ones.
  • Remove duplicate lines keeps the first of each repeated line and preserves the original order. It is case-sensitive, because case usually carries meaning in the lists people paste in.

Invisible differences

Windows line endings are normalised before anything else runs. Text copied from a Windows editor carries a carriage return before every newline, which makes two identical-looking lines compare as different and makes a deduplication look broken for no visible reason.

Where the mess comes from

Text out of a PDF, which arrives with line breaks in the middle of sentences and spacing that was really letter-spacing. Lists out of a spreadsheet or an email thread, which arrive with blank lines and repeats. Anything copied out of a web page, which brings the source’s indentation with it. All of it runs in your browser. Nothing is uploaded.

Questions

Is my text sent anywhere?
No. Every step runs in your browser and nothing is uploaded, logged or stored. The page works with the network off.
Why is nothing switched on to begin with?
Because a cleaner that rewrites what you pasted before you asked it to is one you cannot trust with something that matters. Each step is a deliberate choice, and with none of them on the output is exactly what you put in.
In what order are the steps applied?
Collapse repeated spaces, trim the ends of lines, remove blank lines, then remove duplicates. The order matters: a line containing only spaces does not count as blank until it has been trimmed, and two lines differing only in trailing whitespace are not duplicates until then either.
Does collapsing spaces join my lines together?
No. It collapses runs of spaces and tabs within each line and leaves the line breaks alone. Paragraph structure survives; only the horizontal spacing changes.
Is duplicate removal case-sensitive?
Yes. 'Apple' and 'apple' are kept as two different lines. Case often carries meaning in the lists people paste in, names, codes, identifiers, so treating them as the same would throw away real data.
Which duplicate is kept?
The first. The rest are removed and the remaining lines stay in their original order, so a list keeps whatever ordering it arrived with rather than being sorted.
Why did my duplicate lines not get removed?
Usually invisible differences. Trailing spaces make two identical-looking lines different, which is why turning on 'trim the ends of lines' at the same time normally fixes it. Windows line endings are normalised automatically, so that particular invisible difference is already handled.
What is this useful for?
Text pasted out of a PDF, which arrives with broken spacing and stray line breaks; lists copied from a spreadsheet or an email, which arrive with blank lines and duplicates; and anything copied from a web page, which brings its indentation with it.

More tools