Back

Newline and UTF-8 BOM converter

Text tools

Loading

Loading tool

The tool is loaded only when you open it.

All processing for this tool happens in your browser. Your input is not sent to a server.

About this tool

A file can contain several line-ending conventions even when its editor shows ordinary lines. Inspect LF, CRLF and lone CR separately, then select preserve, LF, CRLF or CR without trimming text or inventing a final newline. This tool keeps the complete original alongside a bounded escaped preview and a separate converted result. The UTF-8 byte signature is tracked separately from U+FEFF text so removing a file BOM does not silently delete meaningful text characters. Choose the output contract required by your receiving system rather than assuming that one convention is always correct.

Common uses

  • Prepare a UTF-8 configuration fixture with CRLF for a consumer that explicitly requires Windows-style separators.
  • Diagnose a mixed-line-ending text export and compare source and output byte counts before replacing a file.
  • Remove a known UTF-8 file signature while retaining an additional leading U+FEFF text character and every internal U+FEFF.

How to use it

  1. 1.Paste literal text or open a local UTF-8 file. Strict decoding rejects invalid UTF-8 and UTF-16/UTF-32 signatures. A BOM-less file is only interpreted as UTF-8; other encodings are never guessed. The first EF BB BF file prefix becomes signature metadata, while subsequent U+FEFF remains text.
  2. 2.Choose the line-ending and BOM policies, then analyze. Review original and output separator counts, mixed status, logical lines, UTF-16 units, code points and UTF-8 bytes. Escaped previews show at most 2,000 source units and 12,000 rendered characters; clipping never shortens the complete stored result.
  3. 3.Explicitly review the output before copying or downloading. A downloaded UTF-8 file gives exact bytes and the selected signature. The ASCII JSON report retains the complete source and output; treat it as containing the original private data. Editing input or policy invalidates the old review. Copy includes text only, excluding the separately tracked signature. Clipboard and receiving apps can normalize line endings. Download UTF-8 for exact bytes.

Exact synthetic conversion examples

Mixed separators to LF

request (ASCII JSON): {"input":"A\r\nB\nC\rD","source":"paste","bom":false,"newline":"lf","bomPolicy":"preserve"}
{"text":"A\nB\nC\nD","bom":false,"crlf":0,"lf":3,"cr":0,"logicalLines":4,"utf16Units":7,"codePoints":7,"utf8Bytes":7}

Three separator styles become three LF tokens; no trailing newline is invented.

LF to CRLF with a signature

request (ASCII JSON): {"input":"a\nb","source":"paste","bom":false,"newline":"crlf","bomPolicy":"add"}
{"text":"a\r\nb","bom":true,"crlf":1,"lf":0,"cr":0,"logicalLines":2,"utf16Units":4,"codePoints":4,"utf8Bytes":7}

The CRLF body is four bytes; the separate signature adds three bytes.

Convert to legacy CR

request (ASCII JSON): {"input":"a\r\nb\n","source":"paste","bom":false,"newline":"cr","bomPolicy":"remove"}
{"text":"a\rb\r","bom":false,"crlf":0,"lf":0,"cr":2,"logicalLines":3,"utf16Units":4,"codePoints":4,"utf8Bytes":4}

A CRLF token and LF token each become one CR. The existing final separator remains.

Preserve a mixed file body

request (ASCII JSON): {"input":"a\r\nb\rc\n","source":"paste","bom":false,"newline":"preserve","bomPolicy":"preserve"}
{"text":"a\r\nb\rc\n","bom":false,"crlf":1,"lf":1,"cr":1,"logicalLines":4,"utf16Units":7,"codePoints":7,"utf8Bytes":7}

Preserve retains every token, including the trailing LF and final empty logical line.

Remove only the file signature

request (ASCII JSON): {"input":"\ufeffA\r\n","source":"file","bom":true,"newline":"lf","bomPolicy":"remove"}
{"text":"\ufeffA\n","bom":false,"crlf":0,"lf":1,"cr":0,"logicalLines":2,"utf16Units":3,"codePoints":3,"utf8Bytes":5}

The file begins with a signature plus a textual U+FEFF. Removing the signature retains the text U+FEFF.

Keep pasted U+FEFF

request (ASCII JSON): {"input":"\ufeffA","source":"paste","bom":false,"newline":"preserve","bomPolicy":"remove"}
{"text":"\ufeffA","bom":false,"crlf":0,"lf":0,"cr":0,"logicalLines":1,"utf16Units":2,"codePoints":2,"utf8Bytes":4}

The initial pasted U+FEFF is text, so remove does not delete it.

Keep Unicode separators as text

request (ASCII JSON): {"input":"a\u0085b\u2028c\u2029d\n","source":"paste","bom":false,"newline":"crlf","bomPolicy":"preserve"}
{"text":"a\u0085b\u2028c\u2029d\r\n","bom":false,"crlf":1,"lf":0,"cr":0,"logicalLines":2,"utf16Units":9,"codePoints":9,"utf8Bytes":14}

Only the final LF becomes CRLF. NEL, LS and PS remain untouched and do not increase logical lines.

Count a supplementary emoji

request (ASCII JSON): {"input":"\ud83d\ude00\r\n","source":"paste","bom":false,"newline":"lf","bomPolicy":"preserve"}
{"text":"\ud83d\ude00\n","bom":false,"crlf":0,"lf":1,"cr":0,"logicalLines":2,"utf16Units":3,"codePoints":2,"utf8Bytes":5}

The emoji occupies two UTF-16 units, one code point and four UTF-8 bytes.

Add a signature to empty text

request (ASCII JSON): {"input":"","source":"paste","bom":false,"newline":"lf","bomPolicy":"add"}
{"text":"","bom":true,"crlf":0,"lf":0,"cr":0,"logicalLines":0,"utf16Units":0,"codePoints":0,"utf8Bytes":3}

The body stays empty with zero logical lines; only the three-byte signature is emitted.

Retain a lone surrogate in JSON

request (ASCII JSON): {"input":"A\ud800\r\n","source":"paste","bom":false,"newline":"lf","bomPolicy":"preserve"}
{"text":"A\ud800\n","bom":false,"crlf":0,"lf":1,"cr":0,"logicalLines":2,"utf16Units":3,"codePoints":3,"utf8Bytes":null}

UTF-8 export is blocked, but ASCII JSON preserves the exact U+D800 code unit.

Keep escape spellings literal

request (ASCII JSON): {"input":"a\\r\\nb","source":"paste","bom":false,"newline":"lf","bomPolicy":"preserve"}
{"text":"a\\r\\nb","bom":false,"crlf":0,"lf":0,"cr":0,"logicalLines":1,"utf16Units":6,"codePoints":6,"utf8Bytes":6}

Backslash escape spellings are ordinary ASCII; the tool never evaluates them.

Common newline and BOM mistakes

  • Counting a CRLF pair as two line breaks.
  • Treating every leading U+FEFF as removable metadata, especially after paste.
  • Assuming a visual preview proves exact byte identity or that copying preserves line endings.
  • Assuming equal byte counts mean equal text.
  • Expecting NEL, LS or PS to become LF automatically.
  • Sharing the ASCII report as though it were redacted.

Limits and notes

  • Only CRLF, LF and CR are converted. CRLF is one separator, not two. NEL (U+0085), LS (U+2028), PS (U+2029), spaces, tabs and all other text are retained. Empty text has 0 logical lines; nonempty text has separator count + 1, including the empty final line after a trailing separator. No final newline is added or removed.
  • File BOM policy means preserve, add or remove one three-byte UTF-8 signature. It never removes textual U+FEFF. Pasted U+FEFF is text; the original file encoding and signature cannot be recovered from paste. Adding a signature before a leading text U+FEFF intentionally produces two EF BB BF sequences. Browser paste and clipboard operations may normalize line endings outside this tool. The output text starts with U+FEFF. Even without a separately added signature its UTF-8 bytes start EF BB BF, which another decoder may consume as a signature. Removing those bytes would remove the retained text character.
  • UTF-8 bytes include the separately tracked file signature; UTF-16 units and code points describe text without that signature. Supplementary characters occupy two UTF-16 units but one code point. Code points are not grapheme clusters or display columns. A lone surrogate is retained diagnostically, counts as one malformed code point, and has no valid UTF-8 byte length. Raw export is blocked rather than replacing it with U+FFFD.
  • Processing limits are 100,000 input UTF-16 units, 400,003 input file bytes, 200,000 output units, 800,003 output bytes and 2,000,000 ASCII report characters. An exceeded limit is an explicit failure, never an incomplete download. Processing is local with no uploads, network lookup, persistence or execution of text. This is not a charset detector, Unicode normalizer, whitespace cleanup or security sanitizer.
  • The ASCII-only JSON report preserves every original and output UTF-16 unit, including lone surrogates. Escaping is reversible representation, not redaction, anonymization or encryption. The report contains complete original text and may contain secrets. UTF-16 and UTF-32 files require a separate explicit encoding conversion before import.

Frequently asked questions

Does removing BOM remove all U+FEFF characters?

No. Only the first EF BB BF file prefix is signature metadata. Later leading or internal U+FEFF is text and stays intact. All pasted U+FEFF is text, even at position zero. The output text starts with U+FEFF. Even without a separately added signature its UTF-8 bytes start EF BB BF, which another decoder may consume as a signature. Removing those bytes would remove the retained text character.

Can preserve reconstruct line endings changed during paste?

No. Preserve keeps the exact string received by the tool. A browser or clipboard may already have normalized it. Import the original UTF-8 file when exact input bytes matter.

Why does a file ending with one newline have an extra logical line?

The count describes segments separated by CRLF, LF or CR. A trailing separator creates a final empty segment. Empty text is the explicit zero-line special case.

Can I use this to convert arbitrary encodings?

No. Invalid UTF-8 and UTF-16/UTF-32 signatures are rejected. A BOM-less stream must decode strictly as UTF-8; successful decoding is not proof of the encoding intended by its producer.

Related tools