Back

List comparison

Text tools

Loading

Loading tool

The tool is loaded only when you open it.

All processing for this tool happens in your browser. Your input is not sent to a server.

About this tool

Compare two lists of identifiers, labels or task names without treating their contents as numbers. Each physical line is one item. The tool reports the intersection, A minus B, B minus A, union and symmetric difference. Comparison keys can optionally trim outer whitespace or ignore case using JavaScript lowercase rules. The displayed record keeps the first original spelling and source line, so a matching key does not silently rewrite the exported value. All processing stays in your browser; a match is a string comparison and does not establish a person’s identity, account ownership or real-world equivalence.

Common uses

  • Compare two inventory-code snapshots while preserving leading zeros and long numeric identifiers.
  • Find tags present in both lists, items added or removed, or a combined deduplicated list in a predictable source order.
  • Review repeated labels, outer whitespace and case differences before preparing a local text export.

How to use it

  1. 1.Paste one item per line in A and B, or choose strict UTF-8 files. CRLF, LF and CR are separators; other Unicode separators remain inside items. A leading BOM is kept as U+FEFF text.
  2. 2.Choose whether to trim keys or compare case-sensitively, then compare. Defaults are no trim and case-sensitive. Inspect each side’s split-line, ignored-empty, nonempty, unique and duplicate counts.
  3. 3.Select one of the five sets and inspect the ASCII-escaped originals, keys, first source lines and occurrence counts. Review the export warning before copying or downloading. TXT exports the selected set; JSON contains both complete inputs and all five sets.

Executable examples: ASCII JSON requests and expected original items

All five sets

{"a":"red\nblue","b":"blue\ngreen","options":{"trim":false,"caseSensitive":true}}
{"intersection":["blue"],"aOnly":["red"],"bOnly":["green"],"union":["red","blue","green"],"symmetricDifference":["red","green"]}

The intersection uses A’s blue; union retains A order then appends B-only green.

First occurrence and duplicate counts

{"a":"b\na\nb\na","b":"a\na\nc","options":{"trim":false,"caseSensitive":true}}
{"intersection":["a"],"aOnly":["b"],"bOnly":["c"],"union":["b","a","c"],"symmetricDifference":["b","c"]}

Repeated keys produce one representative. The common a comes from A line 2; countA and countB are both 2.

Trim keys, preserve originals

{"a":"  Ada \nAda","b":"Ada\n Bob ","options":{"trim":true,"caseSensitive":true}}
{"intersection":["  Ada "],"aOnly":[],"bOnly":[" Bob "],"union":["  Ada "," Bob "],"symmetricDifference":[" Bob "]}

Both A spellings share the key Ada. The exported representative still contains its original outer spaces.

Default case-sensitive keys

{"a":"Ada","b":"ada","options":{"trim":false,"caseSensitive":true}}
{"intersection":[],"aOnly":["Ada"],"bOnly":["ada"],"union":["Ada","ada"],"symmetricDifference":["Ada","ada"]}

Ada and ada remain different when case-sensitive comparison is enabled.

Locale-independent lowercase

{"a":"Ada\nADA","b":"ada\nBOB","options":{"trim":false,"caseSensitive":false}}
{"intersection":["Ada"],"aOnly":[],"bOnly":["BOB"],"union":["Ada","BOB"],"symmetricDifference":["BOB"]}

Ada and ADA share one lowercase key; A’s first spelling represents the common item.

Sharp s is not ss

{"a":"\u00df\nSS","b":"ss","options":{"trim":false,"caseSensitive":false}}
{"intersection":["SS"],"aOnly":["\u00df"],"bOnly":[],"union":["\u00df","SS"],"symmetricDifference":["\u00df"]}

Lowercasing SS yields ss; lowercasing ß does not produce ss.

Greek sigma uses context

{"a":"\u03a3\n\u039f\u03a3","b":"\u03c3\n\u03bf\u03c3","options":{"trim":false,"caseSensitive":false}}
{"intersection":["\u03a3"],"aOnly":["\u039f\u03a3"],"bOnly":["\u03bf\u03c3"],"union":["\u03a3","\u039f\u03a3","\u03bf\u03c3"],"symmetricDifference":["\u039f\u03a3","\u03bf\u03c3"]}

Σ becomes σ, while the final letter of ΟΣ becomes ς. It differs from the explicitly written οσ.

Turkish I without locale tailoring

{"a":"I\n\u0130\n\u0131","b":"i\ni\u0307","options":{"trim":false,"caseSensitive":false}}
{"intersection":["I","\u0130"],"aOnly":["\u0131"],"bOnly":[],"union":["I","\u0130","\u0131"],"symmetricDifference":["\u0131"]}

I becomes i and İ becomes i plus combining dot. Dotless ı stays distinct.

No Unicode normalization

{"a":"\u00e9","b":"e\u0301","options":{"trim":false,"caseSensitive":true}}
{"intersection":[],"aOnly":["\u00e9"],"bOnly":["e\u0301"],"union":["\u00e9","e\u0301"],"symmetricDifference":["\u00e9","e\u0301"]}

Precomposed é and e followed by a combining accent stay separate despite similar appearance.

Identifiers remain exact strings

{"a":"001\n1\n9007199254740993\n__proto__","b":"1\n9007199254740992\nconstructor","options":{"trim":false,"caseSensitive":true}}
{"intersection":["1"],"aOnly":["001","9007199254740993","__proto__"],"bOnly":["9007199254740992","constructor"],"union":["001","1","9007199254740993","__proto__","9007199254740992","constructor"],"symmetricDifference":["001","9007199254740993","__proto__","9007199254740992","constructor"]}

Leading zeros, integers beyond safe Number precision and prototype-shaped names remain unchanged data.

Whitespace is data without trim

{"a":"\n \n\t\n","b":"","options":{"trim":false,"caseSensitive":true}}
{"intersection":[],"aOnly":[" ","\t"],"bOnly":[],"union":[" ","\t"],"symmetricDifference":[" ","\t"]}

Only empty keys are ignored. A space and a tab remain two unique items when trim is off.

Trimmed empty keys are ignored

{"a":"\n \n\t\n","b":"","options":{"trim":true,"caseSensitive":true}}
{"intersection":[],"aOnly":[],"bOnly":[],"union":[],"symmetricDifference":[]}

All A lines become empty keys. The trailing LF still contributes one physical split line.

CRLF, CR and LF delimiters

{"a":"a\r\nb\rc\n","b":"c\na","options":{"trim":false,"caseSensitive":true}}
{"intersection":["a","c"],"aOnly":["b"],"bOnly":[],"union":["a","b","c"],"symmetricDifference":["b"]}

A has four split lines, including its trailing empty line. Matches preserve A order and one-based source positions.

BOM differs from Unicode White_Space

{"a":"\ufeffx\n\u0085x","b":"x","options":{"trim":true,"caseSensitive":true}}
{"intersection":["\ufeffx"],"aOnly":["\u0085x"],"bOnly":[],"union":["\ufeffx","\u0085x"],"symmetricDifference":["\u0085x"]}

String.trim removes the outer U+FEFF but keeps U+0085 NEL, which is not ECMAScript trim whitespace.

NBSP only at the edges

{"a":"\u00a0x\u00a0\nx\u00a0y","b":"x\nx y","options":{"trim":true,"caseSensitive":true}}
{"intersection":["\u00a0x\u00a0"],"aOnly":["x\u00a0y"],"bOnly":["x y"],"union":["\u00a0x\u00a0","x\u00a0y","x y"],"symmetricDifference":["x\u00a0y","x y"]}

Trim removes surrounding NBSP from the key but leaves an internal NBSP distinct from a regular space.

U+2028 is not a list delimiter

{"a":"a\u2028b","b":"a\nb","options":{"trim":false,"caseSensitive":true}}
{"intersection":[],"aOnly":["a\u2028b"],"bOnly":["a","b"],"union":["a\u2028b","a","b"],"symmetricDifference":["a\u2028b","a","b"]}

Only CRLF, LF and CR split lists. The embedded Unicode line separator remains inside one item.

Unsafe text stays literal

{"a":"<img src=x onerror=alert(1)>\n=1+1\n\u202eabc\n\u0000","b":"=1+1","options":{"trim":false,"caseSensitive":true}}
{"intersection":["=1+1"],"aOnly":["<img src=x onerror=alert(1)>","\u202eabc","\u0000"],"bOnly":[],"union":["<img src=x onerror=alert(1)>","=1+1","\u202eabc","\u0000"],"symmetricDifference":["<img src=x onerror=alert(1)>","\u202eabc","\u0000"]}

HTML-like text, formulas, bidi control and NUL are data. ASCII previews reveal them; raw TXT still carries them unchanged.

Unpaired surrogate is lossless in JSON

{"a":"\ud800","b":"\ud800\nx","options":{"trim":false,"caseSensitive":true}}
{"intersection":["\ud800"],"aOnly":[],"bOnly":["x"],"union":["\ud800","x"],"symmetricDifference":["x"]}

The common high surrogate remains an exact UTF-16 unit. TXT for this set is blocked, while escaped JSON preserves it.

Common mistakes

  • Treating duplicate occurrences as repeated set members instead of one key with occurrence counts.
  • Assuming case-insensitive means Unicode case folding, Turkish-specific casing or accent removal.
  • Confusing escaped display text with the original values carried by TXT or JSON.
  • Expecting TXT to preserve original line separators, duplicate rows, ignored lines or source metadata.
  • Sharing the JSON report thinking it contains only the selected result, when it also includes both complete input lists.

Limits and notes

  • Exact trim set: U+0009, U+000A, U+000B, U+000C, U+000D, U+0020, U+00A0, U+1680, U+2000–U+200A, U+2028, U+2029, U+202F, U+205F, U+3000, U+FEFF. This is ECMAScript String.trim whitespace, not the Unicode White_Space property: U+FEFF is included, U+0085 and U+200B are excluded. CR and LF are consumed as list separators before trimming. Only outer whitespace affects the key; original spelling is retained. Case-insensitive mode uses the host’s locale-independent ECMAScript String.prototype.toLowerCase, not Unicode case folding or locale-aware collation. Greek Σ can become σ or final ς by context; ß does not equal ss. Turkish İ lowercases to i plus combining dot. No NFC, NFD, NFKC, NFKD, accent removal, width conversion, fuzzy matching or identity inference is performed.
  • Each side is limited to 131,072 UTF-16 units and 10,000 physical split lines; an empty input counts as one ignored line, and a trailing separator adds another ignored line. Files must be strict UTF-8 and at most 524,288 bytes; invalid bytes are rejected, not replaced. The file BOM is preserved, and only the trim option removes U+FEFF from a comparison key. Browser paste or clipboard behavior may alter line endings before or after this tool. The result is bounded to 20,000 unique entries and 60,000 references across the five sets. The interface shows at most 300 rows and 256 UTF-16 units per displayed cell, with a notice for clipping. TXT is limited to 1,048,576 UTF-8 bytes; ASCII JSON to 8,000,000 bytes. An exceeded limit blocks the export rather than silently shortening it. Changing input or comparison options invalidates previous results.
  • Empty comparison keys are ignored. Extra duplicates equal nonempty occurrences minus unique keys, separately for A and B. Each key keeps its first original and one-based source line. A represents shared keys; A-based results follow A order and B-only records follow B order. Union and symmetric difference append B-only entries after A entries. No numeric conversion, sorting or duplicate frequency expansion is applied.
  • TXT joins selected original items with LF and adds no final LF. It preserves raw controls, bidi characters, HTML-like text and spreadsheet formula prefixes, with no CSV or formula protection. Unpaired UTF-16 surrogates block raw TXT copying and downloading to prevent replacement loss. Review the destination; a safe ASCII preview does not make raw text safe for another application.
  • JSON losslessly preserves both full original inputs, including ignored rows, duplicate spellings and original CRLF/CR/LF separators, together with options, statistics and all five sets. It escapes non-ASCII, controls and unpaired surrogates. It contains more than the selected set and is not anonymization or encryption. Inputs are not uploaded, remotely queried or persisted by this tool; copying or saving is an explicit user action.

Frequently asked questions

How is list comparison different from a line-by-line diff?

Sets compare key membership and collapse repeated keys. They do not align corresponding rows or show edits between neighboring lines. Use Text diff when line position and changed passages matter.

Does trimming or ignoring case change my output?

These options change matching keys. The original representative remains the first occurrence on its source side, including spaces and case. A supplies shared representatives, so swapping A and B can change their spelling and output order.

Why do visually identical strings fail to match?

Different Unicode sequences, non-breaking spaces, zero-width characters or case rules can give different keys. There is no normalization or human-identity inference. Inspect the ASCII key preview before deciding whether a separate cleanup is appropriate.

Can I safely open the raw export in a spreadsheet?

TXT has no formula protection and is not CSV. A spreadsheet may interpret prefixes such as =, +, - or @ and may convert numeric-looking identifiers. JSON preserves string values and metadata, but includes both complete source inputs. Review the destination and content before using either export.

Related tools