CSV split and merge
Data conversion
Loading
Loading tool
The tool is loaded only when you open it.
All processing for this tool happens in your browser. Your input is not sent to a server.
About this tool
Divide an export into smaller CSV files by logical data-row count or exact column value, or append several local CSV files into one table. Quoted newlines stay inside their original fields. Merge follows the visible file order and preserves every data row, including duplicates. Choose strict ordered headers, align identical names to the first file, or form a first-seen union of columns. This is row concatenation, not a key-based join. Parsing, planning and CSV/ZIP creation run in a local browser worker without input uploads or saved input history.
Common uses
- Split a delivery or import file into batches with a fixed maximum number of data records, repeating the real header in each part.
- Create one file per exact department, account or category while keeping group and row order, string identifiers and multiline notes.
- Append monthly exports with identical columns, deliberately align changed column order, or review an explicit union when newer files add columns.
How to use it
- 1.Choose Split for one pasted source or local file, or Merge for two to twenty local files. Select whether every source has a header. Set comma, semicolon, tab or pipe explicitly; file encoding is UTF-8, UTF-16LE or UTF-16BE, with per-file choices in merge mode.
- 2.For Split, select a positive whole-number data-row count or inspect the input and select an exact grouping column. For Merge, review the visible file order, move files up or down as needed, and select strict, align or union. Headerless files require equal widths and are appended by position.
- 3.Run the operation and review counts, per-file errors, output names and a paged preview. Any invalid source blocks the entire merge. Repair or explicitly remove that source and rerun; files are never silently skipped. An input or option change clears stale results. Keep formula protection on when reviewing in a spreadsheet, or explicitly choose raw output after reading the warning. Split downloads a ZIP with ordinal part names. Merge downloads merged.csv or a ZIP containing it. Preview clipping does not truncate a successful bounded export.
Executable split and merge examples
Count CSV records, including quoted newlines
{
"mode": "split",
"header": true,
"safe": true,
"sources": [
{
"text": "id,note\n001,\"first\nsecond\"\n002,\"comma, quote \"\"ok\"\"\"\n003,猫",
"delimiter": ",",
"encoding": "utf-8"
}
],
"splitBy": "rows",
"rowsPerPart": 2
}{
"files": [
{
"name": "part-0001.csv",
"records": [
[
"id",
"note"
],
[
"001",
"first\nsecond"
],
[
"002",
"comma, quote \"ok\""
]
]
},
{
"name": "part-0002.csv",
"records": [
[
"id",
"note"
],
[
"003",
"猫"
]
]
}
]
}The first part has two data records even though one field contains a newline. Each part repeats the header. Leading zeros, the comma, escaped quote and Unicode survive CSV reserialization.
Exact groups use first-seen order
{
"mode": "split",
"header": true,
"safe": true,
"sources": [
{
"text": "group,id\nA,001\n,002\n ,003\na,004\nA,005\n001,006\n1,007",
"delimiter": ",",
"encoding": "utf-8"
}
],
"splitBy": "value",
"column": 0
}{
"files": [
{
"name": "part-0001.csv",
"records": [
[
"group",
"id"
],
[
"A",
"001"
],
[
"A",
"005"
]
]
},
{
"name": "part-0002.csv",
"records": [
[
"group",
"id"
],
[
"",
"002"
]
]
},
{
"name": "part-0003.csv",
"records": [
[
"group",
"id"
],
[
" ",
"003"
]
]
},
{
"name": "part-0004.csv",
"records": [
[
"group",
"id"
],
[
"a",
"004"
]
]
},
{
"name": "part-0005.csv",
"records": [
[
"group",
"id"
],
[
"001",
"006"
]
]
},
{
"name": "part-0006.csv",
"records": [
[
"group",
"id"
],
[
"1",
"007"
]
]
}
]
}A and a, an empty string and one space, and 001 and 1 are six distinct groups. Rows within A keep their source order. Only ordinal names enter the ZIP; private group values never become filenames.
Strict concatenation keeps duplicate and header-like rows
{
"mode": "merge",
"header": true,
"safe": true,
"sources": [
{
"text": "id,value\n001,2\n001,2",
"delimiter": ",",
"encoding": "utf-8"
},
{
"text": "id,value\nid,value\n002,3",
"delimiter": ",",
"encoding": "utf-8"
}
],
"schema": "strict"
}{
"files": [
{
"name": "merged.csv",
"records": [
[
"id",
"value"
],
[
"001",
"2"
],
[
"001",
"2"
],
[
"id",
"value"
],
[
"002",
"3"
]
]
}
]
}Identical ordered headers allow appending. Duplicate data is retained, and the id,value record inside the second file remains data. Only each file’s actual first header record is excluded from the appended data.
Strict mode rejects reordered headers
{
"mode": "merge",
"header": true,
"safe": true,
"sources": [
{
"text": "id,city\n001,Paris",
"delimiter": ",",
"encoding": "utf-8"
},
{
"text": "city,id\nTokyo,002",
"delimiter": ",",
"encoding": "utf-8"
}
],
"schema": "strict"
}{
"error": "schemaStrict"
}The second file has the same names in a different order. Strict mode blocks the whole merge; explicitly choose align-by-name when that reordering is intended.
Align exact names to the first file
{
"mode": "merge",
"header": true,
"safe": true,
"sources": [
{
"text": "id,city\n001,Paris",
"delimiter": ",",
"encoding": "utf-8"
},
{
"text": "city,id\nTokyo,002",
"delimiter": ",",
"encoding": "utf-8"
}
],
"schema": "align"
}{
"files": [
{
"name": "merged.csv",
"records": [
[
"id",
"city"
],
[
"001",
"Paris"
],
[
"002",
"Tokyo"
]
]
}
]
}Align accepts the identical header set and reorders the second row to id,city. It does not trim names, ignore case or infer that differently named columns have the same meaning.
Union adds columns and loses missing-versus-empty identity
{
"mode": "merge",
"header": true,
"safe": true,
"sources": [
{
"text": "id,name\n001,Ada\n002,",
"delimiter": ",",
"encoding": "utf-8"
},
{
"text": "id,city\n003,Tokyo",
"delimiter": ",",
"encoding": "utf-8"
}
],
"schema": "union"
}{
"files": [
{
"name": "merged.csv",
"records": [
[
"id",
"name",
"city"
],
[
"001",
"Ada",
""
],
[
"002",
"",
""
],
[
"003",
"",
"Tokyo"
]
]
}
]
}The new city column follows id,name. Absent source columns become empty strings. The empty name in row 002 and the absent name in row 003 are indistinguishable in the output; keep the source files when that distinction matters.
Headerless files use positions and may use different delimiters
{
"mode": "merge",
"header": false,
"safe": true,
"sources": [
{
"text": "001;Paris\n002;Tokyo",
"delimiter": ";",
"encoding": "utf-8"
},
{
"text": "003\tOsaka",
"delimiter": "\t",
"encoding": "utf-8"
}
],
"schema": "strict"
}{
"files": [
{
"name": "merged.csv",
"records": [
[
"001",
"Paris"
],
[
"002",
"Tokyo"
],
[
"003",
"Osaka"
]
]
}
]
}Both sources have two fields per record. Output contains three data records and no synthetic header. File-specific delimiter choices affect reading; output CSV always uses comma delimiters and UTF-8.
A header-only source has one row-count part
{
"mode": "split",
"header": true,
"safe": true,
"sources": [
{
"text": "id,city",
"delimiter": ",",
"encoding": "utf-8"
}
],
"splitBy": "rows",
"rowsPerPart": 1
}{
"files": [
{
"name": "part-0001.csv",
"records": [
[
"id",
"city"
]
]
}
]
}There are zero data records, so row-count splitting returns one header-only part. The header is not counted as a data row.
A header-only source has no value groups
{
"mode": "split",
"header": true,
"safe": true,
"sources": [
{
"text": "id,city",
"delimiter": ",",
"encoding": "utf-8"
}
],
"splitBy": "value",
"column": 1
}{
"files": []
}No data record has a grouping value. Value splitting reports zero groups and does not offer an empty ZIP. A zero-byte input is instead an error.
Protection changes risky cells in every part
{
"mode": "split",
"header": true,
"safe": true,
"sources": [
{
"text": "=header,value\n=1+1,-42\n@SUM(A1),plain",
"delimiter": ",",
"encoding": "utf-8"
}
],
"splitBy": "rows",
"rowsPerPart": 1
}{
"files": [
{
"name": "part-0001.csv",
"records": [
[
"'=header",
"value"
],
[
"'=1+1",
"'-42"
]
]
},
{
"name": "part-0002.csv",
"records": [
[
"'=header",
"value"
],
[
"'@SUM(A1)",
"plain"
]
]
}
]
}Default protection prefixes risky headers and data with an apostrophe, including a fullwidth sign and a negative number. The protected header appears in both files. This intentionally changes exported values; CSV quoting alone is not formula protection.
Raw export is an explicit current-input choice
{
"mode": "split",
"header": true,
"safe": false,
"sources": [
{
"text": "=header,value\n=1+1,-42",
"delimiter": ",",
"encoding": "utf-8"
}
],
"splitBy": "rows",
"rowsPerPart": 1
}{
"files": [
{
"name": "part-0001.csv",
"records": [
[
"=header",
"value"
],
[
"=1+1",
"-42"
]
]
}
]
}Disabling protection keeps parsed string values exactly. Only do this for a trusted destination: spreadsheet programs may execute formulas or reinterpret identifiers. Changing input or transformation settings resets the protection choice.
One malformed file blocks the whole merge
{
"mode": "merge",
"header": true,
"safe": true,
"sources": [
{
"text": "id,city\n001,Paris",
"delimiter": ",",
"encoding": "utf-8"
},
{
"text": "id,city\n002",
"delimiter": ",",
"encoding": "utf-8"
}
],
"schema": "strict"
}{
"error": "width"
}The second source lacks its city field. All schema policies reject ragged source rows; union fills absent columns across valid schemas, not missing fields inside malformed records. No partial merged download is produced.
Common split and merge mistakes
- Counting physical newlines instead of logical CSV records, especially with multiline notes.
- Selecting align for files with missing or added column names; only union permits that schema difference.
- Assuming union can distinguish a missing source column from an existing empty string after export.
- Mixing header and headerless files under one batch header setting, or expecting a short record to be padded automatically.
- Expecting concatenation to remove duplicates or enrich rows by a shared ID.
- Choosing the wrong delimiter or encoding and expecting automatic correction.
- Opening untrusted raw output in a spreadsheet or assuming ordinary CSV quotes prevent formula execution.
- Sharing an unencrypted ZIP without reviewing the private data inside it.
Limits and notes
- Each source is limited to 2 MiB of raw file bytes and 2 MiB of decoded UTF-8 text, 10,000 data rows, 128 columns and 200,000 data cells. A merge accepts 2–20 files; the batch is limited to 8 MiB of raw bytes, 8 MiB of decoded UTF-8, 40,000 data rows and 500,000 input cells. Output is limited to 40,000 data rows, 128 columns and 500,000 data cells; split allows at most 100 parts or groups. Exceeding a limit fails the operation rather than truncating it.
- The combined serialized CSV output is capped at 12 MiB and ZIP at 13 MiB. Repeated headers, quoting, UTF-8 characters and formula prefixes count toward output limits. ZIP entries are stored without compression, so ZIP is a packaging format here, not a promise of smaller downloads. The worker has a 10-second operation deadline and can be cancelled; allocation checks are not a guarantee of a fixed browser heap size.
- Split counts logical data records, never physical text lines. Header-only input yields one header-only part in row-count mode and zero groups in value mode; zero groups have no ZIP download. Zero-byte input is an error. Group identity is exact: empty strings, spaces, case variants, Unicode variants and 001 versus 1 remain distinct. Groups follow first occurrence and rows retain their within-group order. Group labels and original file names are never used as archive entry names: entries are flat part-0001.csv names, or merged.csv.
- Strict merge requires identical headers in identical order. Align requires identical exact name sets and adopts the first file’s column order. Union appends new names in first-seen file/header order and fills absent columns with empty strings; this loses the distinction between an absent column and an existing empty cell. Every source must have nonblank, unique headers and equal-width records under all policies. Headerless merge supports equal-width positional concatenation only and emits no synthetic header. Duplicate rows and data rows equal to a header are retained.
- Choose the file encoding explicitly; malformed bytes, conflicting BOMs and UTF-32 are rejected. There is no legacy-encoding detection or automatic delimiter guess. Output is UTF-8 comma CSV with CRLF record separators; quoted field newlines remain unchanged. Raw export preserves parsed string values, not original bytes, quoting or record separators. Formula protection prefixes risky headers and data with apostrophes, so protected values differ from the originals. It is not a universal spreadsheet-safety guarantee; spreadsheets may still coerce IDs or dates. Raw CSV can activate formulas. Changing inputs or transformation settings restores protection. Preview is limited to 25 rows per page and 12 columns, with long cells and labels shortened on screen. Downloads keep the complete successful result within the stated limits. Cancel, Reset, new inputs, file reordering or navigation discard pending results and exports. No input is uploaded or persisted by this tool, but downloads contain the processed data, may expose private values and remain on your device after clearing the page.
Frequently asked questions
Is splitting by 1,000 rows the same as splitting every 1,000 lines?
No. A quoted CSV field can contain several line breaks while remaining part of one data record. The tool parses CSV records first, excludes the real header from the count, and repeats that header in every row-count part. A headerless input gets no added header.
Which merge policy should I choose, and what happens to an invalid file?
Use strict when every file should have the same names in the same order. Use align when only the order differs and all exact column names match. Use union only when added or absent columns are intentional: new columns are appended, absent values become empty strings, and missing-versus-empty information is lost. None of these policies repairs a malformed source row. The tool reports bounded per-file diagnostics and blocks the whole merge and its downloads. Fix the invalid source or deliberately remove it from the visible list, then rerun. It never silently creates a partial-success CSV or ZIP.
Will merging remove duplicates or match records by an ID?
No. Merge appends all data records in the displayed file order. It does not deduplicate, sort, trim, coerce values or match keys. Use CSV join for relational enrichment, or explicitly clean a separate copy if duplicate removal is needed.
How should I choose formula protection and share the output?
Formula protection is enabled by default and changes cells that may start spreadsheet formulas, including after certain whitespace or controls and with supported fullwidth signs. It also applies to every repeated header. Explicit raw export keeps parsed strings unchanged, but can execute formulas when opened in a spreadsheet. Neither setting prevents every spreadsheet type-conversion behavior. Split uses deterministic flat names such as part-0001.csv and merge uses merged.csv. Group values are visible in the local manifest and CSV contents, but are not used to name archive entries. The ZIP is not encrypted, so review the file contents before sharing it.