Back

Parquet viewer and converter

Data conversion

Loading

Loading tool

The tool is loaded only when you open it.

All processing for this tool happens in your browser. Your input is not sent to a server.

About this tool

Inspect a small local Parquet file without uploading its contents. The bundled hyparquet reader runs in a browser worker and reads supported flat columns with UNCOMPRESSED or SNAPPY compression. Review physical types, logical annotations, nullability and the representation used by this tool, then page through rows and export an inclusive row range. Exact 64-bit integers, decimals, binary values and temporal units use schema-described strings; ordinary booleans and supported finite numbers retain their types. Unsupported columns stay in the schema with a reason and an unsupported:null representation. Their row values are explicit null placeholders, not decoded source nulls; supported columns remain available. Check these warnings before interpreting or exporting the result. The tool does not fetch remote files, use a third-party CDN or persist your file contents.

Common uses

  • Inspect a bounded analytics export before passing it to another tool: check column names, nullable fields, row counts and the distinction between physical and logical types, including any unsupported-column warnings.
  • Extract a contiguous range of records for a spreadsheet review while keeping the original file unchanged and enabling formula protection for headers and text.
  • Download a schema-aware JSON sample that keeps null distinct from empty text and avoids rounding large identifiers, fixed-scale decimals or timestamp units.

How to use it

  1. 1.Choose one local Parquet file and load it. After validation, inspect the schema, row count and row groups. A .parquet extension alone does not establish a valid file. Review every unsupported-column warning: those columns remain labeled in the schema but contain null placeholders. Supported columns can still be inspected. Use a full local Parquet tool when you need the original values of unsupported columns; corrupt or oversized files fail rather than returning a partial load.
  2. 2.Use Previous and Next to inspect 25-row pages. The preview shows at most the first 12 columns and clips a cell after 500 characters; consult the schema for all columns and the null/type labels for exact interpretation. The page number is a viewing control and does not select exported rows.
  3. 3.Enter a start and end row using 1-based inclusive positions in file order, then choose CSV or JSON. For a zero-row file use the empty range 0–0 to export schema or headers only. Export contains every column and the full values in that valid range, even if the preview hides or clips them. Keep CSV formula protection enabled for spreadsheet review; raw CSV is an explicit choice. Cancel, Reset or leaving the tool discards the in-memory file and results, so select the file again to continue.

Executable examples from synthetic Parquet fixtures

Keep Unicode, null and primitive types

{
  "file": "mixed-snappy-v1.parquet",
  "format": "json",
  "start": 1,
  "end": 5,
  "safe": true,
  "excerptColumns": [
    "name",
    "active",
    "int32_value"
  ]
}
{
  "schemaExcerpt": [
    {
      "name": "name",
      "physicalType": "BYTE_ARRAY",
      "logicalType": "STRING",
      "representation": "text",
      "nullable": true,
      "supported": true
    },
    {
      "name": "active",
      "physicalType": "BOOLEAN",
      "logicalType": "NONE",
      "representation": "boolean",
      "nullable": true,
      "supported": true
    },
    {
      "name": "int32_value",
      "physicalType": "INT32",
      "logicalType": "NONE",
      "representation": "number",
      "nullable": true,
      "supported": true
    }
  ],
  "range": {
    "start": 1,
    "end": 5
  },
  "rowsExcerpt": [
    {
      "name": "東京 · café 😀",
      "active": true,
      "int32_value": -2147483648
    },
    {
      "name": "",
      "active": false,
      "int32_value": 2147483647
    },
    {
      "name": "\\N",
      "active": null,
      "int32_value": 0
    },
    {
      "name": "=SUM(1,2)\r\n\"quoted\",value",
      "active": true,
      "int32_value": -1
    },
    {
      "name": null,
      "active": false,
      "int32_value": null
    }
  ]
}

The synthetic file has 14 columns and 5 rows. This excerpt shows name, active and int32_value. Unicode and empty text survive; null differs from false and numeric zero. The full JSON export contains all 14 schema entries and columns.

Read a Snappy-compressed v2 data page

{
  "file": "mixed-snappy-v2.parquet",
  "format": "json",
  "start": 1,
  "end": 2,
  "safe": true,
  "excerptColumns": [
    "name",
    "int32_value"
  ]
}
{
  "schemaExcerpt": [
    {
      "name": "name",
      "physicalType": "BYTE_ARRAY",
      "logicalType": "STRING",
      "representation": "text",
      "nullable": true,
      "supported": true
    },
    {
      "name": "int32_value",
      "physicalType": "INT32",
      "logicalType": "NONE",
      "representation": "number",
      "nullable": true,
      "supported": true
    }
  ],
  "range": {
    "start": 1,
    "end": 2
  },
  "rowsExcerpt": [
    {
      "name": "東京 · café 😀",
      "int32_value": -2147483648
    },
    {
      "name": "",
      "int32_value": 2147483647
    }
  ]
}

This independently written file uses SNAPPY, dictionary values and data page v2. The first two records keep their original strings and signed INT32 values. excerptColumns only shortens this documentation; it is not an export column selector.

Export both ends of a row range

{
  "file": "mixed-uncompressed-v1.parquet",
  "format": "json",
  "start": 2,
  "end": 3,
  "safe": true,
  "excerptColumns": [
    "name",
    "int32_value"
  ]
}
{
  "schemaExcerpt": [
    {
      "name": "name",
      "physicalType": "BYTE_ARRAY",
      "logicalType": "STRING",
      "representation": "text",
      "nullable": true,
      "supported": true
    },
    {
      "name": "int32_value",
      "physicalType": "INT32",
      "logicalType": "NONE",
      "representation": "number",
      "nullable": true,
      "supported": true
    }
  ],
  "range": {
    "start": 2,
    "end": 3
  },
  "rowsExcerpt": [
    {
      "name": "",
      "int32_value": 2147483647
    },
    {
      "name": "\\N",
      "int32_value": 0
    }
  ]
}

The UNCOMPRESSED/PLAIN version contains the same records. start 2 and end 3 include the second and third data rows, in file order. An empty string and the literal backslash-N remain distinct in JSON. All columns are exported even though only two are shown here.

Retain signed and unsigned 64-bit integers

{
  "file": "mixed-snappy-v1.parquet",
  "format": "json",
  "start": 1,
  "end": 4,
  "safe": true,
  "excerptColumns": [
    "id",
    "uint64_value"
  ]
}
{
  "schemaExcerpt": [
    {
      "name": "id",
      "physicalType": "INT64",
      "logicalType": "NONE",
      "representation": "integer:string",
      "nullable": true,
      "supported": true
    },
    {
      "name": "uint64_value",
      "physicalType": "INT64",
      "logicalType": "UINT_64",
      "representation": "integer:string",
      "nullable": true,
      "supported": true
    }
  ],
  "range": {
    "start": 1,
    "end": 4
  },
  "rowsExcerpt": [
    {
      "id": "-9223372036854775808",
      "uint64_value": "18446744073709551615"
    },
    {
      "id": "9223372036854775807",
      "uint64_value": "9223372036854775808"
    },
    {
      "id": "9007199254740993",
      "uint64_value": "9007199254740993"
    },
    {
      "id": "-9007199254740993",
      "uint64_value": "0"
    }
  ]
}

INT64 and UINT64 are exact decimal strings, including zero. Values beyond 2^53 and the signed/unsigned 64-bit endpoints are not rounded. The schema identifies their physical/logical types and integer:string representation.

Preserve a decimal’s declared scale

{
  "file": "mixed-snappy-v1.parquet",
  "format": "json",
  "start": 1,
  "end": 5,
  "safe": true,
  "excerptColumns": [
    "amount"
  ]
}
{
  "schemaExcerpt": [
    {
      "name": "amount",
      "physicalType": "FIXED_LEN_BYTE_ARRAY",
      "logicalType": "DECIMAL(30,6)",
      "representation": "decimal:string",
      "nullable": true,
      "supported": true
    }
  ],
  "range": {
    "start": 1,
    "end": 5
  },
  "rowsExcerpt": [
    {
      "amount": "9007199254740993.123456"
    },
    {
      "amount": "-9007199254740993.000001"
    },
    {
      "amount": "0.000001"
    },
    {
      "amount": "0.000000"
    },
    {
      "amount": null
    }
  ]
}

DECIMAL(30,6) is fixed-point text. Both 0.000001 and 0.000000 keep six decimal places, large amounts retain every digit, and null remains null. No binary floating-point arithmetic is introduced.

Read temporal units before interpreting the value

{
  "file": "mixed-snappy-v1.parquet",
  "format": "json",
  "start": 1,
  "end": 4,
  "safe": true,
  "excerptColumns": [
    "timestamp_us",
    "timestamp_ns",
    "date",
    "time_us"
  ]
}
{
  "schemaExcerpt": [
    {
      "name": "timestamp_us",
      "physicalType": "INT64",
      "logicalType": "TIMESTAMP(MICROS,UTC)",
      "representation": "timestamp:unit:string",
      "nullable": true,
      "supported": true
    },
    {
      "name": "timestamp_ns",
      "physicalType": "INT64",
      "logicalType": "TIMESTAMP(NANOS,local)",
      "representation": "timestamp:unit:string",
      "nullable": true,
      "supported": true
    },
    {
      "name": "date",
      "physicalType": "INT32",
      "logicalType": "DATE(days since 1970-01-01)",
      "representation": "date:days:string",
      "nullable": true,
      "supported": true
    },
    {
      "name": "time_us",
      "physicalType": "INT64",
      "logicalType": "TIME(MICROS,UTC)",
      "representation": "time:unit:string",
      "nullable": true,
      "supported": true
    }
  ],
  "range": {
    "start": 1,
    "end": 4
  },
  "rowsExcerpt": [
    {
      "timestamp_us": "-1",
      "timestamp_ns": "-1",
      "date": "-719162",
      "time_us": "0"
    },
    {
      "timestamp_us": "-1000001",
      "timestamp_ns": "-1000000001",
      "date": "-1",
      "time_us": "1"
    },
    {
      "timestamp_us": "0",
      "timestamp_ns": "0",
      "date": "0",
      "time_us": "12345678901"
    },
    {
      "timestamp_us": "9007199254740993",
      "timestamp_ns": "9007199254740993",
      "date": "20000",
      "time_us": "86399999999"
    }
  ]
}

-1 is retained as an exact count, not converted to a Date. timestamp_us uses microseconds with UTC adjustment; timestamp_ns uses nanoseconds with local semantics. DATE is days since 1970-01-01 and TIME uses its declared unit. The full exported schema carries these labels.

Distinguish empty binary bytes from null

{
  "file": "mixed-snappy-v1.parquet",
  "format": "json",
  "start": 1,
  "end": 5,
  "safe": true,
  "excerptColumns": [
    "binary",
    "fixed_binary"
  ]
}
{
  "schemaExcerpt": [
    {
      "name": "binary",
      "physicalType": "BYTE_ARRAY",
      "logicalType": "NONE",
      "representation": "binary:hex",
      "nullable": true,
      "supported": true
    },
    {
      "name": "fixed_binary",
      "physicalType": "FIXED_LEN_BYTE_ARRAY",
      "logicalType": "NONE",
      "representation": "binary:hex",
      "nullable": true,
      "supported": true
    }
  ],
  "range": {
    "start": 1,
    "end": 5
  },
  "rowsExcerpt": [
    {
      "binary": "0x00ff80",
      "fixed_binary": "0x00ff8001"
    },
    {
      "binary": "0x",
      "fixed_binary": "0x00000000"
    },
    {
      "binary": "0x000102030405060708090a0b0c0d0e0f",
      "fixed_binary": "0x41424344"
    },
    {
      "binary": "0x5c4e",
      "fixed_binary": "0xfffefdfc"
    },
    {
      "binary": null,
      "fixed_binary": null
    }
  ]
}

Binary values use lowercase hexadecimal with a 0x prefix. Empty bytes become 0x; null remains null. The text is an inspection representation, not UTF-8 decoding of arbitrary bytes.

Retain special floating-point values

{
  "file": "mixed-snappy-v1.parquet",
  "format": "json",
  "start": 1,
  "end": 5,
  "safe": true,
  "excerptColumns": [
    "float64_value",
    "float32_value"
  ]
}
{
  "schemaExcerpt": [
    {
      "name": "float64_value",
      "physicalType": "DOUBLE",
      "logicalType": "NONE",
      "representation": "float:number-or-special-string",
      "nullable": true,
      "supported": true
    },
    {
      "name": "float32_value",
      "physicalType": "FLOAT",
      "logicalType": "NONE",
      "representation": "float:number-or-special-string",
      "nullable": true,
      "supported": true
    }
  ],
  "range": {
    "start": 1,
    "end": 5
  },
  "rowsExcerpt": [
    {
      "float64_value": "NaN",
      "float32_value": "NaN"
    },
    {
      "float64_value": "+Infinity",
      "float32_value": "+Infinity"
    },
    {
      "float64_value": "-Infinity",
      "float32_value": "-Infinity"
    },
    {
      "float64_value": "-0",
      "float32_value": "-0"
    },
    {
      "float64_value": 1.25,
      "float32_value": 1.5
    }
  ]
}

NaN, positive/negative infinity and negative zero are explicit schema-described strings. Finite values 1.25 and 1.5 remain JSON numbers. The representation avoids JSON.stringify silently turning nonfinite numbers into null or losing negative zero.

Protect a multiline formula-like text value

{
  "file": "mixed-snappy-v1.parquet",
  "format": "csv",
  "start": 4,
  "end": 4,
  "safe": true,
  "excerptColumns": [
    "name",
    "int32_value"
  ]
}
{
  "csvRecordsExcerpt": [
    [
      "name",
      "int32_value"
    ],
    [
      "'=SUM(1,2)\r\n\"quoted\",value",
      "-1"
    ]
  ]
}

These are selected records after ordinary CSV parsing, not the complete 14-column CSV. Formula protection adds an apostrophe to the risky text. Commas, quotes and CRLF inside the value survive quoting. Numeric INT32 -1 stays -1; exact values represented as risky strings can receive protection too.

Decode raw CSV nulls and backslashes

{
  "file": "mixed-snappy-v1.parquet",
  "format": "csv",
  "start": 2,
  "end": 5,
  "safe": false,
  "excerptColumns": [
    "name",
    "active"
  ]
}
{
  "csvRecordsExcerpt": [
    [
      "name",
      "active"
    ],
    [
      "",
      "false"
    ],
    [
      "\\\\N",
      "\\N"
    ],
    [
      "=SUM(1,2)\r\n\"quoted\",value",
      "true"
    ],
    [
      "\\N",
      "false"
    ]
  ]
}

These records are shown after CSV parsing. Source text backslash-N gains a second leading backslash; a true null is one backslash-N token. Empty text stays empty. Raw mode preserves the formula-like string and can be unsafe in a spreadsheet; it still applies null/backslash encoding.

Keep a supported column beside a nested placeholder

{
  "file": "unsupported-struct.parquet",
  "format": "json",
  "start": 1,
  "end": 2,
  "safe": true,
  "excerptColumns": [
    "flat_id",
    "nested"
  ]
}
{
  "schemaExcerpt": [
    {
      "name": "flat_id",
      "physicalType": "INT32",
      "logicalType": "NONE",
      "representation": "number",
      "nullable": true,
      "supported": true
    },
    {
      "name": "nested",
      "physicalType": "GROUP",
      "logicalType": "NONE",
      "representation": "unsupported:null",
      "nullable": true,
      "supported": false,
      "reason": "nested"
    }
  ],
  "range": {
    "start": 1,
    "end": 2
  },
  "rowsExcerpt": [
    {
      "flat_id": 1,
      "nested": null
    },
    {
      "flat_id": 2,
      "nested": null
    }
  ]
}

flat_id still contains 1 and 2. The nested struct remains a GROUP schema entry with supported:false, reason nested and unsupported:null. Both shown nulls are placeholders, even if the original nested values differ. Use the JSON schema to distinguish this from a decoded nullable column.

Label an unsupported compression codec

{
  "file": "unsupported-gzip.parquet",
  "format": "json",
  "start": 1,
  "end": 3,
  "safe": true,
  "excerptColumns": [
    "value"
  ]
}
{
  "schemaExcerpt": [
    {
      "name": "value",
      "physicalType": "INT32",
      "logicalType": "NONE",
      "representation": "unsupported:null",
      "nullable": true,
      "supported": false,
      "reason": "codec"
    }
  ],
  "range": {
    "start": 1,
    "end": 3
  },
  "rowsExcerpt": [
    {
      "value": null
    },
    {
      "value": null
    },
    {
      "value": null
    }
  ]
}

GZIP is not decoded by this viewer. The value column remains in the schema with reason codec and all rows use null placeholders. This result says nothing about the original cell values; use a full Parquet reader to recover them. CSV alone would lose this explanation.

Common Parquet inspection mistakes

  • Assuming every .parquet file uses supported types, encodings and codecs, or treating an unsupported-column placeholder as a source null.
  • Treating a schema-described INT64 or DECIMAL string as a rounded floating-point number.
  • Interpreting an epoch-unit timestamp as milliseconds without checking its unit and UTC-adjustment flag.
  • Mistaking a clipped preview or hidden column for a truncated successful export.
  • Using zero-based positions or expecting the end row to be excluded from the export.
  • Reading CSV \N as ordinary text or forgetting to undo the extra leading backslash on escaped text values.
  • Expecting cancellation, Reset or a file change to keep the old worker session available.
  • Opening untrusted raw CSV in a spreadsheet or assuming double quotes prevent formula evaluation.

Limits and notes

  • One input file is limited to 8 MiB, 10,000 rows, 128 columns and 200,000 data cells. Supported decoded page data has a 32 MiB aggregate budget. A text cell and raw binary cell are each limited to 256 KiB; binary hex text can occupy 512 KiB plus its 2-character prefix. The footer is capped at 1 MiB, metadata at 256 row groups and supported data at 8,192 pages; column names are at most 128 UTF-8 bytes. Each export is capped at 8 MiB. Load, page and export work has a 10-second deadline. Limits fail the operation rather than silently returning a partial file or shortened export. The 25-row preview shows only the first 12 columns and clips at 500 characters per cell. Worker isolation, size checks and cancellation reduce load; they are not a hard limit on total browser memory or a guarantee that hostile input cannot exhaust resources.
  • The decoded subset is flat required or optional primitive columns, data page v1/v2, PLAIN or dictionary value encoding, and UNCOMPRESSED or SNAPPY compression. Supported physical values include BOOLEAN, INT32, INT64, FLOAT, DOUBLE, BYTE_ARRAY and FIXED_LEN_BYTE_ARRAY with supported logical annotations. Nested groups, LIST/MAP/repeated columns, INT96, unknown logical types, other value encodings and other codecs are retained as unsupported schema entries with a reason and null placeholder values; supported columns still load. These placeholders are not actual decoded nulls and cannot recover the original values. Encrypted files, external column files, corruption and remote URLs are not accepted. Support here is narrower than the full Parquet specification or the upstream library. Invalid combinations of known physical/logical types fail validation. Unsupported column payloads are skipped and are not decompressed or integrity-validated; readable supported columns do not prove that every byte of the source file is valid.
  • INT64 and UINT64 values are decimal strings, including values that would fit a JavaScript number. DECIMAL is an exact fixed-point string retaining its declared scale. Binary data is a 0x-prefixed hexadecimal string. DATE, TIME and TIMESTAMP values remain exact integer strings in their declared day or time units; read the logical type for the unit and UTC-adjustment flag. They are not formatted dates or time-zone-converted JavaScript Date values. Finite floating-point values remain numbers; NaN, +Infinity, -Infinity and negative zero use explicit strings described by the schema. Unsupported types are marked in the schema and represented by null placeholders; check supported and reason before interpreting nulls.
  • JSON exports an envelope containing schema, the chosen range and row objects. The schema records each column’s name, physicalType, logicalType, representation, nullable, supported and reason fields. A supported source null remains JSON null, empty text stays an empty string and precise values retain the documented string representation. An unsupported column has representation unsupported:null and null placeholders; the schema retains the reason, so those placeholders must not be interpreted as source nulls. The export is a readable inspection format, not a Parquet re-encoder or a guarantee that another application will recreate every original metadata field. A row range is 1-based, inclusive and contiguous; it is not a filter, sort, arbitrary selection or SQL query. A zero-row file still shows its schema; only its empty range start 0 / end 0 is allowed, producing a JSON envelope with rows [] or CSV headers only.
  • CSV uses UTF-8, comma separators, CRLF records and quoted headers/non-null values; null is the unquoted token \N. Every non-null text value beginning with a backslash gains one additional leading backslash, so the literal text \N becomes \\N after CSV parsing. To decode this convention, map a field exactly equal to \N to null; otherwise remove one backslash from a value starting with two. CSV does not carry the Parquet schema or unsupported-column reasons: placeholder nulls and real nulls cannot be distinguished by CSV alone. Choose JSON when these distinctions matter. Default formula protection prefixes risky headers and string-represented values with an apostrophe, including negative INT64, decimal or temporal strings and intentionally changes those strings; raw CSV disables that protection but retains null/backslash encoding. Raw CSV can activate formulas or trigger spreadsheet type coercion. Downloads are parquet-rows.csv or parquet-rows.json. Preview clipping never clips an otherwise successful export. Cancellation, timeout, worker failure, changing the file, Reset and leaving the tool release the session; downloaded files remain on your device.

Frequently asked questions

Why are large integers, decimals and timestamps exported as strings?

Converting them to ordinary JavaScript numbers or Date objects can round an integer, change a decimal scale or lose sub-millisecond timestamp precision. This tool keeps INT64/UINT64 as decimal strings, DECIMAL as fixed-point text and temporal values as exact unit counts. The JSON schema describes how to interpret them, including time units and UTC adjustment. It does not guess a display time zone.

Does the CSV distinguish null from an empty string and the text \N?

Null is the token \N. Every actual text value beginning with a backslash gains another leading backslash, so the string \N becomes \\N after CSV parsing; an empty string stays empty. Decode the exact \N token as null and remove one leading backslash from other values that start with two. CSV still omits type information and the reasons for unsupported-column null placeholders. Use the JSON envelope to preserve those schema distinctions.

What happens to an unsupported Parquet column?

The column remains in the schema with supported:false, a reason and representation unsupported:null. Its exported and displayed values are null placeholders, while supported columns remain readable. Do not treat a placeholder as a genuine source null. Nested columns, INT96 and non-supported logical types, encodings or codecs can trigger this behavior. Corrupt metadata, invalid supported data and resource-limit violations fail the load. Use a full Parquet tool for the unsupported original values.

Does exporting a range include only the visible preview?

No. Start and end are 1-based inclusive row positions in original file order. A successful export contains every column and complete values in that range, regardless of the current preview page, its 12-column cap or 500-character clipping. An invalid range or oversized output is rejected. An empty file has only the special range 0–0, exporting schema or headers without data rows. Cancel or Reset discards the loaded session and requires reloading the file.

Related tools