A JSON array and a JSON Lines file can hold exactly the same records. The difference is how they behave when files get big, when data arrives continuously, or when something goes wrong halfway through.
The two shapes
A JSON document is one value. To store many records you wrap them in an array:
[
{"id": 1, "level": "info", "msg": "started"},
{"id": 2, "level": "warn", "msg": "slow query"}
]
JSON Lines puts one complete JSON value on each line, with no enclosing array and no commas between records:
{"id":1,"level":"info","msg":"started"}
{"id":2,"level":"warn","msg":"slow query"}
The rules, as written at jsonlines.org, are short: the file is UTF-8, each line is a valid JSON value, and lines are separated by \n (a preceding \r is tolerated). The conventional extension is .jsonl. NDJSON (newline-delimited JSON) is the same idea under a different name, usually served as application/x-ndjson; in practice the two are interchangeable. A third, rarer cousin is RFC 7464 “JSON text sequences”, which prefixes each record with an ASCII record-separator character.
Why JSON Lines exists
Streaming. A reader can process a JSON Lines file one line at a time with constant memory. A standard JSON parser must read the whole array before it returns anything; streaming JSON parsers exist, but they are more complex and less common.
Appending. Adding a record to JSON Lines means writing one more line. Adding to a JSON array means rewriting the closing bracket, which is awkward for log writers and unsafe when several processes write to the same file.
Fault tolerance. If a process crashes mid-write, a JSON array is left without its closing ] and the whole file fails to parse. A JSON Lines file loses at most the last partial line; everything before it is still readable.
Unix tools. Because records are lines, ordinary tools work: wc -l counts records, head samples them, split -l shards a file for parallel processing, and grep finds candidates before a real parser looks at them.
That is why JSON Lines is the default for structured logs (pino, Bunyan, many cloud logging agents), for the Elasticsearch and OpenSearch _bulk API, for BigQuery’s newline-delimited JSON import and export, and for machine-learning datasets and batch request files.
Why a plain JSON array is still often better
For an API response, a configuration file or any payload that is read in one piece, a JSON array is simpler. Every JSON library parses it with one call, browsers’ response.json() handles it directly, and it can be pretty-printed with indentation.
That last point is the key rule of JSON Lines: a record must stay on one line. Pretty-printing a JSON Lines file turns it into something that is neither valid JSON nor valid JSON Lines. To read records comfortably, format each line separately without writing the result back; the JSON Lines formatter does that and points at the exact line number when one record is broken.
A JSON array also has an unambiguous empty state ([]), whereas an empty JSON Lines file is just an empty file, and readers differ on whether blank lines are allowed. Most skip them; strict validators reject them.
Summary of the trade-offs:
- whole-document APIs and configs: JSON
- logs, event streams and append-only data: JSON Lines
- large exports, bulk imports and datasets: JSON Lines
- anything a person edits by hand: JSON (or YAML)
Converting between the two
With jq, -c prints each output on one compact line and .[] iterates an array, so turning an array into JSON Lines is one command; -s (slurp) reads every line into one array for the reverse direction:
jq -c '.[]' orders.json > orders.jsonl
jq -s '.' orders.jsonl > orders.json
In Python, write each record with json.dumps(record) followed by a newline, and read with a loop over the file object calling json.loads(line) for non-empty lines. Avoid indent= when writing JSON Lines.
In the browser, PasteKit converts both ways locally: JSON to NDJSON and NDJSON to JSON. Number precision is preserved, which matters for 64-bit IDs that would otherwise be rounded by JSON.parse.
Reading JSON Lines in common tools
Most data tools read the format directly, often under one of its other names:
- pandas:
pd.read_json("events.jsonl", lines=True), withchunksize=to stream large files in pieces. - DuckDB:
SELECT * FROM read_json_auto('events.jsonl')infers columns, and it queries gzipped files too. - Apache Spark:
spark.read.json(path)expects JSON Lines by default; a single multi-line JSON document needs themultiLineoption instead. - BigQuery: load jobs use the
NEWLINE_DELIMITED_JSONsource format. - jq: processes a stream of values natively, so every filter in the jq cheat sheet works line by line without extra flags.
If a tool claims your “JSON” file is invalid, check whether it expected one document and you gave it lines, or the other way round; that mismatch causes most of the confusing errors around these two formats.
Common problems
- A record split over two lines. Usually caused by a pretty-printer or by an unescaped newline inside a string. Strings in JSON must encode line breaks as
\n. - A trailing comma at the end of each line. That is the result of stripping the brackets off an array by hand. Remove the commas.
- A byte-order mark on the first line. Some Windows tools add one, and strict parsers reject the first record. Save as UTF-8 without BOM.
- Mixed record shapes. JSON Lines does not require every record to have the same keys, but tools that turn records into table columns, such as BigQuery schema auto-detection, may infer types from a sample and then fail on a later line.
Frequently asked questions
Is NDJSON the same as JSON Lines?
For practical purposes, yes. Both mean one JSON value per line separated by newlines; the names come from two separate write-ups of the same convention.
Is a .jsonl file valid JSON?
No, unless it contains exactly one line. A JSON parser will stop at the second value and report unexpected content after the first one.
Can a JSON Lines record contain newlines?
Only escaped, as \n inside strings. A literal line break inside a record splits it into two invalid lines.
Should I gzip JSON Lines files?
Often, yes. The format compresses well, and .jsonl.gz files can still be streamed line by line through a decompressor.