Guides / Duplicate keys
Detect Duplicate JSON Keys Before Exporting CSV
Check the raw JSON for repeated object keys before ordinary parsing, and stop the export if you find any. A parser that keeps the last value can erase the earlier value before CSV conversion starts. Keep the original file and ask the producer which values belong in the schema; the CSV writer cannot recover discarded data.
Why a successful conversion can still lose a value
RFC 8259, section 4 says object names should be unique and warns that receivers handle repeated names differently. Duplicate keys are not a universally rejected syntax error. This workflow chooses to reject them before exporting.
Python's json documentation describes its default behavior: the last value survives. Its object_pairs_hook provides key/value pairs before dictionary construction, allowing a duplicate check.
An original two-record counterexample
[{"id":"K01","amount":7,"amount":9},{"id":"K02","amount":7,"\u0061mount":9}]
K01 repeats the literal key. K02 spells the second key using a Unicode escape that decodes to the same name. The local default decode retained amount: 9 for both records. It did not retain either earlier value of 7. Comparing only record counts would miss this loss.
Download the blocked source and a separate clean example
The clean sample uses amount_original and amount_adjusted. This is an invented demonstration schema, not an inference about what the broken source meant. A real repair needs the producer's field contract.
python reject-duplicate-json-keys.py duplicate-keys-source.json blocked.csv
# Stops with ValueError: duplicate object key: 'amount'
python reject-duplicate-json-keys.py duplicate-keys-clean-sample.json clean.csv
Use new output filenames. The original run rejected both the whole source and the escaped-key example when checked separately. The blocked export created no CSV. The clean export produced 2 data records and 3 columns, verified by parsing its CSV. These measurements describe the downloadable fixtures only.
The preflight, explained
- Read the original JSON text without first formatting it through another parser.
- For each object, collect decoded key/value pairs and reject a key already seen in that object.
- Validate the selected array and its fixed schema before opening the CSV file.
- Write only the reviewed source to a new output file.
Show the complete original Python example
"""Small fixed-schema example. Reject repeated keys before writing any CSV.
python reject-duplicate-json-keys.py duplicate-keys-clean-sample.json new-output.csv
"""
import csv,json,sys
from pathlib import Path
from decimal import Decimal
def unique_object(pairs):
result={}
for key,value in pairs:
if key in result:
raise ValueError(f"duplicate object key: {key!r}")
result[key]=value
return result
def invalid_constant(value):
raise ValueError(f"unsupported numeric constant: {value}")
def decode(raw):
return json.loads(raw,object_pairs_hook=unique_object,
parse_float=Decimal,parse_constant=invalid_constant)
def export(source,target):
rows=decode(Path(source).read_text(encoding='utf-8'))
columns=['id','amount_original','amount_adjusted']
if not isinstance(rows,list) or not rows:
raise ValueError('expected a nonempty array')
for row in rows:
if not isinstance(row,dict) or set(row)!=set(columns):
raise ValueError('exactly id, amount_original and amount_adjusted required')
if not isinstance(row['id'],str):raise ValueError('id must be text')
for key in columns[1:]:
if isinstance(row[key],bool) or not isinstance(row[key],(int,Decimal)):
raise ValueError('amounts must be numbers')
with Path(target).open('x',encoding='utf-8',newline='') as file:
writer=csv.writer(file);writer.writerow(columns)
writer.writerows([[row[key] for key in columns] for row in rows])
return {'records':len(rows),'columns':len(columns)}
if __name__=='__main__':print(json.dumps(export(*sys.argv[1:3])))
The hook also checks nested objects, but stops at the first duplicate. This short example reports the decoded key, not a full object path or source line number. It loads the file in memory and exports only the demonstrated schema; it is not a streaming or general-purpose repair tool. It also rejects NaN and Infinity constants and uses Decimal for decimal-number parsing, without claiming to validate every interoperability issue.
Question: Can I detect duplicate keys after JSON has already become a dictionary?
Not from that dictionary alone if the parser already discarded the earlier pairs. Return to the raw file or obtain an unmodified export. Pretty-printing the parsed result cannot bring back missing values. Preserve the raw source as evidence before any conversion.
Watch the workflow
Do not silently choose first or last
Repeated IDs in separate records are a different issue: this check concerns repeated names inside the same object. A duplicate may indicate an upstream bug or a schema that needs distinct fields. Neither “first wins” nor “last wins” proves which value is correct. Record the rejection, resolve the schema, and rerun the export.
After the source is resolved, verify every row and field. A matching record count is necessary for a one-record-per-object export but cannot prove values survived parsing. CSV escaping and spreadsheet interpretation require their own checks.