Guides / Decimal tokens

Keep Decimal Digits, Exponents and Trailing Zeros in CSV

Capture the original JSON number text before converting it to a binary float, then write that text into the CSV. Formatting a parsed float afterward cannot recover digits or spelling already lost. In Python, json.loads number hooks can retain 0.100000000000000005, 1e-7 and 1.2300 as distinct source tokens. Import the resulting CSV column as text when exact spelling matters.

The loss can happen before CSV is written

A conversion has two separate boundaries: JSON decoding and CSV import into the next application. If decoding has already replaced a token with a float, increasing the CSV decimal-place setting will not restore the original token. If the CSV text is correct but a spreadsheet infers a numeric type, its display or stored value can change afterward.

The official Python JSON decoder documentation describes parse_float and parse_int hooks that receive number strings. This example uses a dedicated string subclass so an unquoted JSON number can still be distinguished from a quoted JSON string. It does not use a regular expression to guess numbers inside the document.

An original six-record comparison

The following is local execution evidence for the downloadable fixture, not a benchmark of every converter. The default column is Python's str of its decoded float or integer; the preserved column is the actual parsed CSV cell.

RecordJSON source tokenDefault decoded value as textPreserved CSV cell
D010.1000000000000000050.10.100000000000000005
D021e-71e-071e-7
D031.23001.231.2300
D04-0.00-0.0-0.00
D052E+32000.02E+3
D06121212

The run produced 6 data records and 2 columns, and every amount token matched its source text. The example covers a long decimal, exponent notation, trailing zeros, negative zero, a capital exponent with an explicit plus sign, and an integer. These are different checks: a changed exponent spelling can represent the same numeric value, while rounding can change the represented value.

Download and run the narrow exporter

Original JSON fixtureToken-preserving Python exporterMeasured CSV outputExecution evidence
python decimal-tokens-to-csv.py decimal-token-example.json new-output.csv

Use the unchanged JSON file as input. The script accepts a nonempty array whose records contain exactly a string id and a numeric amount. It writes id,amount_token; its output column name deliberately says that the value is source text. It rejects a quoted amount such as "1.2300", duplicate keys and nonstandard numeric constants. It will not overwrite an existing output file.

Preserving spelling is different from doing arithmetic

Choose a token-preserving export when an audit, comparison or downstream contract needs the same numeric text. If your goal is calculation, choose a declared decimal arithmetic policy instead: allowed scale, rounding mode and output formatting. A decimal value object can help with arithmetic, but a later formatter may still normalize exponent notation or trailing zeros. Keep the original token separately if both requirements apply.

Do not cast the preserved amount back to float before writing it. Do not assume placing quotes around a CSV field forces every spreadsheet to keep it as text. CSV quoting handles field boundaries; the importing application controls data types. In its import workflow, explicitly select text for amount_token, then compare a few cells with the source fixture.

Can I repair a rounded CSV without the original JSON?

You cannot reliably infer the missing original digits or notation from a rounded value alone. Several source tokens can lead to the same decoded value. Retrieve the raw JSON export or ask its producer to export the field as a string under a documented schema. If the producer rounded it before creating the JSON, this workflow cannot recover that earlier information either.

Read the exporter

Show the original Python implementation
"""Export original amount number tokens from a small flat JSON array.
python decimal-tokens-to-csv.py decimal-token-example.json new-output.csv
"""
import csv, json, sys
from pathlib import Path

class NumberToken(str):
    pass

def unique_object(pairs):
    result = {}
    for key, value in pairs:
        if key in result:
            raise ValueError('duplicate key: ' + repr(key))
        result[key] = value
    return result

def reject_constant(value):
    raise ValueError('non-JSON number: ' + value)

def export(source, target):
    source, target = Path(source), Path(target)
    if source.stat().st_size > 2_000_000:
        raise ValueError('this small example caps input at 2,000,000 bytes')
    if target.exists():
        raise FileExistsError('choose a new output filename')
    records = json.loads(source.read_text(encoding='utf-8'),
        parse_float=NumberToken, parse_int=NumberToken,
        parse_constant=reject_constant, object_pairs_hook=unique_object)
    if not isinstance(records, list) or not records:
        raise ValueError('a nonempty array is required')
    for record in records:
        if not isinstance(record, dict) or set(record) != {'id', 'amount'}:
            raise ValueError('each record must contain only id and amount')
        if type(record['id']) is not str:
            raise ValueError('id must be a JSON string')
        if not isinstance(record['amount'], NumberToken):
            raise ValueError('amount must be a JSON number, not a quoted string')
    with target.open('x', encoding='utf-8', newline='') as file:
        writer = csv.writer(file)
        writer.writerow(['id', 'amount_token'])
        writer.writerows((record['id'], record['amount']) for record in records)

if __name__ == '__main__':
    export(sys.argv[1], sys.argv[2])

Scope and handoff checks

This is a small flat-record recipe, not a streaming converter: it reads the source into memory and deliberately caps input at 2,000,000 bytes. It rejects extra fields, nested values and missing fields instead of silently deciding how to flatten them. The cap is this example's policy, not a measured browser or Python limit. It preserves number token text, not the surrounding whitespace or complete JSON file bytes.

The CSV writer escapes commas, quotes and line breaks, but it does not sanitize spreadsheet formulas in user-supplied string IDs. Inspect untrusted data and import it with an appropriate text policy. Check both record IDs and amount tokens after transfer; the local evidence applies only to the supplied fixture.

For other failure modes, detect duplicate keys before decoding, keep distinct nested paths in separate columns, and verify every row and field.

Open the JSON to CSV converter