PDF to CSV
This free PDF to CSV converter reads a PDF and writes a CSV entirely in your browser. Pick a .pdf file or drop one in, choose how the text should be grouped, read the CSV, and download it. Nothing leaves your device, and there is no sign-up.
A PDF does not contain a table. It contains pieces of text, each with a position on the page — which is why copying a table out of a PDF often lands everything in one column. So this page has to rebuild the grid: it groups text runs into rows by their vertical position, then works out where the columns are from the horizontal position of every run. Because that is an inference rather than a fact stored in the file, the page tells you what it inferred: how many rows and columns it found, how many rows were missing a value, and — most importantly — whether any page had no text layer at all, which means it is a scan and there is nothing to read.
How to use this PDF to CSV converter
- Choose a
.pdffile, or drag one onto the panel. The sample table on the right is converted on load so you can see what the output looks like. - Pick Extract as: Table (by position) rebuilds rows and columns; Plain text lines writes one cell per line, which suits reports, invoices and statements that are not really a grid.
- If columns come out merged or split, change the Column gap and Row grouping. Looser values merge more; tighter values split more.
- Pick the Delimiter, press Convert to CSV, then read What the extraction found before you download.
What this tool supports
Tables rebuilt from position
Every text run on the page carries coordinates. Rows are formed by grouping runs whose vertical position is within a small distance of each other, and columns are formed by clustering the left edge of every run. Anything that shares a column lands in the same cell, separated by a space when a cell holds more than one run — which is how Alice Brown stays one cell instead of two.
| Setting | What it changes | When to move it |
|---|---|---|
| Row grouping | how far apart two runs can be vertically and still count as one row | rows came out split in two, or two rows merged into one |
| Column gap | how far apart two runs must be horizontally to count as separate columns | two columns merged into one cell, or one column split into several |
Pages with no text layer are named, not dropped
A scanned document, a photograph or a screenshot saved as a PDF has no text in it at all — only pixels. Extracting from those needs OCR, which this page does not run. Instead of handing you a short file and saying nothing, it lists exactly which pages have no text layer.
| What is in the PDF | What this page does |
|---|---|
| a real text layer (exported from a spreadsheet, a word processor, a reporting tool) | extracted |
| a scan or an image, no text layer | the page number is listed under the output; nothing is invented for it |
| a mix of both | text pages are extracted, image pages are listed |
Every page, in order
All pages are read in order and written into one CSV. A table that continues across pages comes out as continued rows; the header line is whatever the PDF printed at the top of each page, because that is what is on the page.
Delimiters and quoting
Choose comma, semicolon or tab. A value containing the delimiter, a double quote, a line break, or leading and trailing whitespace is wrapped in double quotes and any double quote inside it is doubled — so a total like 1,204.00 survives intact instead of splitting across two columns.
Privacy and limits
The PDF is parsed on your own device by JavaScript running in the page. No file is uploaded, nothing is stored after you close the tab, and there is no account. Because parsing happens in your browser's memory, very large documents (hundreds of pages) take longer; ordinary statements, invoices and reports take a moment. Password-protected files are refused with a clear message rather than being probed.
What this tool does not do
- No OCR. If a page is an image, there is no text to read. The page is named so you know, and nothing is guessed for it.
- No password handling. A protected PDF is refused. Unlock it first, then convert the unlocked copy.
- No ruling-line detection. Columns come from where the text sits, not from drawn table lines, so a borderless table works and a heavily decorated one may need the gap settings adjusted.
- No merged-cell or multi-level-header reconstruction. A header that spans two columns, or a cell that belongs to several rows, comes out as ordinary cells. Check the output before you rely on it.
- No merging of files. One PDF at a time, no batch mode.
- No promise of perfection. Position-based extraction is an inference. The report tells you what it inferred so you can check it, not so you can skip checking.
Frequently asked questions
Is this PDF to CSV converter free?
Yes. No account, no sign-up, no page limit, and no upload — the PDF is read in your browser.
Is my PDF uploaded anywhere?
No. The file you pick is read with the browser's own File API and parsed in the page. There is no request that carries your file or its contents.
Why is my scanned PDF empty?
Because it has no text layer — it is an image. This page does not run OCR, and it says so by listing the page numbers instead of writing empty rows. Use OCR software first, then convert the result.
It says my PDF is password protected. What now?
Open it with the password and save an unlocked copy, then convert that. This page will not attempt to bypass protection.
Everything landed in one column. What should I change?
Set Column gap to Tight. If that splits single values into several columns, go back to Normal or Loose. The two settings are the whole tuning surface.
Two rows merged, or one row split in two. What should I change?
Change Row grouping. Tight splits rows apart, Loose pulls them together.
A number like 1,204.00 broke into two columns.
In the CSV it did not: values containing the delimiter are wrapped in double quotes, so "1,204.00" stays one field. The preview shows the raw CSV, quotes included.
Does it handle a multi-page table?
Yes. Every page is read in order and its rows are appended. If the PDF repeats a header row on each page, that header appears in the CSV each time it does on the page.
How accurate is it?
For PDFs exported from spreadsheets and reporting tools, very good. For scanned pages, nothing at all — they are listed. For anything in between, read the report under the output: it tells you how many rows and columns were found and how many rows were missing a value, which is where errors show up first.
Is there a page limit?
There is no imposed limit. Parsing happens in your browser's memory, so very large documents depend on your device. Ordinary statements and reports take a moment.
What if I need another direction or format?
For the opposite direction use JSON to PDF. To take a table into a spreadsheet file use CSV to Excel, and for spreadsheet input use Excel to CSV or CSV to JSON.