PDF text extractor

Pull the selectable text out of a PDF, then copy it or save it as a .txt file. Worth knowing before you start: a scanned document is a picture of text, not text, so it holds nothing to extract. This page reads only what is really there, and it reads it in your browser.

Everything runs on your device. Your file is opened by this page in your browser and never uploaded, so nothing is sent to FORKSAI or anyone else. Close the tab and it is gone.

Why some PDFs give you nothing

There are two completely different things people call a PDF. One holds characters, with a font and a position for each of them. The other holds a photograph of a page. They look identical on screen and behave nothing alike.

This tool reads the first kind. On the second kind it returns nothing, and it says so rather than handing you an empty box and letting you wonder. The test takes two seconds: open the PDF in any reader and try to select a word with your cursor. If the highlight will not stick, the characters are not there.

Turning a picture of text back into text is optical character recognition, a different job with a different failure mode: it guesses, and it guesses badly on handwriting, on bad lighting and on anything set in columns. This page does not do it, and does not pretend to.

Why the spacing comes out odd

A PDF does not store sentences or paragraphs. It stores fragments of text with coordinates, in whatever order the program that made the file happened to write them. Rebuilding lines from that is reconstruction, not reading, and it is why extracted text so often needs a tidy up.

The usual troublemakers are two column layouts, where the reading order on screen is not the order in the file, tables, which arrive as a stream of cells with the structure gone, and running headers and footers, which turn up in the middle of the text once per page.

Ligatures are worth watching for too. In many fonts fi and fl are single characters, so words like file and flow can come out looking odd in a plain text editor. A search and replace fixes the lot in one pass.

Frequently asked questions

How do I get the text out of a PDF?
Load the file and the text appears in the box below, page by page. Copy it to the clipboard or save it as a .txt file. Nothing is retyped and nothing is guessed: this is the text the PDF already carries.
Why did my PDF return no text at all?
Because it is a scan. A photographed or scanned page is a picture of text, not text, and a picture holds no characters to extract. This tool reads only real, selectable text, so a scan comes back empty. The quick check is to try selecting a word in your usual PDF reader: if the cursor will not highlight it, this tool will find nothing. Getting text out of a scan needs optical character recognition, which this page does not do.
Are my files uploaded to a server?
No. The page opens your file with JavaScript running in your own browser, does the work there, and hands you the result as a download. Nothing is uploaded, so there is no server copy to delete and no queue to wait in. You can check this yourself: open your browser's network tab, run the tool, and you will see no request carrying your file.
Why is the spacing and line breaking odd?
A PDF stores where each fragment of text sits on the page, not sentences and paragraphs. Rebuilding lines from those positions is guesswork, and it goes wrong most often on multi column layouts, tables and headers, where reading order on screen is not the order the fragments are stored in. Expect to tidy the result.
Can I extract text from only some pages?
Yes. Set a page range, such as 12-18, before extracting.
Does this work on a protected PDF?
A PDF that is encrypted with a password cannot be read at all until it is unlocked. A PDF that merely sets a no copying flag is a different matter: the flag is an instruction to reader software, not encryption, and this tool reads the text regardless. Respect whatever licence the document carries.
Is this PDF text extractor free?
Yes, and there is no account, no email box and no watermark on the output. The tool is here because FORKSAI makes study software and this is a useful thing to give away.

You have the text. Now make it stick.

Paste it into FORKSAI and it comes back as flashcards, a summary and a spaced repetition schedule, so the reading turns into recall.

The tool above stays free and needs no account.

More free PDF tools

Every one of them runs in your browser. See the rest of the free student tools.

Plan your next review session

Use the free student study kit to plan a week of revision, check flashcard quality, and record mistakes from practice questions.

Need a deck first? Turn a lecture PDF into editable flashcards, then follow the active recall guide to practise answering before revealing the back of each card.

Working toward an exam? Build a daily target with the exam study planner, protect the session with the Pomodoro study timer, or check your current result with the weighted GPA calculator.