PDF Extractor
Export the images inside a PDF as a ZIP and its text as a .txt file, in one click.
What it does
Screenshots are a poor way to get product photos out of a catalogue or charts out of a report: quality drops and it takes forever. This tool finds the original images embedded in the pages and hands them over in a single ZIP.
You can also export all of the document's text, in page order, to a plain .txt file. Handy for searching a long document or moving its text into another tool.
How to use it
- Drop your PDF file.
- Choose whether you want the images, the text, or both.
- Press "Extract and download".
When is it useful?
- Pulling images out of product catalogues
- Moving charts and diagrams from a report into a presentation
- Making all the text in a document quick to copy
- Recovering logos and photos from old brochures
Tips
- If you want the whole page as a picture, PDF → Image is the better tool.
- For text that keeps headings and paragraphs, try PDF → HTML.
Frequently asked questions
- Which images are extracted?
- Almost every image type pdf.js can render is supported (JPEG, PNG and most scanned image formats) as long as it is embedded in the page. In a few rare cases that use unusual or custom masking, the image in question may be skipped.
- Does this tool take a screenshot of the PDF pages?
- No, this is different from the "PDF → Image" tool: here we extract the original image files EMBEDDED in the page, not the whole page.
- Does the text output keep the formatting?
- No, the text is extracted one block after another as plain text (.txt) in page order. If you want output that preserves headings and paragraph structure, use the "PDF → HTML" tool.
- At what resolution are images extracted?
- At the resolution they were embedded in the PDF. Images that were scaled down when the file was made are not enlarged.