You upload a PDF, ask a question, and get an error. Or the AI says it read the file but misses pages, invents headings, or answers using only the beginning.
The problem may be the document itself. PDFs can look alike on screen while containing very different things: one may hold selectable text, while another is just images of printed pages. Check the file before repeatedly uploading it or rewriting your prompt.
Check whether the PDF contains selectable text
Open the file in a PDF reader and try to select a sentence. If you can highlight individual words, the PDF probably contains a text layer. If you can select only a whole page, or nothing, it may be a scan.
You can also search for a word you can see on the page. If a word on page 3 does not appear in the PDF search, that page may be an image rather than machine-readable text.
Run OCR on scanned pages
Optical character recognition (OCR) converts text in an image into machine-readable text. A 32-page scanned contract from 2014 may look readable to you, while the software sees 32 pictures.
Run the document through an OCR tool before uploading it. Adobe Acrobat and approved scanning apps can create searchable text layers. Google Drive’s mobile scanner can also create searchable PDFs.
Afterward, search for words on several pages to check that the conversion worked. OCR can misread low-resolution or faded scans, handwriting, unusual fonts, skewed pages, tables, small text, and documents in multiple languages. A searchable PDF is easier to process, but its text may still contain errors.
Check protection and permissions
A PDF may contain text but restrict copying or editing. It may also be password-protected, encrypted, or generated by a document-management system with unusual permissions.
If you are authorized to work with the file, open it normally and check its security settings. Do not bypass restrictions on files you are not permitted to modify or extract. For client documents, check whether the client permits you to upload the information to the AI service. Being able to access a file does not automatically give you permission to process it elsewhere.
Check the file size
High-resolution scans and photographs can make a PDF large. A 40-page text document might be only a few megabytes, while a 40-page scan can be much bigger.
If the AI service rejects the file, compression may help. Keep the original and work on a copy. After compressing, view several pages at 100 percent zoom to make sure small text remains readable.
Split long PDFs by subject
Splitting a 300-page report into pages 1–50, 51–100, and 101–150 is simple, but it may cut chapters in half or separate a table from its explanation.
Split by the document’s structure instead:
- Executive Summary
- Background
- Methodology
- Findings
- Financial Analysis
- Appendices
Use clear names such as 01_Executive_Summary.pdf, 02_Methodology.pdf, and 03_Findings.pdf. This helps you keep track of which sections you and the AI have processed.
Watch for complex layouts and tables
Extraction can become unreliable when a PDF contains multiple columns, floating text boxes, sidebars, footnotes, headers, charts, callouts, or an unusual reading order. A person reads down the left column and then the right. An extraction tool might mix lines from both columns into sentences that were never there.
If an answer sounds scrambled, inspect the extracted text before blaming the model.
Tables can lose their rows and columns too. This table:
| Product | January | February |
|---|---|---|
| A | 120 | 145 |
| B | 90 | 106 |
might come out as “Product January February A 120 145 B 90 106.” If the figures matter, export the table to CSV or Excel and give the structured data to the AI separately.
When formatting does not matter, convert the PDF to Word, plain text, or another accessible format, then inspect the result. If your question concerns one section, copy that passage instead of uploading the whole file. With client material, that can also reduce how much information you send to the AI service.
Check what the AI could read
An upload succeeding does not mean the text was extracted correctly. Before asking for analysis, try this:
Before analyzing this document, report:
- The document title.
- The number of pages you can access.
- The main section headings you can identify.
- The heading on the final page.
- One specific detail from the last substantive section.
If any part appears unreadable or missing, tell me before continuing.
Compare the answer with the original. This will not prove that every sentence was extracted correctly, but it can reveal obvious gaps.
If you want to compare tools, use the same non-confidential test PDF in two or three services. Check whether each can identify headings, find information near the end, read tables, distinguish footnotes, and extract scanned text. The failure may come from a particular document parser rather than the language model.
PDF troubleshooting guide
| Symptom | Likely cause | What to try |
|---|---|---|
| AI says the file has no readable text | Scanned or image-only PDF | Run OCR |
| Some pages are missing | Extraction or file-length problem | Split by section and check page coverage |
| Text appears scrambled | Columns or complex layout | Convert to Word or plain text, or extract sections |
| Table values are mixed up | Poor table extraction | Export the table to CSV or Excel |
| Upload fails | File too large | Compress carefully or split it by subject |
| Text cannot be copied | Scan or document restriction | Check OCR and authorized permissions |
| AI answers only from early pages | Partial extraction or context problem | Ask it to verify details from the final sections |
| Names or numbers are wrong | OCR error | Compare the extracted text with the original |
| AI cannot open a protected PDF | Security restrictions | Use an authorized, accessible copy |
| One tool fails but the PDF looks normal | Parser-specific problem | Try another tool |
A reliable workflow
- Check whether the text is selectable.
- Run OCR if needed.
- Check file protection and your permissions.
- Check the file size.
- Split long documents by logical sections.
- Export complicated tables separately.
- Convert the file if you do not need its layout.
- Remove information you do not need or are not permitted to share.
- Ask the AI to show which parts it can access.
- Compare important answers with the original.
Silent extraction failures are harder to catch than an error message. A confident summary based on only part of a PDF may look complete. Treat extraction as something to check.