Supabind

PDF Extract FAQ

Technical answers about AI processing, document limits, and extraction accuracy.

Is PDF Extract just standard OCR?

No. While standard OCR only turns images into text, Supabind uses Large Language Model (LLM) Vision. This means the AI actually understands the document layout, identifying where headers, rows, and nested tables exist, even if they aren't explicitly marked.

What is the difference between Gemini Flash and Gemini Pro?

Gemini 2.5 Flash (Balanced) is optimized for speed and works best for standard digital PDFs. Gemini 2.5 Pro (High Accuracy) has deeper reasoning capabilities, making it the choice for low-resolution scans, handwritten notes, or extremely complex catalog layouts.

What does the 'Confidence Score' actually measure?

It represents the AI's internal probability that the extracted table structure and text content match the source document. A score above 90% is production-ready; scores below 70% usually indicate that the PDF was too blurry or the layout was too ambiguous.

Can I extract multiple different tables from a single PDF?

Yes. The AI identifies all tabular data on a page. In the Extraction Results editor, you can use the 'Table Selector' dropdown at the top of the grid to switch between different tables found in the same document.

How do I handle PDFs where a single table spans multiple pages?

Our Gemini 2.5 models have a large 'context window,' allowing them to see the whole document at once. If you use the Balanced or High Accuracy mode, the AI will automatically stitch rows across page breaks into a single continuous table.

What are 'Extraction Prompts' and when should I use them?

Extraction Prompts allow you to give the AI specific instructions. For example, if you only need columns for 'SKU' and 'Price', you can update the prompt to say: 'Extract only SKU and Price columns.' This significantly improves accuracy on dense catalogs.

Why did my extraction fail with an 'Analysis Error'?

The most common reasons are: 1) The file exceeded the 10MB limit. 2) The PDF is password-protected. 3) The document contains no readable text or tables. Check your 'Processing History' logs for the specific error code.

Can I process documents in languages other than English?

Yes. Because we use advanced LLMs, the extraction engine supports over 50 languages, including complex scripts like Arabic, Chinese, and Japanese, as well as multi-lingual documents.

How secure are my documents during AI processing?

Your files are processed in a secure, isolated environment. We use enterprise-grade APIs where your data is not used to train the underlying models. Files are encrypted at rest and can be deleted permanently from your 'My Files' tab at any time.

What file types are supported?

Currently, we support .pdf files. Support for image formats (.jpg, .png) and multi-page TIFF files is currently in development and will be released in a future update.

Can I automate the extraction process?

Yes. You can connect an FTP/SFTP server to Supabind. Once connected, you can set the AI to automatically process any new PDFs that arrive in a specific folder without manual intervention.

Was this FAQ helpful?

Have a difficult document?

If your catalog or brochure has a highly custom layout that the AI is struggling with, our data team can build a custom extraction template for your account.

Contact Data Specialists