Supabind

AI Models & Configuration

Understanding the trade-offs between Fast, Balanced, and High Accuracy extraction modes.

Overview

Supabind doesn't use static OCR. Instead, we use state-of-the-art Large Language Models (LLMs) with vision capabilities. Depending on the complexity of your PDF—such as scanned images, nested headers, or multi-page tables—you can choose between three distinct processing modes.

Fast Mode

Powered by

Gemini 2.0 Flash

Fast Mode is optimized for near-instant results on digital-first PDFs. It is ideal for high-volume processing where the document layout is standardized and the text is already selectable.

Best for:

  • • Single-page invoices
  • • Simple spreadsheets exported to PDF
  • • High-speed data scraping

Avoid when:

  • • Processing blurry photos or scans
  • • Dealing with complex nested headers

Balanced (Recommended)

Powered by

Gemini 2.5 Flash

The default setting for most Supabind users. This model offers a significant upgrade in spatial reasoning, allowing it to correctly identify table boundaries that span across multiple pages without losing row alignment.

Why we recommend it: It provides 90% of the accuracy of the Pro model at a fraction of the processing time and token cost.

High Accuracy

Powered by

Gemini 2.5 Pro

Our most advanced document processing engine. High Accuracy mode utilizes deep reasoning to handle document layouts that defeat standard OCR engines.

Low-Res Scans

Deciphers text even from handheld photos of spec sheets.

Complex Schema

Handles multi-level headers and merged cells with precision.

Confidence & Validation

Every table extracted by our AI includes a **Confidence Score**. This is visible in the extraction results panel (e.g., "Confidence: 94%").

What the score means:

  • 90%+ : The AI is highly certain about the table structure and content.
  • 70% - 89% : Common in low-res scans. We recommend manually reviewing the "My Jobs" editor for cell typos.
  • Under 70% : The layout was highly ambiguous. Consider switching to High Accuracy mode and reprocessing.

Understanding Token Usage

Processing documents via AI consumes "tokens." The total token count is a combination of the input (PDF size and page count) and the output (number of rows and columns extracted).

You can view the exact Input Tokens, Output Tokens, and Total Tokens used for any job by clicking the "View Logs" button in the Processing History tab.

Was this guide helpful?

Complex Document Layout?

If your PDFs are highly unique or very low resolution, our data team can help you calibrate the AI models for your specific use case.

Contact Engineering Support