Skip to main content

Overview

Every processed document receives a quality score (0–100) and a grade (A–D). These help you filter high-quality records and identify documents that may need attention before inclusion in a dataset.

Quality grades


Factors that affect the score

  • Field fill rate — How many expected fields for the document type were successfully extracted
  • Noise ratio — Proportion of the document that is non-informative content (e.g., page numbers, repeated headers)
  • OCR confidence — For scanned documents, a low OCR confidence score caps the grade at C

Accessing quality data

Quality information is available in the job response:
ocr_confidence is null when OCR was not used (text-based PDF or DOCX).

Quality trend

Track quality over time on the Usage → Quality tab in the platform, or via the API:

Filtering by grade

Use the platform’s Datasets → Ready to Build tab to filter completed jobs by grade before building a dataset. Or filter in your own pipeline:

Leaving feedback

If a result looks wrong, you can submit feedback directly from the Jobs page (thumbs up / thumbs down) or via the API: