Skip to main content

Overview

A dataset is a curated collection of pipeline execution results — structured fields, extracted text, and metadata — ready for LLM training, RAG, or analytics. Building and exporting datasets does not consume credits.

Build a dataset

After your documents are processed (job status completed), build a dataset:

Export formats


Export a dataset

RAG export with quality filtering

Pass min_quality to exclude low-quality chunks before export:

Dataset profile

Get aggregate statistics for a dataset:
Returns quality grade distribution, PII type summary, average score, and available export formats.

Chunks API

Available on Pro and Enterprise plans.
Retrieve individual text chunks with quality and PII filters:
Query parameters: Response:
See the RAG pipeline guide for a complete integration walkthrough.

Semantic indexing

Available on Pro and Enterprise plans.
Index a dataset for semantic search:
Then search:

Dataset retention

Datasets are stored for a period defined by your plan: