# Classify Document
Source: https://docs.unsiloed.ai/docs/api-reference/classification/classify-document
api-reference/openapi.json POST /classify
Classify PDF and image documents into predefined categories with confidence scoring using job-based processing
## Overview
The Classify Document endpoint analyzes documents and assigns them to predefined categories based on their content, structure, and visual characteristics. This endpoint uses job-based processing where files are uploaded to cloud storage and processed asynchronously.
The endpoint supports both single-page and multi-page classification with detailed confidence scoring for each page. Documents longer than 4 pages are classified from their first 4 pages. Files are uploaded to cloud storage and processed in the background.
## Request
The document to classify: a PDF or an image (JPG, PNG, GIF, WebP, BMP, TIFF). Either `pdf_file` or `file_url` must be provided, but not both. Maximum file size: 500MB.
URL to a document to classify (PDF or image). Either `pdf_file` or `file_url` must be provided, but not both.
JSON string containing an array of category objects with a required name (e.g., `[{"name":"invoice"},{"name":"contract"}]`). A `description` key is accepted for compatibility but is not used by classification; only the names guide the result.
## Response
Unique identifier for the classification job
Current status of the job ("processing")
Human-readable status message
Remaining API quota after this request
## Examples
```bash cURL theme={null}
curl -X POST "https://prod.visionapi.unsiloed.ai/classify" \
-H "api-key: your-api-key" \
-F "pdf_file=@document.pdf" \
-F 'categories=[{"name":"invoice"},{"name":"contract"},{"name":"receipt"}]'
```
```python Python theme={null}
import requests
import json
url = "https://prod.visionapi.unsiloed.ai/classify"
headers = {"api-key": "your-api-key"}
categories = [
{"name": "invoice"},
{"name": "contract"},
{"name": "receipt"}
]
files = {"pdf_file": ("document.pdf", open("document.pdf", "rb"), "application/pdf")}
data = {"categories": json.dumps(categories)}
response = requests.post(url, headers=headers, files=files, data=data)
if response.status_code == 202:
result = response.json()
print(f"Job ID: {result['job_id']}")
print(f"Status: {result['status']}")
print(f"Quota Remaining: {result['quota_remaining']}")
else:
print("Error:", response.status_code, response.text)
```
```javascript JavaScript theme={null}
const formData = new FormData();
formData.append('pdf_file', fileInput.files[0]);
const categories = [
{name: 'invoice'},
{name: 'contract'},
{name: 'receipt'}
];
formData.append('categories', JSON.stringify(categories));
const response = await fetch('https://prod.visionapi.unsiloed.ai/classify', {
method: 'POST',
headers: {
'api-key': 'your-api-key'
},
body: formData
});
if (response.status === 202) {
const result = await response.json();
console.log('Job ID:', result.job_id);
console.log('Status:', result.status);
console.log('Quota Remaining:', result.quota_remaining);
} else {
console.error('Error:', response.status, await response.text());
}
```
```json 202 - Success theme={null}
{
"job_id": "7e10f5cd-6d83-4b8e-9f80-ac103fd9b2f1",
"status": "processing",
"message": "Classification started",
"quota_remaining": 450
}
```
```json 400 - Bad Request theme={null}
{
"detail": "Either pdf_file or file_url must be provided"
}
```
```json 402 - Payment Required theme={null}
{
"detail": {
"message": "Insufficient credits. Please contact hello@unsiloed.ai to top up.",
"quota_data": {
"remaining": 0,
"message": "..."
}
}
}
```
## Job Status Checking
After starting a classification job, you can check its status using the job ID by making a GET request to `/classify/{job_id}`.
## Job Status Values
* **`queued`**: Job is waiting to be picked up for processing
* **`processing`**: Job is currently being processed
* **`completed`**: Job completed successfully with results available
* **`failed`**: Job failed with error details
# Get Classification Result
Source: https://docs.unsiloed.ai/docs/api-reference/classification/get-classification-status
api-reference/openapi.json GET /classify/{job_id}
Check the status and progress of classification jobs and retrieve results
## Overview
The Get Classification Job Status endpoint allows you to check the current status of classification jobs and retrieve the final results once processing is complete. Classification jobs process documents asynchronously, uploading files to cloud storage and analyzing them in the background.
Status checks are lightweight and can be polled frequently to monitor progress.
## Path Parameters
The unique identifier of the classification job
## Response
Unique identifier for the classification job
Current job status: "queued" while the job waits to be picked up, "processing" while it runs, then "completed" or "failed"
Human-readable progress message describing current processing stage
Error message (if job failed, otherwise null)
Classification results (only present when status is "completed")
Whether the classification operation succeeded
Dominant document-level classification. For single-page inputs this is the page's category. For multi-page inputs it is the category that appears on the most pages; ties are broken by mean per-page confidence. Pages whose individual confidence is below the internal threshold are ignored when picking the dominant category. For mixed-category documents, read `categories[]` to see every category present rather than relying on `classification` alone.
Confidence of the dominant classification (0.0–1.0). For single-page inputs this matches the page's confidence. For multi-page inputs it is the **mean confidence of the pages that voted for `classification`** — not the fraction of pages that agreed, and not a document-wide certainty score. A document where 3 pages classify as `academic_paper` at \~0.96 each yields `confidence ≈ 0.96`, regardless of how the remaining pages classified. For per-page confidence, read `page_results[].confidence`; for the full multi-category breakdown, read `categories[]`.
Every category present in the document, sorted by prominence (page count, then mean confidence). Use this — not `classification` — when you need the full picture of a mixed-category document.
The category label (one of the inputs from `categories` in the request).
Page numbers (1-indexed, ascending) where this category was the per-page classification with confidence above the threshold.
Length of `pages` — how many pages voted for this category.
Mean per-page confidence across `pages` for this category (0.0–1.0).
Number of pages classified. At most 4: documents longer than 4 pages are classified from their first 4 pages.
Number of pages attempted (equals total\_pages, at most 4); per-page success is reported in page\_results
Per-page classification details
Page number (1-indexed)
Classification result for this page
Confidence score for this page (0.0-1.0)
Whether this page was classified successfully
The model's raw classification output for this page
For image inputs the result has neither `page_results` nor `categories`; it carries `classification`, `confidence`, and `raw_result` at the top level with `total_pages: 1`. When an individual PDF page fails, its `page_results` entry has `success: false`, an `error` message, no `raw_result`, and the classification defaults to the first category.
## Request Examples
```bash cURL theme={null}
curl -X GET "https://prod.visionapi.unsiloed.ai/classify/47c536aa-9fab-48ca-b27c-2fd74d30490a" \
-H "api-key: your-api-key" \
-H "Content-Type: application/json"
```
```python Python theme={null}
import requests
job_id = "47c536aa-9fab-48ca-b27c-2fd74d30490a"
url = f"https://prod.visionapi.unsiloed.ai/classify/{job_id}"
headers = {
"api-key": "your-api-key",
"Content-Type": "application/json"
}
response = requests.get(url, headers=headers)
if response.status_code == 200:
result = response.json()
print(f"Job ID: {result['job_id']}")
print(f"Status: {result['status']}")
print(f"Progress: {result.get('progress', 'N/A')}")
if result['status'] == 'completed':
classification_result = result['result']
print(f"Classification: {classification_result['classification']}")
print(f"Confidence: {classification_result['confidence']:.2f}")
print(f"Total pages: {classification_result['total_pages']}")
# Page-by-page results
for page_result in classification_result['page_results']:
print(f"Page {page_result['page']}: {page_result['classification']} (confidence: {page_result['confidence']:.2f})")
elif result['status'] == 'failed':
print(f"Error: {result.get('error', 'Unknown error')}")
else:
print("Job is still processing...")
else:
print("Error:", response.status_code, response.text)
```
```javascript JavaScript theme={null}
const jobId = '47c536aa-9fab-48ca-b27c-2fd74d30490a';
const response = await fetch(`https://prod.visionapi.unsiloed.ai/classify/${jobId}`, {
method: 'GET',
headers: {
'api-key': 'your-api-key',
'Content-Type': 'application/json'
}
});
if (response.ok) {
const result = await response.json();
console.log('Job ID:', result.job_id);
console.log('Status:', result.status);
console.log('Progress:', result.progress || 'N/A');
if (result.status === 'completed') {
const classificationResult = result.result;
console.log('Classification:', classificationResult.classification);
console.log('Confidence:', classificationResult.confidence);
console.log('Total pages:', classificationResult.total_pages);
// Display page-by-page results
classificationResult.page_results.forEach(pageResult => {
console.log(`Page ${pageResult.page}: ${pageResult.classification} (${pageResult.confidence.toFixed(2)})`);
});
} else if (result.status === 'failed') {
console.error('Error:', result.error || 'Unknown error');
} else {
console.log('Job is still processing...');
}
} else {
console.error('Request failed:', response.status, await response.text());
}
```
## Response Examples
```json Processing Status theme={null}
{
"job_id": "47c536aa-9fab-48ca-b27c-2fd74d30490a",
"status": "processing",
"progress": "Starting classification...",
"error": null,
"result": null
}
```
```json Completed Status theme={null}
{
"job_id": "47c536aa-9fab-48ca-b27c-2fd74d30490a",
"status": "completed",
"progress": "Classification completed",
"error": null,
"result": {
"success": true,
"classification": "invoice",
"confidence": 0.9999996871837232,
"categories": [
{"category": "invoice", "pages": [1], "page_count": 1, "confidence": 0.9999996871837232}
],
"page_results": [
{
"page": 1,
"success": true,
"confidence": 0.9999996871837232,
"raw_result": "invoice",
"classification": "invoice"
}
],
"total_pages": 1,
"processed_pages": 1
}
}
```
```json Multi-Page Completed Status (pages agree) theme={null}
{
"job_id": "47c536aa-9fab-48ca-b27c-2fd74d30490a",
"status": "completed",
"progress": "Classification completed",
"error": null,
"result": {
"success": true,
"classification": "contract",
"confidence": 0.89,
"categories": [
{"category": "contract", "pages": [1, 2], "page_count": 2, "confidence": 0.89}
],
"page_results": [
{
"page": 1,
"success": true,
"confidence": 0.92,
"raw_result": "contract",
"classification": "contract"
},
{
"page": 2,
"success": true,
"confidence": 0.86,
"raw_result": "contract",
"classification": "contract"
}
],
"total_pages": 2,
"processed_pages": 2
}
}
```
```json Multi-Page Completed Status (mixed-category document) theme={null}
// 3 pages classify as academic_paper, 1 as presentation_slide.
// Top-level classification: "academic_paper" (most pages).
// Top-level confidence: 0.96 = mean of the 3 academic_paper page confidences.
// (NOT the fraction of pages agreeing; NOT 3/4.)
// categories[] surfaces both categories present — use it for mixed documents
// instead of relying on `classification` alone.
{
"job_id": "47c536aa-9fab-48ca-b27c-2fd74d30490a",
"status": "completed",
"progress": "Classification completed",
"error": null,
"result": {
"success": true,
"classification": "academic_paper",
"confidence": 0.96,
"categories": [
{"category": "academic_paper", "pages": [1, 2, 3], "page_count": 3, "confidence": 0.96},
{"category": "presentation_slide", "pages": [4], "page_count": 1, "confidence": 0.91}
],
"page_results": [
{"page": 1, "success": true, "confidence": 0.98, "raw_result": "academic_paper", "classification": "academic_paper"},
{"page": 2, "success": true, "confidence": 0.96, "raw_result": "academic_paper", "classification": "academic_paper"},
{"page": 3, "success": true, "confidence": 0.94, "raw_result": "academic_paper", "classification": "academic_paper"},
{"page": 4, "success": true, "confidence": 0.91, "raw_result": "presentation_slide", "classification": "presentation_slide"}
],
"total_pages": 4,
"processed_pages": 4
}
}
```
```json Failed Status theme={null}
{
"job_id": "47c536aa-9fab-48ca-b27c-2fd74d30490a",
"status": "failed",
"progress": "Classification failed",
"error": "This PDF could not be opened — the file may be corrupt or in an unsupported format.",
"result": null
}
```
```json Job Not Found theme={null}
{
"detail": "Job not found"
}
```
## Job Status Values
Job has been created and is waiting to be picked up for processing. Jobs move to processing as soon as a worker is available.
Job is currently being processed. This includes file upload to storage, document analysis, and classification processing. Progress messages will indicate the current stage.
Job has completed successfully. The result field contains the classification results with confidence scores and page-by-page details.
Job failed during processing. The error field contains details about what went wrong. Common causes include file corruption, invalid conditions, or processing errors.
## Polling Strategy
For long-running classification jobs, implement polling with exponential backoff:
```python theme={null}
import time
import asyncio
async def poll_classification_status(job_id, max_wait_time=300):
"""Poll classification job status with exponential backoff"""
base_delay = 1 # Start with 1 second
max_delay = 30 # Maximum delay between polls
current_delay = base_delay
total_wait_time = 0
while total_wait_time < max_wait_time:
try:
response = requests.get(
f"https://prod.visionapi.unsiloed.ai/classify/{job_id}",
headers={"api-key": "your-api-key", "Content-Type": "application/json"}
)
if response.status_code == 200:
result = response.json()
if result['status'] == 'completed':
return result['result']
elif result['status'] == 'failed':
raise Exception(f"Classification failed: {result.get('error', 'Unknown error')}")
else:
print(f"Status: {result['status']}, Progress: {result.get('progress', 'N/A')}")
# Wait before next poll
await asyncio.sleep(current_delay)
total_wait_time += current_delay
# Exponential backoff
current_delay = min(current_delay * 2, max_delay)
except Exception as e:
print(f"Error polling status: {e}")
await asyncio.sleep(current_delay)
total_wait_time += current_delay
raise Exception("Classification job timed out")
```
## Progress Messages
The progress field carries one of three messages:
* **"Starting classification..."** - Job has been picked up and is being processed
* **"Classification completed"** - Job finished successfully
* **"Classification failed"** - Job encountered an error (jobs that fail before classification starts may retain "Starting classification...")
## Error Handling
### Common Error Scenarios
1. **File Processing Errors**: PDF corruption, password protection, or unreadable content (job fails with the reason in `error`)
2. **Model Errors**: Vision model failures or timeouts during page classification
3. **Job Not Found**: Invalid job ID or job has been deleted (HTTP 404)
Input validation errors (missing file, malformed categories, oversize files) are rejected synchronously by `POST /classify` with a 400/413 response, so they never appear as failed jobs here.
### Error Response Example
```json File Processing Error theme={null}
{
"job_id": "47c536aa-9fab-48ca-b27c-2fd74d30490a",
"status": "failed",
"progress": "Classification failed",
"error": "This PDF could not be opened — the file may be corrupt or in an unsupported format.",
"result": null
}
```
# Extract Data
Source: https://docs.unsiloed.ai/docs/api-reference/extraction/extract-data
api-reference/openapi.json POST /v2/extract
Extract structured data from PDF documents using custom schemas
## Overview
The `/v2/extract` endpoint extracts structured data from PDF documents. It supports optional bounding box citations and handles large documents efficiently.
The endpoint returns a job ID for asynchronous processing. Use the job management endpoints to check status and retrieve results.
## Request
The PDF file to process for data extraction. Maximum file size: 500MB. Either `pdf_file` or `file_url` must be provided.
URL to a PDF file to process. Either `pdf_file` or `file_url` must be provided.
JSON schema defining the structure and fields to extract from the document. Must be a valid JSON Schema format with type definitions, properties, and required fields.
Model tier to use for extraction. Available tiers: `alpha`, `beta`, `gamma`, `delta`.
**Recommended: `gamma`** (default), the thorough tier. Tiers: `alpha` (fast), `beta` (balanced), `gamma` (thorough), `delta` (advanced).
Return bounding box coordinates for extracted values. When enabled, each extracted field includes a `citation` object with precise location data in the source document.
Enable an agentic verification and validation layer for higher accuracy and consistency. Extracted values are independently cross-checked against the document, and only fields flagged as uncertain are re-verified and reconciled before the result is returned. Slower; off by default. Set to `true` to enable per request.
Optional name identifying the schema, stored with the job for traceability.
Run a pre-flight PII check before extraction. If PII is found at or above `pii_block_severity`, no extraction job is created and the endpoint returns HTTP 200 with `"status": "pii_blocked"` and `"blocked": true` in the body, so check those fields before polling.
PII detection engine used when `detect_pii` is enabled: `standard` (default) or `advanced` (higher precision).
Severity threshold that blocks extraction when `detect_pii` is enabled: `any`, `low`, `medium`, or `high`.
## Response
Unique identifier for the extraction job
Initial job status (typically "queued")
Descriptive message about the job creation
Number of Credits remaining in your quota
### PII-Blocked Response
When `detect_pii` is enabled and PII at or above the threshold is found, the endpoint returns HTTP 200 with no job created. Check `status` and `blocked` to tell this apart from a normal submission. The `reason` is `"pii_detected"` when `pii_block_severity` is `any`, or `"pii_detected:"` for the other thresholds:
```json theme={null}
{
"job_id": null,
"status": "pii_blocked",
"blocked": true,
"reason": "pii_detected:high",
"severity": "high",
"engine": "standard",
"ocr_used": true,
"total_findings": 4,
"by_category": {"financial": 2, "contact": 2},
"by_type": {"CREDIT_CARD": 2, "EMAIL_ADDRESS": 2},
"summary": null,
"spans": []
}
```
## Extraction Results Format
Once the job is completed, the results will contain the extracted data with additional metadata:
Each extracted field returns an object with the following structure:
The extracted value matching the schema type
Confidence scores for the field:
Confidence (0-1) that the value was located in the document. `0.0` when citations are disabled.
Confidence (0-1) in the extracted value itself.
Where the value was found. `null` when citations are disabled or the value could not be grounded:
Bounding box `[left, top, right, bottom]` in PDF point space (origin: top-left).
Page number where the value was found (1-indexed).
Page width in PDF points.
Page height in PDF points.
### Example Extraction Results
For a simple financial document schema:
```json Schema theme={null}
{
"type": "object",
"properties": {
"Individuals": {
"type": "string",
"description": "Percentage Holding"
},
"LIC of India": {
"type": "string",
"description": "No of Shares Held"
},
"United bank of india": {
"type": "string",
"description": "No of shares held by United bank of india"
}
},
"required": [
"Individuals",
"LIC of India",
"United bank of india"
],
"additionalProperties": false
}
```
```json Extraction Results theme={null}
{
"Individuals": {
"value": "10.57",
"score": {
"grounding_score": 0.97,
"extraction_score": 0.99
},
"citation": {
"bbox": [79, 381, 524, 565],
"page": 2,
"page_width": 612,
"page_height": 792
}
},
"LIC of India": {
"value": "1515000",
"score": {
"grounding_score": 0.98,
"extraction_score": 0.99
},
"citation": {
"bbox": [79, 381, 524, 565],
"page": 2,
"page_width": 612,
"page_height": 792
}
},
"United bank of india": {
"value": "500000",
"score": {
"grounding_score": 0.99,
"extraction_score": 0.99
},
"citation": {
"bbox": [79, 381, 524, 565],
"page": 2,
"page_width": 612,
"page_height": 792
}
}
}
```
```bash cURL theme={null}
curl -X POST "https://prod.visionapi.unsiloed.ai/v2/extract" \
-H "accept: application/json" \
-H "api-key: your-api-key" \
-H "Content-Type: multipart/form-data" \
-F "pdf_file=@document.pdf;type=application/pdf" \
-F "schema_data={\"type\":\"object\",\"properties\":{\"title\":{\"type\":\"string\",\"description\":\"Document title\"},\"date\":{\"type\":\"string\",\"description\":\"Document date\"}},\"required\":[\"title\",\"date\"],\"additionalProperties\":false}"
```
```python Python theme={null}
import requests
import json
url = "https://prod.visionapi.unsiloed.ai/v2/extract"
headers = {
"accept": "application/json",
"api-key": "your-api-key"
}
# Define extraction schema using JSON Schema format
schema = {
"type": "object",
"properties": {
"title": {
"type": "string",
"description": "Document title"
},
"date": {
"type": "string",
"description": "Document date"
}
},
"required": ["title", "date"],
"additionalProperties": False
}
# Prepare file and schema data
files = {
"pdf_file": ("document.pdf", open("document.pdf", "rb"), "application/pdf")
}
data = {
"schema_data": json.dumps(schema)
}
# Optional: Enable citations
# data["enable_citations"] = "true"
# Optional: Select model tier (default: gamma)
# data["model"] = "gamma"
# Optional: Agentic verification & validation layer for higher accuracy (slower)
# data["high_accuracy_mode"] = "true"
response = requests.post(url, headers=headers, files=files, data=data)
if response.status_code == 200:
result = response.json()
print(f"Extraction job created: {result['job_id']}")
print(f"Status: {result['status']}")
print(f"Quota remaining: {result['quota_remaining']}")
else:
print("Error:", response.status_code, response.text)
files["pdf_file"][1].close()
```
```javascript JavaScript theme={null}
const formData = new FormData();
formData.append('pdf_file', fileInput.files[0]);
const schema = {
type: "object",
properties: {
title: {
type: "string",
description: "Document title"
},
date: {
type: "string",
description: "Document date"
}
},
required: ["title", "date"],
additionalProperties: false
};
formData.append('schema_data', JSON.stringify(schema));
// Optional: Enable citations
// formData.append('enable_citations', 'true');
// Optional: Select model tier (default: gamma)
// formData.append('model', 'gamma');
// Optional: Agentic verification & validation layer for higher accuracy (slower)
// formData.append('high_accuracy_mode', 'true');
const response = await fetch('https://prod.visionapi.unsiloed.ai/v2/extract', {
method: 'POST',
headers: {
'accept': 'application/json',
'api-key': 'your-api-key'
},
body: formData
});
if (response.ok) {
const result = await response.json();
console.log(`Extraction job created: ${result.job_id}`);
console.log(`Status: ${result.status}`);
console.log(`Quota remaining: ${result.quota_remaining}`);
// Poll for job completion using GET /extract/{job_id}
} else {
console.error('Extraction failed:', response.status, await response.text());
}
```
```json Success Response theme={null}
{
"job_id": "945b4578-691f-4c74-8184-dde654093b11",
"status": "queued",
"message": "PDF citation processing started",
"quota_remaining": 48988
}
```
```json Error Response theme={null}
{
"detail": "Either pdf_file or file_url must be provided"
}
```
## Citations
The `enable_citations` parameter controls whether bounding box coordinates are returned with extracted data. Citations provide references back to the source document, allowing you to trace where each extracted value was found.
### With Citations Enabled
When `enable_citations` is set to `True`, each extracted field includes a `citation` object with precise location data (or `null` when the value could not be grounded):
```json theme={null}
{
"invoice_number": {
"value": "INV-2025-001",
"score": {
"grounding_score": 0.97,
"extraction_score": 0.99
},
"citation": {
"bbox": [139, 209, 280, 222],
"page": 1,
"page_width": 595,
"page_height": 842
}
}
}
```
**Bbox coordinate system:**
* `bbox`: `[left, top, right, bottom]` in PDF point space (origin: top-left)
* Standard A4 page = 595 x 842 points
* `page_width` / `page_height` included for scaling to any display size
### Without Citations (Default)
When `enable_citations` is `False` (default), no grounding pass runs. On the default `gamma` tier each field still uses the nested shape, with `grounding_score: 0.0` and `citation: null`. On the `alpha`, `beta`, and `delta` tiers the response uses a legacy flat shape instead: `{"value": ..., "score": }` with no `citation` key. The nested `gamma` shape:
```json theme={null}
{
"invoice_number": {
"value": "INV-2025-001",
"score": {
"grounding_score": 0.0,
"extraction_score": 0.97
},
"citation": null
}
}
```
Set `enable_citations` to `True` when you need to trace extracted values back to their exact location in the document, such as for UI highlighting or audit trails.
## JSON Schema Definition
The `schema_data` parameter must be a valid JSON Schema that defines the structure of data to extract. All schemas must follow the JSON Schema specification with proper type definitions, properties, and constraints.
### Basic Schema Structure
All extraction schemas must include:
* `type`: "object" (root level)
* `properties`: Object defining the fields to extract
* `required`: Array of required field names
* `additionalProperties`: Set to `False` for strict validation
### Financial Document Schema Example
This example demonstrates extracting shareholding patterns and board information from financial documents:
```json theme={null}
{
"type": "object",
"properties": {
"Individuals": {
"type": "string",
"description": "Percentage Holding"
},
"LIC of India": {
"type": "number",
"description": "No of Shares Held"
},
"board of directors": {
"type": "array",
"description": "list of names of board of directors",
"items": {
"type": "object",
"required": [
"names of board of directors"
],
"properties": {
"names of board of directors": {
"type": "string",
"description": "names of all the members of board of directors of ACRE"
}
},
"additionalProperties": false
}
},
"shareholding pattern": {
"type": "array",
"description": "shareholding pattern",
"items": {
"type": "object",
"required": [
"name of shareholders",
"number of shares held"
],
"properties": {
"name of shareholders": {
"type": "string",
"description": "name of the shareholders in ACRE Table"
},
"number of shares held": {
"type": "string",
"description": "numbers of shares held by shareholders in ACRE Table"
}
},
"additionalProperties": false
}
}
},
"required": [
"Individuals",
"LIC of India",
"board of directors",
"shareholding pattern"
],
"additionalProperties": false
}
```
### Advanced Financial Schema Example
This example shows a more complex schema for extracting detailed shareholding information:
```json theme={null}
{
"type": "object",
"properties": {
"shares held by Punjab National bank": {
"type": "string",
"description": "shares held by Punjab National bank"
},
"shares held by IFCI": {
"type": "string",
"description": "shares held by IFCI"
},
"shareholding pattern": {
"type": "object",
"description": "shareholding pattern",
"properties": {
"Percentage holding": {
"type": "array",
"description": "percentage holding of shareholders in ACRE",
"items": {
"type": "string",
"description": "percentage holding of shareholders in ACRE"
}
},
"Name of shareholders": {
"type": "array",
"description": "Names of shareholders in ACRE",
"items": {
"type": "string",
"description": "Names of shareholders in ACRE"
}
}
},
"required": ["Percentage holding", "Name of shareholders"],
"additionalProperties": false
},
"names of board of directors": {
"type": "array",
"description": "list of names of members of board of directors in ACRE",
"items": {
"type": "object",
"properties": {
"names of board of directors": {
"type": "string",
"description": "list of names of members of board of directors in ACRE"
}
},
"required": ["names of board of directors"],
"additionalProperties": false
}
}
},
"required": [
"shares held by Punjab National bank",
"shares held by IFCI",
"shareholding pattern",
"names of board of directors"
],
"additionalProperties": false
}
```
### Citation Extraction Schema
```json theme={null}
{
"type": "object",
"properties": {
"title": {
"type": "string",
"description": "Document title or paper title"
},
"authors": {
"type": "array",
"description": "List of author names",
"items": {
"type": "string"
}
},
"publication_date": {
"type": "string",
"description": "Publication date in YYYY-MM-DD format"
},
"journal_name": {
"type": "string",
"description": "Name of journal or publication venue"
},
"doi": {
"type": "string",
"description": "Digital Object Identifier"
},
"abstract": {
"type": "string",
"description": "Document abstract or summary"
},
"keywords": {
"type": "array",
"description": "Key terms and subject keywords",
"items": {
"type": "string"
}
},
"references": {
"type": "array",
"description": "List of cited references",
"items": {
"type": "string"
}
}
},
"required": ["title", "authors"],
"additionalProperties": false
}
```
### Legal Document Schema
```json theme={null}
{
"type": "object",
"properties": {
"document_type": {
"type": "string",
"description": "Type of legal document (contract, agreement, etc.)"
},
"parties": {
"type": "array",
"description": "Names of parties involved",
"items": {
"type": "object",
"properties": {
"name": {
"type": "string",
"description": "Party name"
},
"role": {
"type": "string",
"description": "Party role (e.g., buyer, seller, contractor)"
}
},
"required": ["name", "role"],
"additionalProperties": false
}
},
"effective_date": {
"type": "string",
"description": "Document effective date"
},
"key_terms": {
"type": "array",
"description": "Important terms and conditions",
"items": {
"type": "string"
}
},
"obligations": {
"type": "array",
"description": "Key obligations and responsibilities",
"items": {
"type": "object",
"properties": {
"party": {
"type": "string",
"description": "Party responsible for the obligation"
},
"obligation": {
"type": "string",
"description": "Description of the obligation"
}
},
"required": ["party", "obligation"],
"additionalProperties": false
}
}
},
"required": ["document_type", "parties", "effective_date"],
"additionalProperties": false
}
```
## JSON Schema Field Types
Text content, single values. Use for names, descriptions, dates as text.
Numeric values, amounts, quantities. Use for counts, percentages, monetary values.
Whole numbers only. Use for counts, IDs, years.
True/false values. Use for yes/no questions, flags.
Lists of items. Must include `items` property defining the type of array elements.
Structured data with nested fields. Must include `properties` defining nested structure.
Null values. Can be combined with other types using array notation: `["string", "null"]`
## Job Management Integration
After creating an extraction job, you can poll for completion using the job status endpoints:
```python theme={null}
import requests
import time
# After creating the extraction job, you receive a job_id
job_id = "945b4578-691f-4c74-8184-dde654093b11"
headers = {
"accept": "application/json",
"api-key": "your-api-key"
}
# Poll for job completion
while True:
response = requests.get(
f"https://prod.visionapi.unsiloed.ai/extract/{job_id}",
headers=headers
)
if response.status_code == 200:
result = response.json()
status = result.get("status", "").lower()
print(f"Job status: {status}")
if status == "completed":
print("Extraction completed!")
print("Extracted data:", result.get("result"))
break
elif status == "failed":
print(f"Job failed: {result.get('error', 'Unknown error')}")
break
else:
print(f"Error checking status: {response.status_code}")
break
time.sleep(5) # Wait 5 seconds before checking again
```
## Advanced Schema Patterns
### Nested Object Structures
For complex documents with hierarchical data:
```json theme={null}
{
"type": "object",
"properties": {
"company_info": {
"type": "object",
"description": "Company identification and basic information",
"properties": {
"name": {
"type": "string",
"description": "Full company name"
},
"ticker": {
"type": "string",
"description": "Stock ticker symbol"
},
"sector": {
"type": "string",
"description": "Business sector"
}
},
"required": ["name"],
"additionalProperties": false
},
"financial_data": {
"type": "object",
"description": "Financial metrics and performance data",
"properties": {
"revenue": {
"type": "number",
"description": "Total revenue"
},
"profit_margin": {
"type": "number",
"description": "Profit margin percentage"
}
},
"required": ["revenue"],
"additionalProperties": false
}
},
"required": ["company_info", "financial_data"],
"additionalProperties": false
}
```
### Array of Complex Objects
For extracting lists of structured data:
```json theme={null}
{
"type": "object",
"properties": {
"transactions": {
"type": "array",
"description": "List of financial transactions",
"items": {
"type": "object",
"properties": {
"date": {
"type": "string",
"description": "Transaction date"
},
"amount": {
"type": "number",
"description": "Transaction amount"
},
"description": {
"type": "string",
"description": "Transaction description"
},
"category": {
"type": "string",
"description": "Transaction category"
}
},
"required": ["date", "amount", "description"],
"additionalProperties": false
}
}
},
"required": ["transactions"],
"additionalProperties": false
}
```
## Error Handling
Invalid JSON schema format or missing required parameters
Invalid or missing API key
Insufficient or expired credits
File size exceeds the 500MB limit
Invalid file format, malformed JSON schema, or processing error
Rate limit exceeded
Server error during processing
# Get Extraction Result
Source: https://docs.unsiloed.ai/docs/api-reference/jobs/results
api-reference/openapi.json GET /extract/{job_id}
Retrieve the results of a completed processing job
## Overview
The Get Job Results endpoint retrieves the processed data from a completed job. This endpoint should only be called after confirming the job status is "completed" using the status endpoint.
Results are only available for completed jobs. Check job status first to ensure processing has finished.
## Path Parameters
The unique identifier of the extraction job
## Response
The response structure depends on the job type (extraction, parsing, classification, etc.).
### Extraction Job Results
Unique identifier for the extraction job
Current status of the job, lowercase: "queued", "processing", "completed", "review", or "failed". The result field is populated when status is "completed" or "review".
Original filename of the uploaded document
Presigned download URL for the uploaded document. Expires roughly an hour after the response is generated; re-issue this request to get a fresh URL.
ISO 8601 timestamp when the job was created
ISO 8601 timestamp when the job was last updated
Job metadata containing:
* `order`: Array of extracted field names in order
* `schema`: The JSON schema used for extraction
* `page_count`: Number of pages in the document
The extracted data matching the provided JSON schema. Present when status is "completed" or "review". Each field contains:
* `value`: The extracted value, matching the schema type
* `score`: Confidence object with:
* `grounding_score`: Confidence (0-1) that the value was located in the document; `0.0` when citations are disabled
* `extraction_score`: Confidence (0-1) in the extracted value itself, or `null`
* `citation`: Where the value was found, or `null` when citations are disabled or the value could not be grounded:
* `bbox`: `[left, top, right, bottom]` in PDF point space (origin: top-left)
* `page`: Page number where the value was found (1-indexed)
* `page_width`: Width of the source page in points
* `page_height`: Height of the source page in points
For **array fields**, the `value` is an array of objects whose sub-fields each carry their own `value`, `score`, and `citation`; the array field itself also carries an aggregated `score` and `citation: null`. Arrays nested below the top level omit the `citation` key entirely.
The nested score/citation shape applies to citation-enabled jobs and to the default `gamma` model tier. Jobs run with `enable_citations=false` on the `alpha`, `beta`, or `delta` tiers return a legacy flat shape instead: each field is `{"value": ..., "score": }` with no `citation` key.
```bash cURL theme={null}
curl -X GET "https://prod.visionapi.unsiloed.ai/extract/{job_id}" \
-H "api-key: your-api-key"
```
```python Python theme={null}
import json
import requests
job_id = "b2094b38-e432-44b6-a5d0-67bed07d5de1"
url = f"https://prod.visionapi.unsiloed.ai/extract/{job_id}"
headers = {"api-key": "your-api-key"}
response = requests.get(url, headers=headers)
if response.status_code == 200:
data = response.json()
print(json.dumps(data, indent=4))
```
```javascript JavaScript theme={null}
const jobId = 'b2094b38-e432-44b6-a5d0-67bed07d5de1';
const response = await fetch(`https://prod.visionapi.unsiloed.ai/extract/${jobId}`, {
headers: {
'api-key': 'your-api-key'
}
});
if (response.ok) {
const result = await response.json();
console.log(result);
} else {
console.error(
`Request failed with status ${response.status}: ${response.statusText}`
);
}
```
```json Extraction Results theme={null}
{
"job_id": "36adb597-3c2c-43e8-a259-410553291f47",
"status": "completed",
"file_name": null,
"file_url": null,
"created_at": "2026-03-10T16:41:58.407237+00:00",
"updated_at": "2026-03-10T16:42:26.232009+00:00",
"metadata": {
"order": ["EIN", "Address", "Officers", "Organisation", "telephone_number"],
"schema": {
"type": "object",
"required": ["EIN", "Address", "Officers", "Organisation", "telephone_number"],
"properties": {
"EIN": {
"type": "string",
"description": "employee identification number"
},
"Address": {
"type": "string",
"description": "Full Address of organisation"
},
"Officers": {
"type": "array",
"items": {
"type": "object",
"required": ["Officers"],
"properties": {
"Officers": {
"type": "string",
"description": "List of officers"
}
},
"additionalProperties": false
},
"description": "List of officers"
},
"Organisation": {
"type": "string",
"description": "Name of organisation"
},
"telephone_number": {
"type": "string",
"description": "telephone number"
}
},
"additionalProperties": false
},
"page_count": 27
},
"result": {
"EIN": {
"value": "02-0624253",
"score": {
"grounding_score": 0.98,
"extraction_score": 0.99
},
"citation": {
"bbox": [441, 131, 520, 147],
"page": 2,
"page_width": 612,
"page_height": 792
}
},
"Address": {
"value": "602 S OGDEN ST DENVER, CO 80209",
"score": {
"grounding_score": 0.97,
"extraction_score": 0.99
},
"citation": {
"bbox": [84, 201, 181, 216],
"page": 2,
"page_width": 612,
"page_height": 792
}
},
"Officers": {
"value": [
{
"Officers": {
"value": "KIMBERLY TROGGIO",
"score": {
"grounding_score": 0.98,
"extraction_score": 0.99
},
"citation": {
"bbox": [0, 446, 107, 465],
"page": 7,
"page_width": 612,
"page_height": 792
}
}
}
],
"score": {
"grounding_score": 0.98,
"extraction_score": 0.99
},
"citation": null
},
"Organisation": {
"value": "GLOBAL HUMANITARIAN EXPEDITIONS",
"score": {
"grounding_score": 0.99,
"extraction_score": 0.99
},
"citation": {
"bbox": [84, 120, 239, 135],
"page": 2,
"page_width": 612,
"page_height": 792
}
},
"telephone_number": {
"value": "(303) 858-8857",
"score": {
"grounding_score": 0.98,
"extraction_score": 0.99
},
"citation": {
"bbox": [441, 185, 533, 200],
"page": 2,
"page_width": 612,
"page_height": 792
}
}
}
}
```
```json Processing theme={null}
{
"job_id": "36b2c5dc-942c-4b20-8451-39764246f9aa",
"status": "processing",
"file_name": "document.pdf",
"file_url": "https://example-bucket.s3.amazonaws.com/...?AWSAccessKeyId=...&Signature=...&Expires=...",
"created_at": "2026-03-10T16:41:58.407237+00:00",
"updated_at": "2026-03-10T16:42:02.110532+00:00",
"metadata": {
"order": ["EIN", "Address"],
"schema": {"type": "object", "properties": {}},
"page_count": 27
}
}
```
## Complete Workflow Example
Here's a complete example of submitting a job, monitoring its progress, and retrieving results:
```python theme={null}
import requests
import time
import json
def process_document_with_results(file_path, schema, api_key):
"""Complete workflow: submit job, wait for completion, get results"""
headers = {"api-key": api_key}
# Step 1: Submit extraction job
files = {"pdf_file": open(file_path, "rb")}
data = {"schema_data": json.dumps(schema)}
response = requests.post(
"https://prod.visionapi.unsiloed.ai/v2/extract",
files=files,
data=data,
headers=headers
)
if response.status_code != 200:
raise Exception(f"Failed to submit job: {response.text}")
job_id = response.json()["job_id"]
print(f"Job submitted: {job_id}")
# Step 2: Poll for completion
while True:
status_response = requests.get(
f"https://prod.visionapi.unsiloed.ai/extract/{job_id}",
headers=headers
)
if status_response.status_code != 200:
raise Exception(f"Failed to check status: {status_response.text}")
job = status_response.json()
status = job["status"]
print(f"Job status: {status}")
if status in ("completed", "review"):
break
elif status == "failed":
raise Exception(f"Job failed: {job.get('error', 'Unknown error')}")
time.sleep(5) # Wait 5 seconds before checking again
# Step 3: Get results
results_response = requests.get(
f"https://prod.visionapi.unsiloed.ai/extract/{job_id}",
headers=headers
)
if results_response.status_code != 200:
raise Exception(f"Failed to get results: {results_response.text}")
return results_response.json()
# Usage example
schema = {
"type": "object",
"properties": {
"company_name": {
"type": "string",
"description": "Name of the company"
},
"total_amount": {
"type": "number",
"description": "Total financial amount"
}
},
"required": ["company_name"],
"additionalProperties": False
}
try:
results = process_document_with_results("document.pdf", schema, "your-api-key")
print("Extraction completed!")
print("Results:", results)
except Exception as e:
print(f"Error: {e}")
```
## Result Data Structure
### Extracted Field Format
Each extracted field in the `result` object contains:
* `value`: The extracted value, matching the schema type (string, number, boolean, or an array of objects for array fields)
* `score`: A confidence object with `grounding_score` (0-1 confidence the value was located in the document; `0.0` when citations are disabled) and `extraction_score` (0-1 confidence in the value itself, or `null`)
* `citation`: A citation object indicating where the value was found, or `null` when citations are disabled or the value could not be grounded
For **array fields**, the `value` is an array of objects whose sub-fields each carry their own `value`, `score`, and `citation`.
### Citation Format
Each citation object contains:
* `bbox`: `[left, top, right, bottom]` coordinates in PDF point space (origin: top-left)
* `page`: The page number where the value was found (1-indexed)
* `page_width`: Width of the source page in points
* `page_height`: Height of the source page in points
### Confidence Scores
* **0.9-1.0**: Very high confidence, extraction is very likely correct
* **0.8-0.9**: High confidence, extraction is likely correct
* **0.7-0.8**: Good confidence, may warrant review for critical applications
* **0.6-0.7**: Medium confidence, should be reviewed
* **Below 0.6**: Low confidence, likely needs manual verification
## Error Handling
The endpoint always returns 200 for an existing job. While the job is queued or processing, the response simply has no `result` field; keep polling until the status is "completed" or "review".
A failed job also returns 200, with status "failed" and the failure reason in the `error` field.
The job ID is invalid, belongs to another organization, or the job has been deleted. Verify you're using the correct job ID.
# Capacity and Backpressure
Source: https://docs.unsiloed.ai/docs/api-reference/limits/capacity-and-backpressure
What a 503 from the API means, how batch extraction counts against your limits, and how to poll jobs responsibly
## Overview
Unsiloed rejects requests for two different reasons, and the status code tells you which one applies:
* **`429 Too Many Requests`** comes from **your organization's policy**. You sent requests faster than your plan's per-second limit. See [Rate Limits](/docs/api-reference/limits/rate-limits).
* **`503 Service Unavailable`** comes from **platform capacity**. The platform is busy right now, whatever your own request rate. It is not caused by your organization's usage.
Both responses include a `Retry-After` header. In both cases the request was rejected before a job was created, so you can safely resubmit it once the wait is over.
Unsiloed does not queue excess requests on your behalf. A request is either accepted and processed, or it is rejected with a `429` or `503` that tells you when to try again. Accepted jobs run to completion.
## Platform busy
Before a new job is accepted, the API checks how many jobs are in progress across the platform. If that number is at its limit, the submission is rejected with a `503` and a `Retry-After` between **5 and 60 seconds**. The busier the platform, the longer the suggested wait.
This check applies to job submissions: [`POST /v2/extract`](/docs/api-reference/extraction/extract-data), [`POST /classify`](/docs/api-reference/classification/classify-document), [`POST /splitter`](/docs/api-reference/splitting/split-document), and batch extraction. Jobs that were already accepted keep running and are not affected.
```http theme={null}
HTTP/1.1 503 Service Unavailable
Content-Type: application/json
Retry-After: 12
```
```json theme={null}
{
"error": {
"code": "service_unavailable",
"message": "Server is overloaded. Retry after 12s.",
"details": {
"reason": "queue_full",
"retry_after": 12
}
}
}
```
`error.details.reason` is `queue_full` and `error.details.retry_after` matches the `Retry-After` header.
## Platform-wide capacity limit
The API also has a platform-wide capacity limit on total request volume across all customers. It protects the service during unusual traffic spikes, and you are unlikely to see it in normal use. When it is reached, any request to the processing and job status endpoints returns a `503` with `Retry-After: 5`.
```http theme={null}
HTTP/1.1 503 Service Unavailable
Content-Type: application/json
Retry-After: 5
```
```json theme={null}
{
"error": {
"code": "service_unavailable",
"message": "Service temporarily at capacity. Please retry."
}
}
```
## Responding to a 503
1. Wait at least the number of seconds in `Retry-After`.
2. Resubmit the same request. No job was created, so this does not cause duplicates.
3. If you get another `503`, keep waiting for `Retry-After` and add backoff with jitter so many clients do not retry at the same moment.
4. Cap the number of retries and surface an error if the platform stays busy.
The retry helper in the [API FAQ](/docs/faq/api#how-do-i-handle-rate-limiting) handles both `429` and `503` this way.
Do not retry a `503` immediately or in a tight loop. It delays recovery for everyone, including your own jobs that are already running.
## Batch extraction
`POST /batch/extract` submits several files in one request. Two rules apply:
* **At most 20 files per batch.** A larger batch is rejected with `400` and the message `At most 20 files per batch`.
* **Each file counts as one extraction request** against your organization's Extraction limit. A batch of 10 files uses the same allowance as 10 separate `POST /v2/extract` calls.
If a batch needs more requests than your remaining Extraction allowance, the whole batch is rejected with a `429` and **no jobs are created**. The files counted before the limit was reached still use up allowance, so wait for `Retry-After` before resubmitting.
If a batch is larger than your plan's Extraction allowance, keep batches at or below that allowance, or submit files individually and pace them.
If the platform is busy, a batch is rejected with the same `503` described in [Platform busy](#platform-busy) before any of its files are counted against your limit.
## Polling job status
Extraction, classification, and splitting run asynchronously: you submit a job, receive a `job_id`, and poll for the result.
* **Poll at a steady interval.** Every 5 seconds is a good default; jobs usually take seconds to minutes, so polling faster does not get you results sooner.
* **Set a time limit.** Stop polling after a maximum number of attempts and treat the job as timed out on your side.
* **Honor `Retry-After`.** Status endpoints have no per-organization rate limit, but they are covered by the platform-wide capacity limit. If a status request returns `503`, wait for `Retry-After` before polling again instead of counting it as a failed job.
* **Spread out many jobs.** When tracking a large number of jobs, stagger the polls instead of checking every job at the same instant.
```python Python theme={null}
import os
import random
import time
import requests
API_KEY = os.environ["UNSILOED_API_KEY"]
BASE_URL = "https://prod.visionapi.unsiloed.ai"
def wait_for_extraction(job_id, poll_seconds=5, max_attempts=60):
"""Poll an extraction job until it completes, fails, or times out (about 5 minutes)."""
for _ in range(max_attempts):
response = requests.get(f"{BASE_URL}/extract/{job_id}", headers={"api-key": API_KEY})
if response.status_code in (429, 503):
retry_after = response.headers.get("Retry-After", "")
time.sleep(int(retry_after) if retry_after.isdigit() else poll_seconds)
continue
response.raise_for_status()
job = response.json()
if job["status"] in ("completed", "review"):
return job
if job["status"] in ("failed", "cancelled"):
raise RuntimeError(job.get("error", f"extraction job {job['status']}"))
time.sleep(poll_seconds + random.uniform(0, 1))
raise TimeoutError(f"Job {job_id} did not finish in time")
```
```javascript JavaScript theme={null}
const API_KEY = process.env.UNSILOED_API_KEY;
const BASE_URL = "https://prod.visionapi.unsiloed.ai";
const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms));
// Poll an extraction job until it completes, fails, or times out (about 5 minutes).
async function waitForExtraction(jobId, { pollMs = 5000, maxAttempts = 60 } = {}) {
for (let attempt = 0; attempt < maxAttempts; attempt++) {
const response = await fetch(`${BASE_URL}/extract/${jobId}`, {
headers: { "api-key": API_KEY },
});
if (response.status === 429 || response.status === 503) {
const retryAfter = Number.parseInt(response.headers.get("Retry-After") ?? "", 10);
await sleep(Number.isNaN(retryAfter) ? pollMs : retryAfter * 1000);
continue;
}
if (!response.ok) throw new Error(`${response.status}: ${await response.text()}`);
const job = await response.json();
if (job.status === "completed" || job.status === "review") return job;
if (job.status === "failed" || job.status === "cancelled") {
throw new Error(job.error || `extraction job ${job.status}`);
}
await sleep(pollMs + Math.random() * 1000);
}
throw new Error(`Job ${jobId} did not finish in time`);
}
```
## Related
Per-plan limits, rate limit headers, and the 429 response
Confirm the API is up and your key works
Process many documents concurrently with a worker pool
Contact the team if 503s persist
# Checking API Health
Source: https://docs.unsiloed.ai/docs/api-reference/limits/checking-api-health
Confirm the Unsiloed API is reachable and your API key works, and what each response means
## Overview
Two quick checks tell you whether a problem is on your side or ours:
1. **Is the API up?** Call the public liveness endpoint. It needs no API key.
2. **Does my key work?** Make a cheap authenticated request, such as fetching the status of a job you already created.
## Check that the API is reachable
`GET /health` returns `200` as long as the API is up and responding. It needs no API key, does not count against your rate limits, and does no processing.
```bash theme={null}
curl -i https://prod.visionapi.unsiloed.ai/health
```
```json theme={null}
{
"status": "ok"
}
```
`/health` only confirms the API is responding. It does not check your API key and it does not tell you whether the platform is busy. Use the authenticated check below for that.
## Check that your API key works
Fetch the status of a job from one of your earlier requests. This is a lightweight read that does not start any processing or use credits, and it has no per-organization rate limit.
```bash theme={null}
curl -i https://prod.visionapi.unsiloed.ai/extract/YOUR_JOB_ID \
-H "api-key: $UNSILOED_API_KEY"
```
Your API key is checked before anything else, so the status code answers the question even if the job ID is wrong:
| Status | Meaning | What to do |
| - | - | - |
| `200` | The API is up and your key is valid. | Nothing. |
| `404` with `error.code` `not_found` | Your key is valid, but no job with that ID belongs to your organization. | Nothing, if you were only checking the key. Otherwise check the job ID. |
| `401` with `error.code` `auth_failed` | The `api-key` header is missing, or the key is invalid. | Check that the header is named `api-key` and contains the full key. If the key was revoked or rotated, use a current one. |
| `403` with `error.code` `forbidden` | Your credentials were recognized, but access is not permitted. | Email [support@unsiloed.ai](mailto:support@unsiloed.ai). |
| `503` with `error.code` `service_unavailable` | The API is up but temporarily at capacity. | Wait for `Retry-After` seconds, then retry. See [Capacity and Backpressure](/docs/api-reference/limits/capacity-and-backpressure). |
## Interpreting the results
| `/health` | Authenticated request | Likely cause |
| - | - | - |
| `200` | `200` or `404` | Everything is working. Look at the specific request that failed. |
| `200` | `401` | The API is up; the problem is your API key. |
| `200` | `503` | The API is up but busy. Retry after `Retry-After` seconds. |
| No response or connection error | No response | Check your network, proxy, and DNS. If they are fine and the API stays unreachable, email [support@unsiloed.ai](mailto:support@unsiloed.ai). |
## Check usage and limits
Credits used and remaining in your billing cycle
Per-plan request limits and the 429 response
What a 503 means and how to retry
# Get Organization Usage
Source: https://docs.unsiloed.ai/docs/api-reference/organization/usage
api-reference/openapi.json GET /org/get_usage
Check your organization's current usage, limits, and remaining quota
## Overview
The Get Organization Usage endpoint allows you to check your organization's current API usage, monthly limits, and remaining quota. This endpoint uses API key authentication and automatically handles billing cycle resets when applicable.
This endpoint requires a valid API key in the request header. The usage information is specific to the organization associated with the API key.
## Authentication
Your organization's API key for authentication. Format: `unsiloed_[hash]`
## Response
Unique identifier for your organization
Name of your organization
Number of credits used in the current billing cycle
Total monthly usage limit (credits) for your organization
Number of credits remaining in the current billing cycle. Calculated as: `usage_limit - current_usage`
Date of the last billing cycle reset in ISO 8601 format (YYYY-MM-DD). Usage resets automatically every 30 days.
# Get Parse Result
Source: https://docs.unsiloed.ai/docs/api-reference/parser/get-parse-job-status
api-reference/parser/openapi-v1.json GET /parse/{job_id}
Check the status and retrieve results of parsing jobs
## Overview
The Get Parse Job Status endpoint allows you to check the current status of parsing jobs and retrieve the complete results when processing is complete. This endpoint is specifically designed for the parsing API and returns comprehensive document analysis including text extraction, image recognition, table parsing, and OCR data.
Parsing jobs are processed asynchronously. Use this endpoint to poll for completion and retrieve results when the job status is "Succeeded".
## Parameters
Job ID returned by `POST /parse`.
Return segment images as base64-encoded data URIs instead of S3 presigned URLs. Defaults to `false`.
Include the `chunks` array in the response. Defaults to `true`.
Return a presigned S3 URL to the raw output JSON file instead of inlining the full response body. Defaults to `false`.
Opt in to receiving file URLs (`pdf_url`, `file_url`, `output_file_url`, segment `image`, `configuration.input_file_url`) in the response. Defaults to `false`, in which case these fields are returned as `null` so the response (and any log of it) does not expose the storage bucket, region, or path. `exports` URLs are always returned regardless of this setting. Can also be set via the `include-url` header.
Comma-separated segment types to keep in the response (alias: `keep_segment_types`). Omit to include every type.
Return the cached cross-page merged result for jobs that were submitted with `merge_tables`. No merging happens at read time; the job's own setting takes precedence, so passing this for a job created without `merge_tables` has no effect. Defaults to `false`.
## Response
Job identifier.
Current job status: `Starting`, `Queued`, `Processing`, `Succeeded`, `Failed`, or `Cancelled`. Jobs created through the [v2 presigned-upload flow](/docs/api-reference/parser/parse-document-v2) can also report `AwaitingUpload` before the file is uploaded.
ISO 8601 timestamp when the job was created.
Citation or job metadata. Populated when `xml_citation` is enabled or from the job record.
ISO 8601 timestamp when processing started. `null` until a worker picks the job up (`Starting` and `Queued`).
ISO 8601 timestamp when processing completed. Present when status is `Succeeded` or `Failed`.
Total number of document chunks. `0` until status is `Succeeded`.
Number of pages in the document. `0` until processing begins; populated while status is `Processing`.
Array of document chunks with segments and extracted content. Empty until status is `Succeeded` (chunk content is only downloaded for succeeded jobs).
Presigned S3 URL to the generated PDF. Returned as `null` unless `include_url=true` is set.
Presigned download URLs for exported file formats. Only present when `export_format` was specified in the parse request and the export has completed. Keys are format names (e.g. `"docx"`), values are presigned S3 URLs valid for 1 hour. If export failed, contains `{"docx_error": "..."}` instead. Always returned with URLs, regardless of `include_url`.
Original file name from the job record.
MIME type of the uploaded file.
S3 URL of the original uploaded file. Returned as `null` unless `include_url=true` is set.
Credits used for this job.
Status detail message, always present (e.g. "Task queued"). Carries the failure reason when status is `Failed`.
Configuration used for this job, mirroring the parameters submitted at creation time.
Whether table merging was enabled for this job.
Two flags change the response shape for succeeded jobs: with `output_file=true` the parsed content is replaced by a `result` object carrying `output_file_url` (a presigned URL to the raw output JSON) instead of inline `chunks`; and a `merge_tables` job whose merged result is already available may return the content nested under an `output` key with a `task_id` instead of the standard top-level fields.
```bash cURL theme={null}
curl -X 'GET' \
'https://prod.visionapi.unsiloed.ai/parse/04a7a6d8-5ef7-465a-b22a-8a98e7104dd9' \
-H 'accept: application/json' \
-H 'api-key: your-api-key'
```
```python Python theme={null}
import requests
import time
def get_parse_job_status(job_id, api_key):
"""Get the status of a parsing job"""
url = f"https://prod.visionapi.unsiloed.ai/parse/{job_id}"
headers = {"api-key": api_key}
response = requests.get(url, headers=headers)
if response.status_code == 200:
job = response.json()
print(f"Job Status: {job['status']}")
print(f"Created: {job['created_at']}")
if job['status'] == 'Succeeded':
print("Job completed successfully!")
print(f"Total chunks: {job['total_chunks']}")
return job
elif job['status'] == 'Failed':
print(f"Job failed: {job.get('message', 'Unknown error')}")
return None
elif job['status'] in ['Starting', 'Processing']:
print("Job is currently being processed...")
return job
else:
print("Error:", response.status_code, response.text)
return None
# Usage
job_id = "04a7a6d8-5ef7-465a-b22a-8a98e7104dd9"
result = get_parse_job_status(job_id, "your-api-key")
```
```javascript JavaScript theme={null}
async function getParseJobStatus(jobId, apiKey) {
const response = await fetch(`https://prod.visionapi.unsiloed.ai/parse/${jobId}`, {
headers: {
'accept': 'application/json',
'api-key': apiKey
}
});
if (response.ok) {
const job = await response.json();
console.log('Job Status:', job.status);
console.log('Created:', job.created_at);
switch (job.status) {
case 'Succeeded':
console.log('Job completed successfully!');
console.log('Total chunks:', job.total_chunks);
return job;
case 'Failed':
console.log('Job failed:', job.message || 'Unknown error');
return null;
case 'Starting':
case 'Processing':
console.log('Job is currently being processed...');
return job;
default:
console.log('Unknown status:', job.status);
return job;
}
} else {
console.error('Failed to get job status:', await response.text());
return null;
}
}
// Usage
const jobId = '04a7a6d8-5ef7-465a-b22a-8a98e7104dd9';
const result = await getParseJobStatus(jobId, 'your-api-key');
```
```json Starting Job theme={null}
{
"job_id": "04a7a6d8-5ef7-465a-b22a-8a98e7104dd9",
"status": "Starting",
"created_at": "2025-10-22T06:51:16.870302Z",
"metadata": {}
}
```
```json Processing Job theme={null}
{
"job_id": "04a7a6d8-5ef7-465a-b22a-8a98e7104dd9",
"status": "Processing",
"created_at": "2025-10-22T06:51:16.870302Z",
"started_at": "2025-10-22T06:51:16.966136Z",
"metadata": {}
}
```
```json Succeeded Job theme={null}
{
"job_id": "04a7a6d8-5ef7-465a-b22a-8a98e7104dd9",
"status": "Succeeded",
"created_at": "2025-10-22T06:51:16.870302Z",
"started_at": "2025-10-22T06:51:16.966136Z",
"finished_at": "2025-10-22T06:57:19.821541Z",
"metadata": {},
"total_chunks": 25,
"page_count": 10,
"pdf_url": "https://s3.us-east-1.amazonaws.com/unsiloed-bucket/output.pdf?...",
"exports": {
"docx": "https://s3.us-east-1.amazonaws.com/unsiloed-bucket/output.docx?..."
},
"file_name": "document.pdf",
"file_type": "application/pdf",
"file_url": "https://s3.us-east-1.amazonaws.com/unsiloed-bucket/input/document.pdf",
"credit_used": 10,
"merge_tables": false,
"configuration": {
"high_resolution": false,
"ocr_strategy": "auto_detection",
"ocr_engine": "UnsiloedHawk",
"layout_analysis": "smart_layout_detection",
"segment_type_naming": "Unsiloed",
"error_handling": "Continue",
"agentic_ocr": null,
"merge_tables": false,
"extract_colors": false,
"extract_links": false,
"extract_strikethrough": false,
"xml_citation": false,
"chunk_processing": {},
"segment_analysis": {}
},
"chunks": [
{
"chunk_id": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
"chunk_length": 128,
"embed": "eternal logo the image displays a logo...",
"segments": [
{
"segment_type": "Picture",
"content": "eternal",
"image": "https://s3.us-east-1.amazonaws.com/unsiloed-bucket/...",
"page_number": 1,
"segment_id": "1c60ecbd-b6da-493c-b3f6-6849337a981f",
"confidence": 0.6062846,
"page_width": 1191.0,
"page_height": 1684.0,
"html": "
The image displays a logo...
",
"markdown": "Eternal Logo\n\nThe image displays a logo...",
"bbox": {
"left": 72.92226,
"top": 62.030334,
"width": 230.36308,
"height": 55.395317
},
"ocr": [
{
"bbox": {
"left": 63.753525,
"top": 5.395447,
"width": 164.45312,
"height": 42.757812
},
"text": "eternal",
"confidence": 0.9999992
}
]
}
]
}
]
}
```
```json Failed Job theme={null}
{
"job_id": "04a7a6d8-5ef7-465a-b22a-8a98e7104dd9",
"status": "Failed",
"created_at": "2025-10-22T06:51:16.870302Z",
"started_at": "2025-10-22T06:51:16.966136Z",
"finished_at": "2025-10-22T06:51:25.123456Z",
"metadata": {},
"message": "Failed to process document: File appears to be corrupted"
}
```
```json Error Response - Unauthorized theme={null}
"Unauthorized"
```
```text Error Response - Forbidden theme={null}
You don't have permission to access this task
```
```json Error Response - Task Not Found theme={null}
{
"error": {
"code": "task_not_found",
"message": "Task 04a7a6d8-5ef7-465a-b22a-8a98e7104dd9 not found",
"details": {
"task_id": "04a7a6d8-5ef7-465a-b22a-8a98e7104dd9"
}
}
}
```
```json Error Response - Server Error theme={null}
{
"error": {
"code": "internal_error",
"message": "Unable to load this parse task right now. Please try again shortly."
}
}
```
## Job Status Values
Job has been created and is waiting to be processed. This is the initial status when a parsing job is first created.
Job was created through the v2 presigned-upload flow and is waiting for the file to be uploaded. It transitions to Queued once the upload completes.
Job is waiting to be picked up by a worker.
Job is currently being processed. This includes PDF parsing, text extraction, image analysis, table detection, and OCR processing.
Job has completed successfully. The response includes the complete analysis results with all extracted data, images, and metadata.
Job failed during processing. Check the message field for details about what went wrong.
Job was cancelled before processing completed.
## Polling Strategy
For long-running parsing jobs, implement a polling strategy to check status periodically:
```python theme={null}
import requests
import time
def poll_parse_job(job_id, api_key, max_wait_time=300, poll_interval=5):
"""Poll a parsing job until completion or timeout"""
start_time = time.time()
headers = {"api-key": api_key}
while time.time() - start_time < max_wait_time:
response = requests.get(
f"https://prod.visionapi.unsiloed.ai/parse/{job_id}",
headers=headers
)
if response.status_code == 200:
job = response.json()
if job['status'] == 'Succeeded':
return job
elif job['status'] == 'Failed':
raise Exception(f"Job failed: {job.get('message', 'Unknown error')}")
elif job['status'] in ['Starting', 'Processing']:
print(f"Job status: {job['status']} - waiting...")
time.sleep(poll_interval)
else:
print(f"Unknown status: {job['status']}")
time.sleep(poll_interval)
else:
print(f"Error checking status: {response.status_code}")
time.sleep(poll_interval)
raise Exception("Job polling timed out")
# Usage
try:
result = poll_parse_job("04a7a6d8-5ef7-465a-b22a-8a98e7104dd9", "your-api-key")
print("Job completed successfully!")
print(f"Total chunks: {result['total_chunks']}")
except Exception as e:
print(f"Error: {e}")
```
## Segment Types
When a job succeeds, the response includes detailed analysis of different document segments:
### Title
Top-level document titles, distinct from section headers.
### SectionHeader
Document headers and titles that define section boundaries.
### Text
Regular text content including paragraphs, sentences, and individual text elements.
### ListItem
Individual items within ordered or unordered lists.
### Table
Tabular data with structured rows and columns.
### Picture
Images and graphics within the document, including logos, charts, and illustrations.
### Caption
Text captions associated with images or figures.
### Formula
Mathematical or chemical formulas detected within the document.
### Footnote
Footnote text appearing at the bottom of a page.
### PageHeader
Recurring header content appearing at the top of pages.
### PageFooter
Recurring footer content appearing at the bottom of pages.
### Page
A full-page segment when the document is processed without fine-grained layout analysis.
Each segment includes:
* **segment\_type**: Type of content detected
* **content**: Extracted text content
* **image**: URL to extracted image (if applicable)
* **page\_number**: Page where the segment appears
* **confidence**: Confidence score for the extraction
* **bbox**: Precise coordinates of the segment
* **html**: HTML-formatted content
* **markdown**: Markdown-formatted content
* **ocr**: Detailed OCR data with individual text elements
* **chart\_data**: Structured chart data (chart type, series, labels, legend) for Picture segments identified as charts, when chart extraction is enabled
* **cell\_references**: Spreadsheet cell-range references (`{sheet, address, ref}`) for Excel segments
* **references**: Research-paper citations attached to the segment, when `xml_citation` is enabled
* **merged\_page\_bboxes**: Per-page bounding boxes for tables merged across page breaks, when `merge_tables` is enabled
## Error Handling
### Common Error Scenarios
1. **Job Not Found**: Invalid or expired job ID returns a `404` response.
2. **Unauthorized**: Missing or invalid API key returns a `401` response.
3. **Forbidden**: Valid API key but no permission to access this task returns a `403` response.
4. **Rate Limiting**: This GET endpoint is not rate limited by the application; only the submit endpoints are. Poll responsibly regardless.
5. **Client-Side Polling Timeout**: The job did not complete within the time your polling logic allows. This is not a server-returned error; implement a reasonable client-side timeout and handle it gracefully.
6. **Server Error**: Internal processing error returns a `500` response.
### Best Practices
* **Polling Frequency**: Check status every 5-10 seconds for long-running jobs
* **Timeout Handling**: Implement reasonable timeouts to prevent infinite polling
* **Error Recovery**: Handle failed jobs gracefully with retry logic
* **API Key Security**: Keep your API key secure and never expose it in client-side code
## Rate Limits
* **Concurrent Jobs**: Limited number of active parsing jobs per API key
* **Request Frequency**: Avoid excessive polling (recommended: 5-10 second intervals)
Check your API plan for specific limits and quotas.
# Parse Document
Source: https://docs.unsiloed.ai/docs/api-reference/parser/parse-document
api-reference/parser/openapi-v1.json POST /parse
Upload a document for parsing and retrieve ordered Markdown chunks with layout metadata.
## Overview
The Parse Document endpoint processes PDFs, images (PNG, JPEG, TIFF), presentations (PPT, PPTX), and word-processing files (DOC, DOCX). It breaks each document into sections with text, images, tables, and optical character recognition (OCR) data. You can provide a document by direct file upload or by URL.
The parse endpoint uploads your document and configuration in a single request:
1. **POST** to `/parse` with your file and configuration: the API uploads the document and creates a parse job.
2. The job is automatically enqueued for processing.
3. **Poll** `GET /parse/{job_id}` to track progress and retrieve results.
Use [Large File Uploads](/docs/api-reference/parser/parse-document-v2) when a file is too large to send directly to `POST /parse`. The large-file flow returns a presigned URL for uploading the file to Unsiloed-managed storage.
## Request
You must provide either `file` or `url`. If both are provided, `file` takes precedence.
Document file to process. Supported formats: PDF, images (PNG, JPEG, TIFF), presentations (PPT, PPTX), and word-processing files (DOC, DOCX). Required if `url` is not provided. Use `POST /parse/excel` for XLS and XLSX workbooks.
Presigned or public URL of the document to fetch and process. Required if `file` is not provided.
Use high-resolution images for cropping and post-processing. Improves OCR accuracy on low-quality scans by enhancing clarity and contrast. Latency penalty: \~2–3 seconds per page. Defaults to `false`.
How the system analyzes and segments document structure.
* `"smart_layout_detection"` **(default)**: Intelligently identifies document structure, headers, sections, and content relationships across the entire document using bounding boxes.
* `"page_by_page"`: Analyzes each page independently as a single segment. Faster for simple documents.
* `"advanced_layout_detection"`: Uses a vision-language model for exhaustive page segmentation. Detects 14 element types (Caption, Footnote, Formula, ListItem, PageFooter, PageHeader, Picture, SectionHeader, Table, Text, Title, KeyValuePair, Signature, Seal). Best for visually complex or unusual layouts.
Which layout-detection model to run.
* `"v1"` **(default)**: The standard 11-class layout model. Unrecognized values also fall back to this.
* `"v2"`: A form-aware 14-class layout model. Detects the same regions as v1 plus `KeyValuePair`, `Signature`, and `Seal`. Use it for forms, applications, and any document where labeled fields matter.
Choose whether OCR runs automatically on detected images or processes all content.
* `"auto_detection"` **(default)**: Intelligently detects bad quality PDFs, scanned documents, and images, then applies OCR only where needed.
* `"force_ocr"`: Runs OCR on the entire document regardless of quality.
OCR engine to use for text recognition:
* `"UnsiloedHawk"` **(default)**: Higher accuracy for complex layouts and mixed content. Unrecognized values also fall back to this engine.
* `"UnsiloedBeta"`: Handles rotated/warped text and irregular bounding boxes.
* `"UnsiloedStorm"`: Enterprise-grade accuracy optimized for 50+ languages.
Per-segment OCR enhancement: re-runs a dedicated agentic OCR model on each detected segment after layout detection for higher accuracy. Omit or leave empty to disable.
* `"standard"`: Good balance of speed and accuracy.
* `"advanced"`: Higher quality, best for complex layouts, rotated text, and mixed-language content.
Detect and preserve strikethrough formatting in HTML and Markdown output. Defaults to `false`.
Detect and combine table segments across page breaks, reconstructing complete table structure by matching headers and columns. Defaults to `false`.
Maximum number of tables per merge group when `merge_tables` is enabled. Groups larger than this are split. Defaults to `20`.
Reorder detected segments to follow the document's natural reading order. Recommended for **multi-column layouts** (newspapers, two- or three-column papers) where the default smart layout occasionally interleaves text across columns. Has no effect when `layout_analysis="page_by_page"`. Defaults to `false`.
Run a PII detection pass before parsing. If PII is found at or above `pii_block_severity`, the task is rejected and no parsing occurs. Defaults to `false`.
Severity threshold at which the task is rejected when `detect_pii` is enabled: `any` (default) blocks on any PII found; `low` blocks on quasi-identifiers (names, dates, locations) or higher; `medium` blocks on contact PII (email, phone) or higher; `high` blocks only on direct identifiers (SSN, passport, credit card). Ignored if `detect_pii` is false.
PII detection engine: `standard` (default) or `advanced` (higher precision, additional processing cost). Ignored if `detect_pii` is false.
JSON array string of segment types to validate and correct using a Vision Language Model, fixing misclassified segments. Example: `["Table", "Formula", "Picture"]`. Defaults to `["Table", "Picture"]`; an empty or unparseable value also falls back to that default, so Table and Picture validation runs even when this field is omitted.
Legacy parameter that validates table segment classifications using a Vision Language Model. Prefer `validate_segments: ["Table"]` instead. Defaults to `false`.
Choose which types of content to include in the parsed output. Comma-separated segment types, or `"all"` to include everything. Defaults to `"all"`.
Available segment types:
* `table`: Tabular data segments
* `picture`: Image and graphic segments
* `formula`: Mathematical equations
* `text`: Regular text content
* `sectionheader`: Section headers
* `title`: Document titles
* `listitem`: List items
* `caption`: Image captions
* `footnote`: Footnotes
* `pageheader`: Page headers
* `pagefooter`: Page footers
* `keyvaluepair`: Key-value pairs (advanced layout detection)
* `signature`: Signatures (advanced layout detection)
* `seal`: Seals and stamps (advanced layout detection)
* `page`: Full-page segments
Examples: `"table"`, `"table,picture"`, `"table,formula"`, `"picture,formula"`.
Extract and hyperlink bibliography citations in the markdown output. PDFs only. Defaults to `false`.
JSON object controlling which fields are included in the response. Set fields to `false` to exclude them and reduce response size. All fields default to `true`. Ignored when `response_profile` is `slim` or `full` (the profile wins).
Available fields:
* `html`: HTML representation of segments
* `markdown`: Markdown representation of segments
* `ocr`: Raw OCR text data with bounding boxes and confidence scores
* `image`: Cropped segment images (base64 encoded)
* `content`: Text content of segments
* `bbox`: Bounding box coordinates
* `confidence`: Confidence scores for segments
* `embed`: Vector embeddings / embed text
* `chart_data`: Extracted chart data for Picture segments identified as charts
Example: `{"html": true, "markdown": true, "ocr": false, "image": false}`.
Response shape selector: `slim`, `full`, or `custom`. Omit to return the full shape.
* `"slim"`: Returns only the essentials per chunk — `embed`, `bbox`, `page_number`, `segment_id`, `segment_type`, and HTML for tables / Markdown for everything else. Drops `content`, `image`, `ocr`, `confidence`, `chart_data`, `page_height`, `page_width`. Best for embedding-only workflows where you want the smallest payload.
* `"full"`: Every field returned (equivalent to omitting this param).
* `"custom"`: Honor `output_fields` verbatim.
When both `response_profile` and `output_fields` are provided, the profile wins — `output_fields` is only consulted for `custom` or when the profile is omitted. Applies to inline JSON responses only; `GET /parse/{job_id}?output_file=true` returns a presigned URL to the stored full-shape output file.
JSON object controlling HTML/Markdown generation strategy and AI model per segment type. Configure how different segment types are processed, including table processing models, image description models, and formula processing.
Example:
```json theme={null}
{
"Table": {"html": "VLM", "markdown": "VLM", "model_id": "us_table_v2"},
"Picture": {"html": "VLM", "markdown": "VLM", "model_id": "nova"},
"Formula": {"html": "Auto", "markdown": "VLM", "model_id": "nova"},
"KeyValuePair": {"html": "Auto", "markdown": "VLM", "model_id": "astra_v2"}
}
```
Options per segment type:
* `html`: `"VLM"` or `"Auto"`
* `markdown`: `"VLM"` or `"Auto"`
* `model_id` (Table): `"astra"`, `"us_table_v1"`, `"us_table_v2"`
* `model_id` (Picture/Formula): `"nova"`, `"luna"`, `"sol"`
* `model_id` (KeyValuePair): `"astra_v2"`
* `use_table_ocr` (Table only): Advanced OCR optimized for tabular data. Better handles bordered cells, gridlines, and complex table layouts.
* `vlm`: Custom prompt for the VLM model. Use this to give the model specific instructions for extracting or describing these segment types.
* `translation`: Optional per-segment translation, e.g. `{"provider": "Auto", "target_language": "en"}`. `provider` is `"Auto"` for fast machine translation or `"VLM"`/`"LLM"` for model-based translation; `target_language` is an ISO 639-1 code, or `"auto"` to auto-detect the source and translate to English. Optional `model_id` and `prompt` apply to model-based translation.
Defaults for `html` and `markdown` are not uniform across segment types. `Table.html`, `Picture.markdown`, `Formula.markdown`, `KeyValuePair.markdown`, and both `Page` strategies default to `"VLM"`; everything else defaults to `"Auto"`. See [Segment Analysis Configuration](#segment-analysis-configuration) below for the full table and override examples.
Alias for `segment_analysis`. If both are provided, `segment_processing` takes precedence.
Specify which pages to process. Formats: `"1-5"`, `"2,4,6"`, `"[1,3,5]"`. Defaults to all pages.
Segment type naming convention. `"Unsiloed"` (default) uses names like `PageHeader`, `ListItem`, `Picture`. `"Other"` uses alternative names like `Header`, `List Item`, `Figure`.
Transfer text color from the PDF text layer to OCR results. Defaults to `false`.
Attach hyperlink URLs from PDF annotations to OCR results. Defaults to `false`.
JSON array string of export formats to generate after processing, e.g. `["docx"]`. When set, the pipeline generates the requested export files after parsing completes. The exported files are available as presigned URLs in the `exports` field of the response. Supported values: `"docx"`, `"markdown"`, `"json"`.
This is a multipart form field, so the value must be a JSON-encoded string (`["docx"]`), not a repeated field. Passing a bare value like `docx` will fail to parse and silently skip the export.
Error handling strategy for non-critical processing errors. `"Continue"` (default) proceeds despite errors (e.g., LLM refusals on individual segments). `"Fail"` stops and fails the task on any error.
Reserved field. Persisted in the task configuration but currently has no effect on retention for `POST /parse` — the task is not auto-deleted. To get a presigned-upload TTL, use [`POST /v2/parse/upload`](/docs/api-reference/parser/parse-document-v2) instead, where `expires_in` controls the upload URL's validity.
JSON object for chunk processing configuration.
JSON object for LLM processing configuration.
## Configuration Best Practices
Click to expand each scenario below to view detailed configuration settings, recommendations, and trade-offs for your specific use case.
Use this configuration when accuracy is critical and processing time is less important.
**Configuration:**
```json theme={null}
{
"use_high_resolution": true,
"layout_analysis": "smart_layout_detection",
"ocr_strategy": "force_ocr",
"merge_tables": true,
"validate_segments": ["Table", "Picture", "Formula"],
"segment_analysis": {
"Table": {
"html": "VLM",
"markdown": "VLM",
"extended_context": true,
"crop_image": "All",
"model_id": "us_table_v2"
}
}
}
```
**When to Use:**
* Legal documents requiring precise text extraction
* Financial statements with complex tables
* Archival documents with low-quality scans
* Documents where accuracy is more important than speed
**Trade-offs:**
* **Latency**: +2-3 seconds per page for high resolution
* **Latency**: +1-2 seconds per page for segment validation
Use this configuration when speed is prioritized over maximum accuracy.
**Configuration:**
```json theme={null}
{
"use_high_resolution": false,
"layout_analysis": "page_by_page",
"ocr_strategy": "auto_detection",
"merge_tables": false
}
```
**When to Use:**
* High-volume document processing
* Real-time applications requiring quick results
* Documents with simple layouts
* Pre-screened high-quality digital documents
**Benefits:**
* Fastest processing time
* Lower cost per document
* Suitable for batch processing large volumes
Extract only tables and charts from financial reports and statements.
**Configuration:**
```json theme={null}
{
"merge_tables": true,
"segment_filter": "table,picture",
"validate_segments": ["Table", "Picture"],
"layout_analysis": "smart_layout_detection",
"ocr_strategy": "auto_detection",
"segment_analysis": {
"Table": {
"html": "VLM",
"markdown": "VLM",
"model_id": "us_table_v2"
}
}
}
```
**When to Use:**
* Balance sheets and P\&L statements
* Quarterly/annual financial reports
* Investment reports with charts
* Documents where only structured data matters
**Benefits:**
* Reduced response size (text content filtered out)
* Focus on data-rich content
* Merged multi-page tables for complete datasets
Extract only tabular data with minimal response size for maximum efficiency.
**Configuration:**
```json theme={null}
{
"merge_tables": true,
"segment_filter": "table",
"validate_segments": ["Table"],
"output_fields": {
"html": true,
"markdown": true,
"ocr": false,
"image": false,
"content": true,
"bbox": false,
"confidence": false
}
}
```
**When to Use:**
* Extracting data from invoices
* Processing structured forms
* Database population from documents
* CSV/Excel export workflows
**Benefits:**
* Minimal response payload
* Faster data transfer
* Easy integration with data pipelines
Extract content with structured citations from research papers and academic documents.
**Configuration:**
```json theme={null}
{
"use_high_resolution": true,
"layout_analysis": "smart_layout_detection",
"ocr_strategy": "auto_detection",
"xml_citation": true
}
```
**When to Use:**
* Research papers with bibliographies
* Academic articles with citations
* Scientific documents
* Literature reviews
**Benefits:**
* Automatic citation extraction and linking
* Structured bibliography metadata
* In-text citation hyperlinks in markdown
* Preserves academic document structure
Citation extraction is only available for PDF documents.
Optimize for scanned documents and images with poor text quality.
**Configuration:**
```json theme={null}
{
"use_high_resolution": true,
"ocr_strategy": "force_ocr",
"layout_analysis": "smart_layout_detection"
}
```
**When to Use:**
* Scanned paper documents
* Low-quality photocopies
* Historical documents
* Image-based PDFs
**Benefits:**
* Maximum OCR coverage
* Better text extraction from poor quality sources
* Higher accuracy for challenging documents
Optimize response size and performance by selectively including only the fields you need.
**For Minimal Response Size:**
```json theme={null}
{
"output_fields": {
"html": false,
"markdown": false,
"ocr": false,
"image": false,
"content": true,
"bbox": false,
"confidence": false,
"embed": true
}
}
```
**For Text-Only Processing:**
```json theme={null}
{
"output_fields": {
"html": false,
"markdown": true,
"ocr": false,
"image": false,
"content": true,
"bbox": true,
"confidence": false,
"embed": true
}
}
```
**For Full Analysis (Default):**
Omit `output_fields` or set all fields to `True` to include all available data.
**Benefits:**
* Reduced response size and bandwidth usage
* Faster processing and data transfer
* Cost optimization for high-volume processing
## Parameter Details
### File Input Options
The API supports two methods for providing the document to process:
1. **Direct File Upload** (`file` parameter): Upload the document file directly as multipart/form-data
2. **Presigned URL** (`url` parameter): Provide a publicly accessible URL or presigned URL to the document
**Important Notes:**
* You must provide **either** `file` or `url`, but not both
* When using `url`, the document will be downloaded from the provided URL before processing
* Presigned URLs are ideal for documents already stored in cloud storage (S3, GCS, Azure Blob, etc.)
* The URL must be publicly accessible or include necessary authentication parameters (e.g., S3 presigned URLs with signatures)
* Supported formats are the same for both methods: PDF, images (PNG, JPEG, TIFF), presentations (PPT, PPTX), and word-processing files (DOC, DOCX). Use `POST /parse/excel` for XLS and XLSX workbooks.
**Use Cases for Presigned URLs:**
* Documents already stored in cloud storage
* Avoiding duplicate file uploads
* Integration with existing document management systems
* Processing documents already available at a URL without uploading them again
### Segmentation Method
The `layout_analysis` parameter controls how the document is analyzed and segmented:
* **`"smart_layout_detection"`** (default): Analyzes pages for layout elements (e.g., Table, Picture, Formula, etc.) using bounding boxes. Provides fine-grained segmentation and better chunking for complex documents.
* **`"page_by_page"`**: Treats each page as a single segment. Faster processing, ideal for simple documents without complex layouts.
* **`"advanced_layout_detection"`**: Uses a vision-language model to exhaustively segment each page into 14 element types (including `KeyValuePair`, `Signature`, and `Seal` in addition to the standard set). Recommended for documents with dense, non-standard, or visually complex layouts where VGT-based detection misses regions.
### Segmentation Version
The `segmentation_version` parameter picks which layout-detection model runs:
* **`"v1"`** (default): The standard layout model. Returns the 11 core element types (Caption, Footnote, Formula, ListItem, PageFooter, PageHeader, Picture, SectionHeader, Table, Text, Title).
* **`"v2"`**: A form-aware model trained jointly on document layout and forms. Returns the same 11 types plus `KeyValuePair`, `Signature`, and `Seal`.
Reach for `v2` when the document is a form, application, or any layout where labeled fields (`Name:`, `Policy No:`, checkbox groups) carry the meaning — under `v1` those regions come back as ordinary `Text` segments, with labels and values often split apart. Everything else about the request stays the same; only the segment types you get back change.
```bash theme={null}
curl -X POST "https://prod.visionapi.unsiloed.ai/parse" \
-H "api-key: $UNSILOED_API_KEY" \
-F "file=@application-form.pdf" \
-F "segmentation_version=v2"
```
### Agentic OCR
The `agentic_ocr` parameter enables per-segment OCR enhancement after layout detection, yielding higher accuracy on small text, stylized fonts, and mathematical formulas.
**Values:**
* `"standard"`: Fast, good for most documents.
* `"advanced"`: Higher quality, better for complex layouts, rotated or irregular text, and multilingual content.
### OCR Mode
The `ocr_strategy` parameter controls optical character recognition processing:
* **`"auto_detection"`** (default): Intelligently determines when OCR is needed based on the document content. Balances accuracy and performance.
* **`"force_ocr"`**: Applies OCR to all content regardless of existing text layer. Use this for scanned documents or when maximum text extraction is required.
### Table Merging
The `merge_tables` parameter enables merging of tables that span across multiple pages:
**How It Works:**
* Analyzes consecutive table segments across pages
* Identifies tables with matching column headers
* Merges them into a single unified table structure
* Preserves table formatting and data integrity
**When to Use:**
* **Multi-Page Financial Statements**: Consolidate P\&L statements or balance sheets spanning multiple pages
* **Large Data Tables**: Merge inventory lists, transaction records, or data sets split across pages
* **Reports with Continuation Tables**: Automatically combine tables marked with "continued on next page"
**Example:**
```json theme={null}
{
"merge_tables": true
}
```
**Benefits:**
* **Simplified Data Processing**: Work with complete tables instead of fragments
* **Better Context**: Maintain full table context for analysis and extraction
* **Reduced Post-Processing**: Eliminates need for manual table stitching
### Citation Extraction (Research Papers)
This is a specialized feature for academic and research PDF documents. Not needed for general document parsing.
The `xml_citation` parameter enables automatic extraction and linking of citations from research papers, academic articles, and scientific documents.
**How It Works:**
* Extracts structured bibliography from the document
* Identifies in-text citation references (e.g., "Chen et al., 2021")
* Hyperlinks citations in the markdown output to their bibliography entries
* Returns structured citation metadata in the response
**Example:**
```json theme={null}
{
"xml_citation": true
}
```
**Response Metadata:**
When enabled, the response includes a `metadata` field with structured citation data:
```json theme={null}
{
"metadata": {
"citations": [
{
"id": 1,
"title": "Deep Learning for NLP",
"authors": ["John Smith", "Jane Doe"],
"year": "2021",
"journal": "Nature",
"volume": "15",
"pages": "123-145",
"doi": "10.1000/example"
}
],
"document_metadata": {
"title": "Document Title",
"authors": ["Author Name"]
}
}
}
```
**Markdown Enhancement:**
In-text citations are automatically hyperlinked:
* Original: `"As shown by Chen et al. (2021)..."`
* Enhanced: `"As shown by [Chen et al. (2021)](#ref-5)..."`
Only available for PDF documents. This parameter is ignored for other file types (images, Office documents).
### Content Type Filtering
The `segment_filter` parameter allows you to filter the output to include only specific segment types, reducing response size and focusing on relevant content:
**How It Works:**
* Accepts a comma-separated list of segment types (case-insensitive)
* Filters segments after processing is complete
* Removes chunks that have no segments after filtering
**Available Options:**
* `"all"` (default): Include all segment types
* `"table"`: Only table segments
* `"picture"`: Only image/graphic segments
* `"table,picture"`: Tables and pictures only
* `"table,formula"`: Tables and formulas only
* Custom combinations using any segment type
**Supported Segment Types:**
* `table`, `picture`, `formula`, `text`, `sectionheader`, `title`, `listitem`, `caption`, `footnote`, `pageheader`, `pagefooter`
**Example Usage:**
```json theme={null}
{
"segment_filter": "table,picture"
}
```
**Use Cases:**
* **Tables Only**: Extract only tabular data from financial documents
* **Pictures Only**: Extract charts, graphs, and diagrams for visual analysis
* **Tables + Pictures**: Get structured data and visualizations, skip text content
* **Custom Combinations**: Mix any segment types based on your needs
**Benefits:**
* **Reduced Response Size**: Filter out unwanted content before receiving results
* **Faster Processing**: Less data to transfer and parse
* **Focused Extraction**: Get only the content types you need
* **Cost Optimization**: Smaller responses reduce bandwidth usage
### Output Fields Configuration
The `output_fields` parameter allows you to control which fields are included in the API response. This is useful for reducing response size, improving performance, and optimizing bandwidth usage when you don't need all available data.
**Available Fields:**
* **`html`** (default: `true`): Include HTML representation of segments
* **`markdown`** (default: `true`): Include Markdown representation of segments
* **`ocr`** (default: `true`): Include OCR results with bounding boxes and confidence scores
* **`image`** (default: `true`): Include cropped segment images (base64 encoded)
* **`content`** (default: `true`): Include text content of segments
* **`bbox`** (default: `true`): Include bounding box coordinates
* **`confidence`** (default: `true`): Include confidence scores for segments
* **`embed`** (default: `true`): Include embed text in chunk responses
**Usage:**
Set fields to `false` to exclude them from the response. Fields not specified default to `true` for backward compatibility.
**Example Configuration:**
```json theme={null}
{
"html": false,
"markdown": true,
"ocr": false,
"image": false,
"content": true,
"bbox": true,
"confidence": false,
"embed": true
}
```
**Benefits:**
* **Reduced Response Size**: Excluding large fields like `image` and `html` can significantly reduce payload size
* **Faster Processing**: Less data to serialize and transfer
* **Cost Optimization**: Smaller responses reduce bandwidth costs
* **Selective Data**: Only retrieve the fields you need for your use case
**When to Use:**
* **Minimal Response**: Set most fields to `false` when you only need basic content
* **Text-Only Processing**: Exclude `image` and `ocr` when processing text content
* **Embedding Generation**: Include only `content` and `embed` when generating embeddings
* **Full Analysis**: Keep all fields enabled (default) for comprehensive document analysis
### Segment Analysis Configuration
The `segment_analysis` parameter allows you to customize how different segment types are processed, including HTML/Markdown generation strategies and which field should populate the `content` field.
**Available Segment Types:**
You can configure processing for any of the following segment types:
* `Table`: Tabular data segments
* `Picture`: Image and graphic segments
* `Formula`: Mathematical equations
* `Title`: Document titles
* `SectionHeader`: Section headers
* `Text`: Regular text content
* `ListItem`: List items
* `Caption`: Image captions
* `Footnote`: Footnotes
* `PageHeader`: Page headers
* `PageFooter`: Page footers
* `Page`: Full page segments
* `KeyValuePair`: Form field regions (labels with their filled-in values). Only produced when `segmentation_version=v2`; the config is accepted either way.
**Configuration Options:**
For each segment type, you can specify:
* **`html`**: Generation strategy for HTML representation
* `"Auto"`: Build HTML directly from the OCR/recognized text (skips the VLM call — faster and cheaper)
* `"VLM"`: Use a VLM to generate HTML
* **`markdown`**: Generation strategy for Markdown representation
* `"Auto"`: Build Markdown directly from the OCR/recognized text (skips the VLM call — faster and cheaper)
* `"VLM"`: Use a VLM to generate Markdown
Defaults are **not uniform** across segment types. Four cases use `VLM` by default so the output is usable without extra configuration:
| Segment type | Default `html` | Default `markdown` |
| - | - | - |
| `Table` | `VLM` | `Auto` |
| `Picture` | `Auto` | `VLM` |
| `Formula` | `Auto` | `VLM` |
| `Page` | `VLM` | `VLM` |
| `KeyValuePair` | `Auto` | `VLM` |
| All other types (`Title`, `Text`, `ListItem`, `Caption`, `Footnote`, `PageHeader`, `PageFooter`, `SectionHeader`) | `Auto` | `Auto` |
If you want a `Picture` segment's Markdown to contain only the transcribed text (rather than a VLM-written description), pass `segment_analysis={"Picture": {"markdown": "Auto"}}`. The same override applies to `Table.html` and `Formula.markdown`. `Auto` is faster and cheaper, but loses VLM summarization for image-heavy segments.
`KeyValuePair` keeps each label with its value in one segment, checkbox and radio state included. It also carries an `image` crop URL when you poll with `include_url=true`.
* **`content_source`**: Defines which field should populate the `content` field in the response
* `"OCR"` (default): Use OCR text for content
* `"HTML"`: Use HTML representation as content
* `"Markdown"`: Use Markdown representation as content
* `"VLM"` (alias `"LLM"`): Use the VLM-generated representation as content
* **`model_id`**: Specifies which AI model to use for the VLM call on this segment type
* Table: `"us_table_v1"` (standard) or `"us_table_v2"` (enhanced accuracy)
* Picture / Formula: `"nova"`, `"luna"`, `"sol"`
* KeyValuePair: `"astra_v2"`
Omit to use the default model for the type.
* **`vlm`**: Custom prompt for the VLM model. Use this to give the model specific instructions for extracting or describing these segment types.
* **`translation`**: Optional per-segment translation configuration:
* `provider`: `"Auto"` (fast machine translation) or `"VLM"`/`"LLM"` (model-based translation)
* `target_language`: ISO 639-1 code (e.g. `"en"`, `"es"`, `"fr"`, `"ko"`), or `"auto"` to auto-detect the source language and translate to English
* `model_id` (optional): model for VLM/LLM translation; defaults to the provider default
* `prompt` (optional): custom instructions appended to the translation system prompt
**Example Configuration:**
```json theme={null}
{
"Table": {
"html": "VLM",
"markdown": "VLM",
"content_source": "HTML",
"model_id": "us_table_v2",
"vlm": "Preserve all merged cells. Use empty strings for missing values."
},
"Picture": {
"html": "VLM",
"markdown": "VLM",
"content_source": "Markdown",
"vlm": "Focus on chart axes, legend labels, and key data trends."
},
"KeyValuePair": {
"html": "VLM",
"markdown": "VLM",
"content_source": "Markdown",
"model_id": "astra_v2"
}
}
```
**How `content_source` Works:**
The `content_source` parameter determines which field's value will be used to populate the `content` field in the segment response:
* When `content_source` is set to `"HTML"`, the `content` field will contain the HTML representation, and the separate `html` and `markdown` fields will be empty
* When `content_source` is set to `"Markdown"`, the `content` field will contain the Markdown representation, and the separate `html` and `markdown` fields will be empty
* When `content_source` is set to `"OCR"` (default), the `content` field contains OCR text, and `html` and `markdown` fields are populated separately
**Use Cases:**
* **HTML as Content**: Set `content_source: "HTML"` for Table segments when you want HTML-formatted table data directly in the `content` field
* **Markdown as Content**: Set `content_source: "Markdown"` for Picture segments when you want Markdown-formatted descriptions in the `content` field
* **VLM-Enhanced Output**: Use `"VLM"` for both `html` and `markdown` generation strategies to get AI-enhanced representations in those fields
## Response
Job identifier. Pass this to `GET /parse/{job_id}` to poll for results.
Initial job status. Always `"Starting"` on creation.
Name of the uploaded file. For URL submissions this is the last path segment of the URL, or `"unknown"` when no usable segment exists.
ISO 8601 timestamp when the job was created.
Human-readable status message with a polling hint.
Number of pages deducted from your quota for this job.
Remaining page quota after this job was deducted.
Whether table merging is enabled for this job (reflects the submitted `merge_tables` value).
## Document Analysis Features
The parsing endpoint provides comprehensive document analysis including:
### Text Extraction
Extracts text content with high accuracy, preserving formatting and structure.
### Image Recognition
Identifies and analyzes images within documents, providing descriptions and metadata.
### Table Parsing
Extracts tabular data with proper structure and formatting.
### OCR Processing
Performs optical character recognition on text elements with confidence scores.
### Section Detection
Automatically identifies different document sections like headers, body text, and captions.
### Bounding Box Information
Provides precise coordinates for all extracted elements.
### Advanced Content Processing
* **VLM-Enhanced Analysis**: Uses vision-language models for better content understanding
* **Multi-Format Output**: Generates HTML, Markdown, and plain text versions
* **Context-Aware Processing**: Maintains document context across segments
* **Intelligent Chunking**: Creates semantically meaningful document chunks
```bash cURL theme={null}
curl -X 'POST' \
'https://prod.visionapi.unsiloed.ai/parse' \
-H 'accept: application/json' \
-H 'api-key: your-api-key' \
-H 'Content-Type: multipart/form-data' \
-F 'file=@document.pdf;type=application/pdf' \
-F 'use_high_resolution=true' \
-F 'layout_analysis=smart_layout_detection' \
-F 'ocr_strategy=auto_detection' \
-F 'ocr_engine=UnsiloedHawk' \
-F 'extract_strikethrough=false' \
-F 'merge_tables=true' \
-F 'enhance_reading_order=false' \
-F 'segment_filter=all' \
-F 'validate_segments=["Table","Picture","Formula"]' \
-F 'export_format=["docx"]' \
-F 'segment_analysis={"Table":{"html":"VLM","markdown":"VLM","extended_context":true,"crop_image":"All","model_id":"us_table_v2"}}'
# Alternative: Use presigned URL instead of file upload
# Replace the file parameter with url parameter:
# -F 'url=https://your-bucket.s3.amazonaws.com/document.pdf?signature=...' \
```
```python Python theme={null}
import requests
import json
url = "https://prod.visionapi.unsiloed.ai/parse"
headers = {
"accept": "application/json",
"api-key": "your-api-key"
}
# Basic file upload with all options
files = {
"file": ("document.pdf", open("document.pdf", "rb"), "application/pdf")
}
data = {
"use_high_resolution": "true",
"layout_analysis": "smart_layout_detection",
"ocr_strategy": "auto_detection",
"ocr_engine": "UnsiloedHawk",
"extract_strikethrough": "false",
"merge_tables": "true",
"enhance_reading_order": "false",
"segment_filter": "all",
"export_format": json.dumps(["docx"]),
"segment_analysis": json.dumps({
"Table": {
"html": "VLM",
"markdown": "VLM",
"extended_context": True,
"crop_image": "All",
"model_id": "us_table_v2"
}
})
}
response = requests.post(url, headers=headers, files=files, data=data)
if response.status_code == 200:
result = response.json()
print(f"Job ID: {result['job_id']}")
print(f"Status: {result['status']}")
print(f"File Name: {result['file_name']}")
print(f"Credit Used: {result['credit_used']}")
print(f"Message: {result['message']}")
else:
print("Error:", response.status_code, response.text)
files["file"][1].close()
# ========== Alternative: Use Presigned URL ==========
# Instead of uploading a file, you can provide a presigned URL
# Remove the 'files' parameter and add 'url' to data:
# data = {
# "url": "https://your-bucket.s3.amazonaws.com/document.pdf?signature=...",
# "use_high_resolution": "true",
# "layout_analysis": "smart_layout_detection",
# "ocr_strategy": "auto_detection",
# "merge_tables": "true",
# "segment_filter": "all"
# }
# response = requests.post(url, headers=headers, data=data)
# ========== Use Case: Extract Tables Only ==========
# For financial reports or documents where you only need tables:
# data = {
# "merge_tables": "true",
# "segment_filter": "table",
# "validate_segments": json.dumps(["Table"]),
# "layout_analysis": "smart_layout_detection",
# "ocr_strategy": "auto_detection"
# }
# ========== Use Case: Citation Extraction (Academic Papers) ==========
# For research papers, enable citation extraction:
# data = {
# "use_high_resolution": "true",
# "layout_analysis": "smart_layout_detection",
# "ocr_strategy": "auto_detection",
# "xml_citation": "true"
# }
# ========== Use Case: Advanced Segment Analysis ==========
# Configure how different segment types are processed:
# segment_analysis_config = {
# "Table": {
# "html": "VLM",
# "markdown": "VLM",
# "content_source": "HTML",
# "model_id": "us_table_v2"
# },
# "Picture": {
# "html": "VLM",
# "markdown": "VLM",
# "content_source": "Markdown"
# }
# }
# data = {
# "use_high_resolution": "true",
# "layout_analysis": "smart_layout_detection",
# "ocr_strategy": "auto_detection",
# "merge_tables": "true",
# "segment_filter": "table,picture",
# "segment_analysis": json.dumps(segment_analysis_config)
# }
```
```javascript JavaScript theme={null}
const formData = new FormData();
// Basic file upload
const fileInput = document.querySelector('input[type="file"]');
if (fileInput.files[0]) {
formData.append('file', fileInput.files[0]);
}
// Add configuration parameters
formData.append('use_high_resolution', 'true');
formData.append('layout_analysis', 'smart_layout_detection');
formData.append('ocr_strategy', 'auto_detection');
formData.append('ocr_engine', 'UnsiloedHawk');
formData.append('extract_strikethrough', 'false');
formData.append('merge_tables', 'true');
formData.append('enhance_reading_order', 'false');
formData.append('segment_filter', 'all');
formData.append('validate_segments', JSON.stringify(["Table", "Picture", "Formula"]));
formData.append('export_format', JSON.stringify(["docx"]));
formData.append('segment_analysis', JSON.stringify({
Table: {
html: 'VLM',
markdown: 'VLM',
extended_context: true,
crop_image: 'All',
model_id: 'us_table_v2'
}
}));
const response = await fetch('https://prod.visionapi.unsiloed.ai/parse', {
method: 'POST',
headers: {
'accept': 'application/json',
'api-key': 'your-api-key'
},
body: formData
});
if (response.ok) {
const result = await response.json();
console.log(`Job ID: ${result.job_id}`);
console.log(`Status: ${result.status}`);
console.log(`File Name: ${result.file_name}`);
console.log(`Credit Used: ${result.credit_used}`);
console.log(`Message: ${result.message}`);
} else {
console.error('Parsing failed:', response.status, await response.text());
}
// ========== Alternative: Use Presigned URL ==========
// Instead of file upload, use a presigned URL:
// formData.append('url', 'https://your-bucket.s3.amazonaws.com/document.pdf?signature=...');
// Remove the file upload lines above and use the url parameter instead
// ========== Use Case: Extract Tables Only ==========
// For financial documents where you only need tables:
// formData.append('merge_tables', 'true');
// formData.append('segment_filter', 'table');
// formData.append('layout_analysis', 'smart_layout_detection');
// formData.append('ocr_strategy', 'auto_detection');
// ========== Use Case: Citation Extraction (Academic Papers) ==========
// For research papers, enable citation extraction:
// formData.append('use_high_resolution', 'true');
// formData.append('layout_analysis', 'smart_layout_detection');
// formData.append('ocr_strategy', 'auto_detection');
// formData.append('xml_citation', 'true');
```
```json Success Response theme={null}
{
"job_id": "e77a5c42-4dc1-44d0-a30e-ed191e8a8908",
"status": "Starting",
"file_name": "document.pdf",
"created_at": "2025-07-18T10:42:10.545832520Z",
"message": "Task created successfully. Use GET /parse/{job_id} to check status and retrieve results.",
"credit_used": 5,
"quota_remaining": 23695,
"merge_tables": false
}
```
```text Error Response - Missing Input theme={null}
Either file or url must be provided
```
```json Error Response - File Too Large theme={null}
{
"error": "file_too_large",
"message": "File size exceeds the configured limit (default 500MB for uploads)",
"file_size_bytes": 614400000,
"max_file_size_bytes": 524288000
}
```
```text 400 - Bad Request theme={null}
Bad request: missing file/url, unsupported file type, or invalid parameters.
```
```json 402 - Insufficient Quota theme={null}
{
"message": "Insufficient quota",
"status": "INSUFFICIENT_QUOTA",
"quota_data": {
"remaining": 0
}
}
```
```text 429 - Usage Limit Exceeded theme={null}
Usage limit exceeded
```
```json 429 - Rate Limit Exceeded theme={null}
{
"error": "rate_limit_exceeded",
"message": "Rate limit of 100 requests per second exceeded. Retry after 1s.",
"retry_after": 1
}
```
```text 500 - Internal Server Error theme={null}
Internal server error.
```
```text 503 - Service Unavailable theme={null}
Service unavailable: job queue is at capacity. Retry after the duration indicated in the Retry-After header.
```
## Retrieving Results
After the job is created, use the GET /parse/ endpoint to check status and retrieve results:
```bash cURL theme={null}
curl -X 'GET' \
'https://prod.visionapi.unsiloed.ai/parse/{job_id}' \
-H 'accept: application/json' \
-H 'api-key: your-api-key'
```
```python Python theme={null}
import requests
import time
def get_parse_results(job_id, api_key):
"""Monitor job and retrieve results when complete"""
headers = {"api-key": api_key}
status_url = f"https://prod.visionapi.unsiloed.ai/parse/{job_id}"
# Poll for completion
while True:
response = requests.get(status_url, headers=headers)
if response.status_code == 200:
status_data = response.json()
print(f"Job Status: {status_data['status']}")
if status_data['status'] == 'Succeeded':
return status_data # Results are included in the same response
elif status_data['status'] == 'Failed':
raise Exception(f"Job failed: {status_data.get('message', 'Unknown error')}")
time.sleep(5) # Check every 5 seconds
# Usage
job_id = "e77a5c42-4dc1-44d0-a30e-ed191e8a8908"
results = get_parse_results(job_id, "your-api-key")
```
## Expected Results Structure
When the job completes successfully, the response contains comprehensive document analysis with enhanced processing:
```json theme={null}
{
"job_id": "04a7a6d8-5ef7-465a-b22a-8a98e7104dd9",
"status": "Succeeded",
"created_at": "2025-10-22T06:51:16.870302Z",
"started_at": "2025-10-22T06:51:16.966136Z",
"finished_at": "2025-10-22T06:57:19.821541Z",
"total_chunks": 25,
"chunks": [
{
"segments": [
{
"segment_type": "Title",
"content": "Disinvestment of IFCI's entire stake in Assets Care & Reconstruction Enterprise Ltd (ACRE)",
"image": null,
"page_number": 1,
"segment_id": "cc5f8dff-31be-4ccf-885d-4f9062fcee17",
"confidence": 0.90187776,
"page_width": 1191.0,
"page_height": 1684.0,
"html": "
Disinvestment of IFCI's entire stake in Assets Care & Reconstruction Enterprise Ltd (ACRE)
Background and context information about the disinvestment process...
",
"markdown": "Background and context information about the disinvestment process...",
"bbox": {
"left": 486.9685,
"top": 139.61847,
"width": 241.29932,
"height": 48.451706
},
"ocr": [
{
"bbox": {
"left": 50.9729,
"top": 3.4557495,
"width": 46.046875,
"height": 19.734375
},
"text": "Background",
"confidence": 0.99999654
}
]
}
]
}
]
}
```
## Segment Types
The parsing API identifies and processes different types of document segments with enhanced processing:
### Picture
Images and graphics within the document, including logos, charts, and illustrations. Enhanced with VLM-based description generation.
### SectionHeader
Document headers and titles that define section boundaries. Processed with semantic understanding.
### Text
Regular text content including paragraphs, sentences, and individual text elements. Enhanced with context-aware processing.
### Table
Tabular data with structured rows and columns. Enhanced with VLM-based formatting and extended context options. You can configure the table processing model using `model_id` in the `segment_analysis` parameter:
* **`us_table_v1`**: Standard table processing model
* **`us_table_v2`**: Enhanced table processing model with improved accuracy
### Caption
Text captions associated with images or figures. Processed with relationship awareness.
### Formula
Mathematical equations and expressions. Enhanced with specialized formula processing.
### Title
Document titles and main headings. Processed with enhanced formatting.
### Footnote
Document footnotes and references. Processed with context linking.
### ListItem
Bulleted and numbered list items. Processed with structure preservation.
### KeyValuePair
Form field regions — a printed label together with the value filled in against it. Produced when you pass `segmentation_version=v2`. See [Element Types](/docs/document-processing/parsing/element-types#keyvaluepair) for a full response sample.
Each segment includes detailed metadata such as confidence scores, bounding boxes, OCR data, and formatted output in both HTML and Markdown with VLM enhancement.
## Error Handling
### Common Error Scenarios
1. **Invalid API Key**: Authentication failed
2. **File Too Large**: File exceeds size limits
3. **Invalid Configuration**: Malformed processing parameters
4. **Server Error**: Internal processing error
5. **Processing Timeout**: Task took too long to complete
6. **Missing File or URL**: Neither `file` nor `url` parameter provided
7. **Both File and URL Provided**: Cannot provide both `file` and `url` simultaneously
8. **Invalid URL**: URL is not accessible or malformed
9. **URL Download Failed**: Unable to download document from provided URL
10. **Insufficient Quota** (`402`): Not enough page credits remaining.
11. **Usage Limit Exceeded** (`429`): Billing usage cap reached. Returns plain text: `Usage limit exceeded`. No `Retry-After` header.
12. **Rate Limit Exceeded** (`429`): Org exceeded its per-second request budget (100 requests per second per organization by default; higher limits are available on request). Returns JSON `{"error": "rate_limit_exceeded", "message": ..., "retry_after": 1}` with a `Retry-After: 1` header.
13. **Internal Server Error** (`500`): An unexpected error occurred during processing.
14. **Service Unavailable** (`503`): Job queue is at capacity. Retry after the duration indicated in the `Retry-After` header.
15. **Forbidden** (`403`): Access has been revoked.
# Upload Large Files for Parsing
Source: https://docs.unsiloed.ai/docs/api-reference/parser/parse-document-v2
api-reference/parser/openapi-v2.json POST /v2/parse/upload
Upload a large document to Unsiloed-managed storage with a presigned URL, then poll the standard Parse status endpoint.
## Overview
The v2 presigned upload endpoint decouples document delivery from job creation. Instead of uploading your file through the API server, you:
1. **POST** to `/v2/parse/upload` with your configuration: the API creates a parse job and returns a short-lived presigned URL.
2. **PUT** your file directly to the presigned URL: the API server is never in the transfer path.
3. Once the upload completes, the job is automatically enqueued for processing.
4. **Poll** `GET /parse/{job_id}` to track progress and retrieve results (the same endpoint as the standard Parse flow).
Use this flow when a file is too large to upload directly to the standard [Parse Document](/docs/api-reference/parser/parse-document) endpoint. For documents that fit in a direct upload, use `POST /parse`, or `POST /parse/excel` for XLS and XLSX workbooks. Excel workbooks uploaded through this flow are processed by the Excel pipeline automatically.
The job status starts as `AwaitingUpload`. It transitions to `Queued` once the upload completes, then to `Processing`, and finally `Succeeded` or `Failed`. See [Get Parse Job Status](/docs/api-reference/parser/get-parse-job-status) for the full response schema and polling examples.
## Request
File name with extension (e.g. `"report.pdf"`). Determines the MIME type for the presigned URL and selects the Excel pipeline for XLS and XLSX names. Supported formats: PDF, images (PNG, JPG/JPEG, TIFF/TIF, WebP), office documents (DOC, DOCX, PPT, PPTX, XLS, XLSX), and HTML (HTML/HTM).
Enhanced image quality: enables high-resolution processing with upscaling algorithms. Improves OCR accuracy on low-quality scans by enhancing clarity and contrast. Latency penalty: \~2–3 seconds per page. Defaults to `false`.
Layout analysis strategy: `smart_layout_detection` (default), `page_by_page`, or `advanced_layout_detection`.
* `"smart_layout_detection"` **(default)**: Identifies document structure, headers, sections, and content relationships across the entire document using bounding boxes.
* `"page_by_page"`: Analyzes each page independently as a single segment. Faster for simple documents.
* `"advanced_layout_detection"`: Higher-accuracy layout detection for complex pages — multi-column layouts, dense tables/figures. Slower than `smart_layout_detection`.
Which layout-detection model to run: `v1` (default) or `v2`. `v2` is form-aware and additionally returns `KeyValuePair`, `Signature`, and `Seal` segments. See the [standard Parse endpoint](/docs/api-reference/parser/parse-document#segmentation-version) for details.
OCR strategy: `auto_detection` (default) or `force_ocr`.
* `"auto_detection"` **(default)**: Intelligently detects bad quality PDFs, scanned documents, and images, then applies OCR only where needed.
* `"force_ocr"`: Runs OCR on the entire document regardless of quality.
OCR engine to use for text recognition: `UnsiloedHawk` (default, recommended), `UnsiloedBeta`, or `UnsiloedStorm`.
* `"UnsiloedHawk"` **(Recommended, default)**: Higher accuracy for complex layouts and mixed content.
* `"UnsiloedBeta"`: Handles rotated/warped text and irregular bounding boxes.
* `"UnsiloedStorm"`: Enterprise-grade accuracy optimized for 50+ languages.
Per-segment OCR enhancement: re-runs a dedicated agentic OCR model on each detected segment after layout detection for higher accuracy. Omit or leave empty to disable.
* `"standard"`: Good balance of speed and accuracy.
* `"advanced"`: Higher quality, best for complex layouts, rotated text, and mixed-language content.
Strikethrough detection: detects and preserves strikethrough formatting in HTML and Markdown output. Defaults to `false`.
Cross-page table consolidation: detects and combines table segments across page breaks. Reconstructs complete table structure by matching headers and columns. Defaults to `false`.
Maximum number of tables per merge group. Groups larger than this are split. Defaults to `20`; values below 2 are clamped to 2.
Reorder detected segments to follow the document's natural reading order. Recommended for **multi-column layouts** (newspapers, two- or three-column papers) where the default smart layout occasionally interleaves text across columns. Has no effect when `layout_analysis="page_by_page"`. Defaults to `false`.
Run a PII pre-check before parsing. If PII is found at the configured severity, the task is rejected without parsing. Defaults to `false`.
Severity threshold to block on when `detect_pii` is enabled: `any` (default), `low`, `medium`, or `high`. `any` blocks on any PII found; `low` adds quasi-identifiers (names, dates, locations); `medium` blocks on contact PII (email, phone) or higher; `high` blocks only on direct identifiers (SSN, passport, credit card). Ignored if `detect_pii` is false.
PII detector engine: `standard` (default) or `advanced` (higher precision, additional processing cost). Any other value falls back to `standard`. Ignored if `detect_pii` is false.
Attach hyperlink URLs from PDF annotations to OCR results. Defaults to `false`.
Transfer text color from the PDF text layer to OCR results. Defaults to `false`.
Citation extraction: extracts academic citations from PDF documents and hyperlinks them in the markdown output. Generates structured bibliography metadata. PDFs only. Defaults to `false`.
Specify which pages to process. Leave empty to process all pages.
Examples: `"1-5"` (pages 1 to 5), `"2,4,6"` (specific pages), `"[1,3,5]"` (array format), `"1-3,7,10-15"` (combination).
Seconds the returned presigned upload URL stays valid. Defaults to 15 minutes (900 seconds). Once the client uploads the file, the task's expiry is cleared and the task lifetime is no longer bounded by this value.
Export format(s) to generate after processing. Supported: `["docx", "markdown", "json"]`. The exported files are available via the `exports` field in the task response.
Segment type naming convention. `"Unsiloed"` (default) uses names like `PageHeader`, `ListItem`, `Picture`. `"Other"` uses alternative names like `Header`, `List Item`, `Figure`.
Comma-separated segment types to keep in the output, or `"all"` to include every type. Defaults to `"all"`.
Segment validation: uses a vision model to validate and correct segment types, fixing misclassified segments. Example: `["Table", "Formula", "Picture"]`. Defaults to `["Table", "Picture"]`; an empty or unparseable value also falls back to that default, so Table and Picture validation runs even when this field is omitted.
Legacy option to validate table segments using a vision model. Prefer `validate_segments` instead.
JSON object controlling how segments are grouped into chunks. See the [standard Parse endpoint](/docs/api-reference/parser/parse-document) for full schema details.
Configure how different segment types are processed: table models, image descriptions, and formula processing.
On this endpoint the field is named `segment_processing`. The standard endpoint field `segment_analysis` is not accepted here and is silently ignored.
```json theme={null}
{
"Table": {"html": "VLM", "markdown": "VLM", "model_id": "us_table_v2"},
"Picture": {"html": "VLM", "markdown": "VLM", "model_id": "nova"},
"Formula": {"html": "Auto", "markdown": "VLM", "model_id": "nova"}
}
```
Options per segment type:
* `html` / `markdown`: `"VLM"` (use a VLM to generate the output) or `"Auto"` (skip the VLM call and build from the OCR/recognized text)
* `model_id` (Table): `"astra"`, `"us_table_v1"`, `"us_table_v2"`
* `model_id` (Picture/Formula): `"nova"`, `"luna"`, `"sol"`
* `model_id` (KeyValuePair): `"astra_v2"`
* `use_table_ocr` (Table only): Advanced OCR for bordered cells and complex table layouts.
* `vlm`: Custom prompt for the VLM model. Use this to give the model specific instructions for extracting or describing these segment types.
* `translation`: Optional per-segment translation, e.g. `{"provider": "Auto", "target_language": "en"}`. `provider` is `"Auto"` for fast machine translation or `"VLM"`/`"LLM"` for model-based translation; `target_language` is an ISO 639-1 code, or `"auto"` to auto-detect the source and translate to English.
Defaults are **not uniform** across segment types. `Table.html`, `Picture.markdown`, `Formula.markdown`, `KeyValuePair.markdown`, and both `Page` strategies default to `"VLM"`; everything else defaults to `"Auto"`. If you want a Picture segment's Markdown to contain only the transcribed text (rather than a VLM-written description), pass `{"Picture": {"markdown": "Auto"}}`.
`KeyValuePair` is configurable here too (form-field regions, produced when `segmentation_version=v2`). It returns an `image` crop URL alongside the text.
For full details, see [Parse Document](/docs/api-reference/parser/parse-document).
JSON object for LLM processing configuration.
JSON object controlling which fields are included in the output. Set fields to `false` to exclude them and reduce response size. Available fields: `html`, `markdown`, `ocr`, `image`, `content`, `bbox`, `confidence`, `embed`, `chart_data`. All default to `true`. Ignored when `response_profile` is `slim` or `full` (the profile wins).
Response shape selector: `slim`, `full`, or `custom`. Omit to return the full shape.
* `"slim"`: Returns only the essentials per chunk — `embed`, `bbox`, `page_number`, `segment_id`, `segment_type`, and HTML for tables / Markdown for everything else. Drops `content`, `image`, `ocr`, `confidence`, `chart_data`, `page_height`, `page_width`. Best for embedding-only workflows where you want the smallest payload.
* `"full"`: Every field returned (equivalent to omitting this param).
* `"custom"`: Honor `output_fields` verbatim.
When both `response_profile` and `output_fields` are provided, the profile wins — `output_fields` is only consulted for `custom` or when the profile is omitted.
How to handle per-page errors during processing. `"Continue"` (default) skips failed pages and continues processing the rest. `"Fail"` aborts the entire job on the first error.
## Response
Unique identifier for the parse job. Use this with `GET /parse/{job_id}` to poll status and retrieve results.
Presigned PUT URL. Upload your file directly to this URL; no `api-key` header is needed for the upload itself. Valid until `expires_at`.
RFC 3339 timestamp when `upload_url` expires. Upload must complete before this time. The presigned URL is valid for 15 minutes by default; when `expires_in` is set (and positive), the upload URL lifetime is set to that same value, which can shorten or extend the window.
Always `"PUT"`. Use an HTTP PUT request when uploading to `upload_url`.
Key-value headers you MUST include in your PUT request. Always includes `Content-Type` set to the MIME type inferred from `file_name`.
Number of page credits deducted as an initial reservation, reconciled on upload.
Remaining page credits after deduction.
## Step-by-Step Guide
### Step 1: Create the parse job
POST to `/v2/parse/upload` with your file name and configuration. The API returns a `job_id` and a short-lived presigned URL.
```python Python theme={null}
import requests
API_KEY = "your-api-key"
BASE_URL = "https://prod.visionapi.unsiloed.ai"
response = requests.post(
f"{BASE_URL}/v2/parse/upload",
headers={"api-key": API_KEY, "Content-Type": "application/json"},
json={
"file_name": "report.pdf",
"use_high_resolution": True,
"layout_analysis": "smart_layout_detection",
"ocr_strategy": "auto_detection",
"extract_strikethrough": False,
"merge_tables": False,
},
)
response.raise_for_status()
data = response.json()
job_id = data["job_id"]
upload_url = data["upload_url"]
upload_headers = data["upload_headers"]
print(f"Job ID: {job_id}")
print(f"Upload URL expires: {data['expires_at']}")
```
```javascript JavaScript theme={null}
const API_KEY = "your-api-key";
const BASE_URL = "https://prod.visionapi.unsiloed.ai";
const response = await fetch(`${BASE_URL}/v2/parse/upload`, {
method: "POST",
headers: { "api-key": API_KEY, "Content-Type": "application/json" },
body: JSON.stringify({
file_name: "report.pdf",
use_high_resolution: true,
layout_analysis: "smart_layout_detection",
ocr_strategy: "auto_detection",
extract_strikethrough: false,
merge_tables: false,
}),
});
const { job_id, upload_url, upload_headers, expires_at } = await response.json();
console.log(`Job ID: ${job_id}`);
console.log(`Upload URL expires: ${expires_at}`);
```
```bash cURL theme={null}
curl -X POST https://prod.visionapi.unsiloed.ai/v2/parse/upload \
-H "api-key: your-api-key" \
-H "Content-Type: application/json" \
-d '{
"file_name": "report.pdf",
"use_high_resolution": true,
"layout_analysis": "smart_layout_detection",
"ocr_strategy": "auto_detection",
"extract_strikethrough": false,
"merge_tables": false
}'
```
The response contains the presigned URL and required headers:
```json theme={null}
{
"job_id": "a3f1c2d4-7e8b-4a9f-b2c1-123456789abc",
"upload_url": "https://upload.visionapi.unsiloed.ai/uploads/a3f1c2d4.../report.pdf?signature=...",
"expires_at": "2025-10-22T07:06:16Z",
"upload_method": "PUT",
"upload_headers": {
"Content-Type": "application/pdf"
},
"credit_used": 1,
"quota_remaining": 1000
}
```
### Step 2: Upload the file
Use the `upload_url` and `upload_headers` from the response to PUT your file. No API key is needed for this request.
```python Python theme={null}
with open("report.pdf", "rb") as f:
put_response = requests.put(upload_url, headers=upload_headers, data=f)
put_response.raise_for_status()
print("Upload complete, job is now queued for processing")
```
```javascript JavaScript theme={null}
import fs from "fs";
const fileBuffer = fs.readFileSync("report.pdf");
const putResponse = await fetch(upload_url, {
method: "PUT",
headers: upload_headers,
body: fileBuffer,
});
if (!putResponse.ok) {
throw new Error(`Upload failed: ${putResponse.status}`);
}
console.log("Upload complete, job is now queued for processing");
```
```bash cURL theme={null}
curl -X PUT "$UPLOAD_URL" \
-H "Content-Type: application/pdf" \
--data-binary @report.pdf
```
You must include every header listed in `upload_headers`. A missing or mismatched `Content-Type` will cause the upload to be rejected with a `403`.
### Step 3: Poll for results
Once the upload completes, the job transitions from `AwaitingUpload` → `Queued` → `Processing` → `Succeeded`. Poll the job through the standard `GET /parse/{job_id}` endpoint.
```python Python theme={null}
import time
while True:
status_response = requests.get(
f"{BASE_URL}/parse/{job_id}",
headers={"api-key": API_KEY},
)
status_response.raise_for_status()
job = status_response.json()
print(f"Status: {job['status']}")
if job["status"] == "Succeeded":
print(f"Done! {job['total_chunks']} chunks extracted.")
break
elif job["status"] == "Failed":
raise RuntimeError(f"Job failed: {job.get('message')}")
time.sleep(5)
```
```javascript JavaScript theme={null}
while (true) {
const statusResponse = await fetch(`${BASE_URL}/parse/${job_id}`, {
headers: { "api-key": API_KEY },
});
const job = await statusResponse.json();
console.log(`Status: ${job.status}`);
if (job.status === "Succeeded") {
console.log(`Done! ${job.total_chunks} chunks extracted.`);
break;
} else if (job.status === "Failed") {
throw new Error(`Job failed: ${job.message}`);
}
await new Promise((r) => setTimeout(r, 5000));
}
```
```bash cURL theme={null}
curl -X GET "https://prod.visionapi.unsiloed.ai/parse/{job_id}" \
-H "api-key: your-api-key"
```
See [Get Parse Job Status](/docs/api-reference/parser/get-parse-job-status) for the full response schema.
```bash cURL theme={null}
# Step 1: Create job and get presigned URL
RESPONSE=$(curl -s -X POST https://prod.visionapi.unsiloed.ai/v2/parse/upload \
-H "api-key: your-api-key" \
-H "Content-Type: application/json" \
-d '{
"file_name": "report.pdf",
"use_high_resolution": true,
"layout_analysis": "smart_layout_detection",
"ocr_strategy": "auto_detection",
"extract_strikethrough": false,
"merge_tables": false
}')
JOB_ID=$(echo "$RESPONSE" | jq -r '.job_id')
UPLOAD_URL=$(echo "$RESPONSE" | jq -r '.upload_url')
# Step 2: Upload file
curl -X PUT "$UPLOAD_URL" \
-H "Content-Type: application/pdf" \
--data-binary @report.pdf
# Step 3: Poll for results
curl -X GET "https://prod.visionapi.unsiloed.ai/parse/$JOB_ID" \
-H "api-key: your-api-key"
```
```python Python theme={null}
import requests
import time
API_KEY = "your-api-key"
BASE_URL = "https://prod.visionapi.unsiloed.ai"
# Step 1: Create job and get presigned URL
resp = requests.post(
f"{BASE_URL}/v2/parse/upload",
headers={"api-key": API_KEY, "Content-Type": "application/json"},
json={
"file_name": "report.pdf",
"use_high_resolution": True,
"layout_analysis": "smart_layout_detection",
"ocr_strategy": "auto_detection",
"extract_strikethrough": False,
"merge_tables": False,
},
)
resp.raise_for_status()
data = resp.json()
job_id = data["job_id"]
upload_url = data["upload_url"]
upload_headers = data["upload_headers"]
print(f"Job created: {job_id}")
print(f"Upload URL expires: {data['expires_at']}")
# Step 2: Upload file
with open("report.pdf", "rb") as f:
put_resp = requests.put(upload_url, headers=upload_headers, data=f)
put_resp.raise_for_status()
print("Upload complete, job is now queued for processing")
# Step 3: Poll for results
while True:
status_resp = requests.get(
f"{BASE_URL}/parse/{job_id}",
headers={"api-key": API_KEY},
)
status_resp.raise_for_status()
job = status_resp.json()
print(f"Status: {job['status']}")
if job["status"] == "Succeeded":
print(f"Done! {job['total_chunks']} chunks extracted.")
break
elif job["status"] == "Failed":
raise RuntimeError(f"Job failed: {job.get('message')}")
time.sleep(5)
```
```javascript JavaScript theme={null}
import fs from "fs";
const API_KEY = "your-api-key";
const BASE_URL = "https://prod.visionapi.unsiloed.ai";
// Step 1: Create job and get presigned URL
const createResp = await fetch(`${BASE_URL}/v2/parse/upload`, {
method: "POST",
headers: { "api-key": API_KEY, "Content-Type": "application/json" },
body: JSON.stringify({
file_name: "report.pdf",
use_high_resolution: true,
layout_analysis: "smart_layout_detection",
ocr_strategy: "auto_detection",
extract_strikethrough: false,
merge_tables: false,
}),
});
const { job_id, upload_url, upload_headers, expires_at } = await createResp.json();
console.log(`Job created: ${job_id}, URL expires: ${expires_at}`);
// Step 2: Upload file
const fileBuffer = fs.readFileSync("report.pdf");
await fetch(upload_url, {
method: "PUT",
headers: upload_headers,
body: fileBuffer,
});
console.log("Upload complete, job is now queued");
// Step 3: Poll for results
while (true) {
const statusResp = await fetch(`${BASE_URL}/parse/${job_id}`, {
headers: { "api-key": API_KEY },
});
const job = await statusResp.json();
console.log(`Status: ${job.status}`);
if (job.status === "Succeeded") {
console.log(`Done! ${job.total_chunks} chunks extracted.`);
break;
} else if (job.status === "Failed") {
throw new Error(`Job failed: ${job.message}`);
}
await new Promise((r) => setTimeout(r, 5000));
}
```
```json Success Response theme={null}
{
"job_id": "a3f1c2d4-7e8b-4a9f-b2c1-123456789abc",
"upload_url": "https://upload.visionapi.unsiloed.ai/uploads/a3f1c2d4.../report.pdf?signature=...",
"expires_at": "2025-10-22T07:06:16Z",
"upload_method": "PUT",
"upload_headers": {
"Content-Type": "application/pdf"
},
"credit_used": 1,
"quota_remaining": 1000
}
```
```json Error Response - Unsupported File Type theme={null}
{
"error": "Unsupported file type: .exe"
}
```
```json 429 - Rate Limit Exceeded theme={null}
{
"error": "rate_limit_exceeded",
"message": "Rate limit of 100 requests per second exceeded. Retry after 1s.",
"retry_after": 1
}
```
## Standard and Large-File Upload Flows
| Step | Standard (`POST /parse`) | Large file (`POST /v2/parse/upload`) |
| - | - | - |
| Submit configuration | Send it with the multipart file | Send it before uploading the file |
| Upload document | Include it in the request | PUT it to the returned presigned URL |
| Job starts | After the request is accepted | After the presigned upload completes |
| Use when | The file fits in a direct API upload | The file is too large for a direct API upload |
## Error Handling
| Status | Cause | Action |
| - | - | - |
| `400` | Invalid or missing `file_name`, unsupported extension | Fix `file_name` and retry |
| `401` | Missing or invalid `api-key` | Check your API key |
| `402` | Insufficient quota or expired credits | Add page credits to your account or renew your plan |
| `403` | Access has been revoked | Contact support |
| `429` | Rate limit exceeded | Back off and retry after 1 second |
| `500` | Internal server error | Retry with exponential backoff |
| `503` | Job queue is at capacity | Back off and retry after the `Retry-After` header value |
| `403` on upload | Missing or wrong headers, expired URL | Check `upload_headers`, get a new URL |
| Job status `Failed` | Processing error | Check `message` field in the status response |
A job can move from `Queued` straight to `Failed` without ever reaching `Processing`: after the upload, the document is page-counted and rejected if it exceeds the page limit (default 2000 pages; split large documents first) or if the remaining page quota cannot cover it.
# Parse Excel
Source: https://docs.unsiloed.ai/docs/api-reference/parser/parse-excel
api-reference/parser/openapi-v1.json POST /parse/excel
Dedicated ingestion endpoint for Excel workbooks (.xls / .xlsx) with spreadsheet-specific configuration.
## Overview
The Parse Excel endpoint processes Excel workbooks (`.xls`, `.xlsx`) using a dedicated spreadsheet pipeline. It shares auth, billing, quota, and rate-limit infrastructure with [Parse Document](/docs/api-reference/parser/parse-document), and is polled via the same `GET /parse/{job_id}`.
1. **POST** to `/parse/excel` with your file (or `url`) and any spreadsheet-specific configuration.
2. The job is automatically enqueued for processing.
3. **Poll** `GET /parse/{job_id}` to track progress and retrieve results.
PDF parsing options (`ocr_strategy`, `layout_analysis`, `segment_processing`, etc.) do not apply here — they are silently ignored. Non-Excel uploads are rejected with `400`; submit those to [`POST /parse`](/docs/api-reference/parser/parse-document) instead.
## Request
Provide either `file` (multipart binary upload) or `url` (presigned/public URL). The `file` field is multipart-only; JSON callers must use `url`.
Excel workbook to process. Supported formats: `.xls`, `.xlsx`. Required if `url` is not provided.
Presigned or public URL of the workbook to fetch and process. Required if `file` is not provided.
### Hidden content
Drop hidden content from the output. Defaults to `false`. The five sub-toggles below are only honored when this is `true`.
When excluding hidden content, also drop entire hidden sheets. Defaults to `true`.
When excluding hidden content, also drop hidden rows from visible sheets. Defaults to `true`.
When excluding hidden content, also drop hidden columns from visible sheets. Defaults to `true`.
When excluding hidden content, also drop styling. Defaults to `true`.
When excluding hidden content, also drop embedded and pasted images. Defaults to `false`.
### Table extraction
Split large tables into smaller segments to keep individual response items manageable. Defaults to `true`.
Maximum number of rows in each split segment. Only effective when `split_large_tables` is `true`. Defaults to `50`.
How aggressively to detect distinct logical tables on the same sheet.
* `"accurate"` **(default)**: Best fidelity at the cost of latency.
* `"fast"`: Quicker clustering, may merge nearby tables.
* `"off"`: Treat each sheet as a single table.
### Lifecycle
Reserved field. Persisted on the task configuration but currently has no effect on retention — Excel tasks are not auto-deleted. To get a presigned-upload TTL for PDFs and other documents, use [`POST /v2/parse/upload`](/docs/api-reference/parser/parse-document-v2) instead.
## Response
The endpoint returns HTTP 200 with the same envelope as [`POST /parse`](/docs/api-reference/parser/parse-document):
Job identifier. Pass this to `GET /parse/{job_id}` to poll for results.
Initial job status. Always `"Starting"` on creation.
Name of the uploaded workbook. For URL submissions this is the last path segment of the URL, or `"unknown"` when no usable segment exists.
ISO 8601 timestamp when the job was created.
Human-readable status message with a polling hint.
Number of credits deducted from your quota for this job.
Remaining quota after this job was deducted.
Reflects the table-merging flag stored on the job. Always `false` for Excel jobs — table merging is a PDF-only feature.
```bash cURL theme={null}
curl -X POST 'https://prod.visionapi.unsiloed.ai/parse/excel' \
-H 'accept: application/json' \
-H 'api-key: your-api-key' \
-H 'Content-Type: multipart/form-data' \
-F 'file=@workbook.xlsx;type=application/vnd.openxmlformats-officedocument.spreadsheetml.sheet' \
-F 'exclude_hidden=true' \
-F 'split_large_tables=true' \
-F 'max_rows_per_segment=100' \
-F 'table_clustering=accurate'
# Alternative: presigned / public URL instead of file upload
# -F 'url=https://your-bucket.s3.amazonaws.com/workbook.xlsx?signature=...'
```
```python Python theme={null}
import requests
url = "https://prod.visionapi.unsiloed.ai/parse/excel"
headers = {"accept": "application/json", "api-key": "your-api-key"}
files = {"file": ("workbook.xlsx", open("workbook.xlsx", "rb"),
"application/vnd.openxmlformats-officedocument.spreadsheetml.sheet")}
data = {
"exclude_hidden": "true",
"split_large_tables": "true",
"max_rows_per_segment": "100",
"table_clustering": "accurate",
}
response = requests.post(url, headers=headers, files=files, data=data)
if response.status_code == 200:
result = response.json()
print(f"Job ID: {result['job_id']}")
print(f"Status: {result['status']}")
print(f"Credit used: {result['credit_used']}")
else:
print("Error:", response.status_code, response.text)
files["file"][1].close()
# ========== Alternative: Use Presigned URL ==========
# data = {
# "url": "https://your-bucket.s3.amazonaws.com/workbook.xlsx?signature=...",
# "exclude_hidden": "true",
# }
# response = requests.post(url, headers=headers, data=data)
```
```javascript JavaScript theme={null}
const formData = new FormData();
const fileInput = document.querySelector('input[type="file"]');
if (fileInput.files[0]) {
formData.append('file', fileInput.files[0]);
}
formData.append('exclude_hidden', 'true');
formData.append('split_large_tables', 'true');
formData.append('max_rows_per_segment', '100');
formData.append('table_clustering', 'accurate');
const response = await fetch('https://prod.visionapi.unsiloed.ai/parse/excel', {
method: 'POST',
headers: {
'accept': 'application/json',
'api-key': 'your-api-key',
},
body: formData,
});
if (response.ok) {
const result = await response.json();
console.log(`Job ID: ${result.job_id}`);
console.log(`Status: ${result.status}`);
} else {
console.error('Parsing failed:', response.status, await response.text());
}
// ========== Alternative: Use Presigned URL ==========
// formData.append('url', 'https://your-bucket.s3.amazonaws.com/workbook.xlsx?signature=...');
```
```json Success Response theme={null}
{
"job_id": "9b1f7a04-2c33-4f8e-9c92-6f8a2e84b3d1",
"status": "Starting",
"file_name": "workbook.xlsx",
"created_at": "2026-06-17T14:22:08.901234Z",
"message": "Task created successfully. Use GET /parse/{job_id} to check status and retrieve results.",
"credit_used": 3,
"quota_remaining": 23692,
"merge_tables": false
}
```
```text 400 - Non-Excel File theme={null}
Bad request: file is not a supported Excel format (.xls / .xlsx).
```
```text 400 - Missing Input theme={null}
Either file or url must be provided
```
```json 402 - Insufficient Quota theme={null}
{
"message": "Insufficient quota",
"status": "INSUFFICIENT_QUOTA",
"quota_data": {
"remaining": 0
}
}
```
```json 429 - Rate Limit Exceeded theme={null}
{
"error": "rate_limit_exceeded",
"message": "Rate limit of 100 requests per second exceeded. Retry after 1s.",
"retry_after": 1
}
```
```text 503 - Service Unavailable theme={null}
Service unavailable: job queue is at capacity. Retry after the duration indicated in the Retry-After header.
```
## Retrieving Results
Use `GET /parse/{job_id}` (the shared polling endpoint) to check status and retrieve results. The result envelope is the same as for PDF jobs — chunks containing segments — and Excel segments include a `cell_references` field linking each segment back to its source sheet, address, and range.
```bash cURL theme={null}
curl -X GET "https://prod.visionapi.unsiloed.ai/parse/{job_id}" \
-H "accept: application/json" \
-H "api-key: your-api-key"
```
```python Python theme={null}
import time, requests
def get_excel_results(job_id, api_key):
headers = {"api-key": api_key}
status_url = f"https://prod.visionapi.unsiloed.ai/parse/{job_id}"
while True:
response = requests.get(status_url, headers=headers)
response.raise_for_status()
job = response.json()
print(f"Status: {job['status']}")
if job["status"] == "Succeeded":
return job
if job["status"] == "Failed":
raise RuntimeError(f"Job failed: {job.get('message')}")
time.sleep(5)
```
See [Get Parse Job Status](/docs/api-reference/parser/get-parse-job-status) for the full response schema and query parameters.
## Error Handling
| Status | Cause | Action |
| - | - | - |
| `400` | Missing `file`/`url`, non-Excel file type, or malformed parameters | Check the file extension and required fields |
| `401` | Missing or invalid `api-key` | Check your API key |
| `402` | Insufficient quota | Add credits to your account or renew your plan |
| `403` | Access has been revoked | Contact support |
| `429` | Rate limit (100 req/s per organization) or billing usage cap hit | Back off and retry after the `Retry-After` header value |
| `500` | Internal server error | Retry with exponential backoff |
| `503` | Job queue at capacity | Retry after the duration indicated in the `Retry-After` header |
# Get Split Result
Source: https://docs.unsiloed.ai/docs/api-reference/splitting/get-split-status
api-reference/openapi.json GET /splitter/{job_id}
Check the status and retrieve results of document splitting jobs
## Overview
The Get Split Status endpoint retrieves the status and results of a document splitting job. Use this endpoint to check if splitting is complete and download the resulting split documents.
Splitting jobs are processed asynchronously. Poll this endpoint to check completion status.
## Path Parameters
The unique identifier of the splitting job
## Response
Unique identifier for the splitting job
Current job status: "queued" while the job waits to be picked up, "processing" while it runs, then "completed" or "failed"
Progress message describing current processing stage
Presigned download URL for the original uploaded file. Expires roughly an hour after the response is generated; re-issue this request to get a fresh URL.
Name of the original file
Echo of the parameters used for this job
The category names submitted with the job
The descriptions submitted for each category, when provided
Whether pages were reordered within each category after classification
Number of pages in the uploaded document
Split results. Always present as a key; `null` until status is "completed"
Whether the splitting operation succeeded
Success/failure message (e.g., "Successfully split PDF into 3 files")
Array of split PDF files
Name of the split file (e.g., "Invoice.pdf")
Relative path to the file
File type (always "file")
Unique file identifier
Presigned download URL for the split file. Expires roughly an hour after the response is generated, so download the file straight away rather than storing the URL. Re-issue this request to get a fresh URL.
Classification confidence score (0.0-1.0)
Error message (if job failed, otherwise null)
```bash cURL theme={null}
curl -X GET "https://prod.visionapi.unsiloed.ai/splitter/04a7a6d8-5ef7-465a-b22a-8a98e7104dd9" \
-H "accept: application/json" \
-H "api-key: your-api-key"
```
```python Python theme={null}
import requests
def get_split_status(job_id, api_key):
"""Get the status of a splitting job"""
url = f"https://prod.visionapi.unsiloed.ai/splitter/{job_id}"
headers = {"api-key": api_key}
response = requests.get(url, headers=headers)
if response.status_code == 200:
job = response.json()
print(f"Job Status: {job['status']}")
if job['status'] == 'completed':
print("Split completed successfully!")
if job.get('result', {}).get('success'):
print(f"Message: {job['result']['message']}")
for file_info in job['result']['files']:
print(f" File: {file_info['name']}")
print(f" Download: {file_info['full_path']}")
return job
elif job['status'] == 'failed':
print(f"Job failed: {job.get('error', 'Unknown error')}")
return None
elif job['status'] == 'processing':
print(f"Job is processing: {job.get('progress', '')}")
return job
else:
print("Error:", response.status_code, response.text)
return None
# Usage
job_id = "04a7a6d8-5ef7-465a-b22a-8a98e7104dd9"
result = get_split_status(job_id, "your-api-key")
```
```javascript JavaScript theme={null}
async function getSplitStatus(jobId, apiKey) {
const response = await fetch(`https://prod.visionapi.unsiloed.ai/splitter/${jobId}`, {
headers: {
'accept': 'application/json',
'api-key': apiKey
}
});
if (response.ok) {
const job = await response.json();
console.log('Job Status:', job.status);
switch (job.status) {
case 'completed':
console.log('Split completed successfully!');
if (job.result?.success) {
console.log('Message:', job.result.message);
job.result.files.forEach(file => {
console.log(` File: ${file.name}`);
console.log(` Download: ${file.full_path}`);
});
}
return job;
case 'failed':
console.log('Job failed:', job.error || 'Unknown error');
return null;
case 'processing':
console.log('Job is processing:', job.progress || '');
return job;
default:
console.log('Unknown status:', job.status);
return job;
}
} else {
console.error('Failed to get job status:', await response.text());
return null;
}
}
// Usage
const jobId = '04a7a6d8-5ef7-465a-b22a-8a98e7104dd9';
const result = await getSplitStatus(jobId, 'your-api-key');
```
## Polling for Job Completion
Here are complete examples that poll the splitting job until it completes:
```bash cURL (with polling) theme={null}
#!/bin/bash
JOB_ID="04a7a6d8-5ef7-465a-b22a-8a98e7104dd9"
API_KEY="your-api-key"
MAX_ATTEMPTS=60
SLEEP_INTERVAL=5
echo "Polling split job status..."
for ((i=1; i<=MAX_ATTEMPTS; i++)); do
echo "Attempt $i..."
RESPONSE=$(curl -s -X GET "https://prod.visionapi.unsiloed.ai/splitter/$JOB_ID" \
-H "accept: application/json" \
-H "api-key: $API_KEY")
STATUS=$(echo $RESPONSE | grep -o '"status":"[^"]*"' | cut -d'"' -f4)
echo "Status: $STATUS"
if [ "$STATUS" = "completed" ]; then
echo "Split completed successfully!"
echo $RESPONSE | python -m json.tool
break
elif [ "$STATUS" = "failed" ]; then
echo "Split failed!"
echo $RESPONSE | python -m json.tool
break
else
echo "Still processing... waiting $SLEEP_INTERVAL seconds"
sleep $SLEEP_INTERVAL
fi
done
```
```python Python (with polling) theme={null}
import requests
import time
def poll_split_job(job_id, api_key, max_attempts=60, sleep_interval=5):
"""
Poll a splitting job until completion or failure
Args:
job_id: The job ID returned from split request
api_key: Your API key
max_attempts: Maximum number of polling attempts
sleep_interval: Seconds to wait between polls
Returns:
Final job result or None if failed
"""
url = f"https://prod.visionapi.unsiloed.ai/splitter/{job_id}"
headers = {"api-key": api_key}
print(f"Polling split job {job_id}...")
for attempt in range(1, max_attempts + 1):
print(f"Attempt {attempt}...")
try:
response = requests.get(url, headers=headers)
if response.status_code == 200:
job = response.json()
status = job.get('status')
progress = job.get('progress', '')
print(f"Status: {status}")
if progress:
print(f"Progress: {progress}")
if status == 'completed':
print("\n✓ Split completed successfully!")
if job.get('result', {}).get('success'):
print(f"Message: {job['result']['message']}")
print(f"\nSplit files:")
for file_info in job['result']['files']:
print(f" • {file_info['name']}")
print(f" Download: {file_info['full_path']}")
print(f" Confidence: {file_info['confidence_score']:.2%}")
return job
elif status == 'failed':
print(f"\n✗ Split failed: {job.get('error', 'Unknown error')}")
return None
elif status == 'processing':
print(f"Still processing... waiting {sleep_interval} seconds\n")
time.sleep(sleep_interval)
elif response.status_code == 404:
print(f"Job not found. It may have expired.")
return None
else:
print(f"Error: {response.status_code} - {response.text}")
return None
except Exception as e:
print(f"Error polling job: {e}")
return None
print(f"Max attempts ({max_attempts}) reached. Job may still be processing.")
return None
# Usage
job_id = "04a7a6d8-5ef7-465a-b22a-8a98e7104dd9"
result = poll_split_job(job_id, "your-api-key")
if result:
print("Split job finished.")
```
```javascript JavaScript (with polling) theme={null}
async function pollSplitJob(jobId, apiKey, maxAttempts = 60, sleepInterval = 5000) {
/**
* Poll a splitting job until completion or failure
*
* @param {string} jobId - The job ID returned from split request
* @param {string} apiKey - Your API key
* @param {number} maxAttempts - Maximum number of polling attempts
* @param {number} sleepInterval - Milliseconds to wait between polls
* @returns {Promise
```json Completed Job theme={null}
{
"job_id": "c8a86841-beb1-4d00-ac4f-2f9fb9de9d5a",
"status": "completed",
"progress": "Starting document splitting...",
"file_url": "https://example-bucket.s3.amazonaws.com/user_uploads/...?AWSAccessKeyId=...&Signature=...&Expires=...",
"file_name": "mixed_documents.pdf",
"parameters": {
"classes": ["Invoice", "Receipt", "Contract"],
"category_descriptions": {
"Invoice": "Financial invoices with itemized charges",
"Receipt": "Purchase receipts"
},
"enable_reordering": false,
"page_count": 10
},
"result": {
"success": true,
"message": "Successfully split PDF into 3 files",
"files": [
{
"name": "Invoice.pdf",
"path": "Invoice.pdf",
"type": "file",
"fileId": "580e091d-c354-4558-8318-89e600346691",
"full_path": "https://example-bucket.s3.amazonaws.com/files/...?AWSAccessKeyId=...&Signature=...&Expires=...",
"confidence_score": 0.999147126422347
},
{
"name": "Contract.pdf",
"path": "Contract.pdf",
"type": "file",
"fileId": "680e091d-c354-4558-8318-89e600346692",
"full_path": "https://example-bucket.s3.amazonaws.com/files/...?AWSAccessKeyId=...&Signature=...&Expires=...",
"confidence_score": 0.987654321098765
}
]
},
"error": null
}
```
```json Processing Job theme={null}
{
"job_id": "c8a86841-beb1-4d00-ac4f-2f9fb9de9d5a",
"status": "processing",
"progress": "Starting document splitting...",
"file_url": "https://example-bucket.s3.amazonaws.com/user_uploads/...?AWSAccessKeyId=...&Signature=...&Expires=...",
"file_name": "mixed_documents.pdf",
"parameters": {
"classes": ["Invoice", "Receipt", "Contract"],
"category_descriptions": {},
"enable_reordering": false,
"page_count": 10
},
"result": null,
"error": null
}
```
```json Failed Job theme={null}
{
"job_id": "c8a86841-beb1-4d00-ac4f-2f9fb9de9d5a",
"status": "failed",
"progress": "Starting document splitting...",
"file_url": "https://example-bucket.s3.amazonaws.com/user_uploads/...?AWSAccessKeyId=...&Signature=...&Expires=...",
"file_name": "mixed_documents.pdf",
"parameters": {
"classes": ["Invoice", "Receipt", "Contract"],
"category_descriptions": {},
"enable_reordering": false,
"page_count": 10
},
"result": null,
"error": "Failed to process document: Invalid PDF format"
}
```
# Split Document
Source: https://docs.unsiloed.ai/docs/api-reference/splitting/split-document
api-reference/openapi.json POST /splitter
Split PDF documents by classifying pages into different categories
## Overview
The Split Document endpoint analyzes PDF pages, classifies them into predefined categories, and creates separate PDF files for each category. This is ideal for processing mixed document batches like scanned files containing invoices, contracts, and reports.
The endpoint processes documents asynchronously via a job-based system. It returns a job\_id immediately and processes the document in the background. Poll the status endpoint to retrieve results when complete.
## Request
The PDF file to split. Either file or file\_url must be provided; sending both returns a 400.
URL to a PDF file to split. Either file or file\_url must be provided.
JSON string containing array of category objects with name and optional description (e.g., `[{"name":"invoice","description":"Financial invoices"}]`). Descriptions help the classifier disambiguate similar categories. Categories that match no pages are skipped; no file is created for them.
Reorder pages within each category after classification, using content and page numbers to infer the logical order. Only applied to categories that match more than one page.
## Response
The endpoint returns HTTP 200 with the job identifier:
Unique identifier for the splitting job
Current status of the job ("processing")
Remaining API quota after this request
## Split Result
When the job completes, `GET /splitter/{job_id}` returns the split files inside its `result` object. The fields below describe that `result` object:
Whether the splitting operation succeeded
Descriptive message about the splitting operation
Array of split PDF files with their metadata
Category-derived filename of the split PDF (e.g., "invoice.pdf" for the "invoice" category)
Unique identifier for the file in storage
File type (always "file")
Relative path to the file in storage
Presigned download URL for the split PDF file. Expires roughly an hour after the response is generated, so download the file straight away rather than storing the URL. Re-issue `GET /splitter/{job_id}` to get a fresh URL.
Classification confidence for this file (0-1), averaged across all pages assigned to the category
## Request Examples
```bash cURL theme={null}
curl -X POST "https://prod.visionapi.unsiloed.ai/splitter" \
-H "api-key: your-api-key" \
-F "file=@mixed_documents.pdf" \
-F 'categories=[{"name":"invoice","description":"Business invoices with itemized charges"},{"name":"contract","description":"Legal agreements and binding documents"}]'
```
```python Python theme={null}
import requests
import json
url = "https://prod.visionapi.unsiloed.ai/splitter"
# Define categories with descriptions for better accuracy
categories = [
{"name": "invoice", "description": "Business invoices with itemized charges and payment terms"},
{"name": "contract", "description": "Legal agreements, terms of service, and binding documents"},
{"name": "report", "description": "Analytical reports, summaries, and data presentations"},
{"name": "letter", "description": "Correspondence, memos, and communication documents"}
]
files = {"file": ("mixed_documents.pdf", open("mixed_documents.pdf", "rb"), "application/pdf")}
data = {"categories": json.dumps(categories)}
headers = {"api-key": "your-api-key"}
response = requests.post(
url,
files=files,
data=data,
headers=headers
)
if response.status_code == 200:
result = response.json()
print(f"Job ID: {result['job_id']}")
print(f"Status: {result['status']}")
print(f"Quota Remaining: {result['quota_remaining']}")
else:
print("Error:", response.status_code, response.text)
# Close file
files["file"][1].close()
```
```javascript JavaScript theme={null}
const formData = new FormData();
formData.append('file', fileInput.files[0]);
const categories = [
{name: "invoice", description: "Business invoices with itemized charges"},
{name: "contract", description: "Legal agreements and binding documents"},
{name: "report", description: "Analytical reports and data presentations"}
];
formData.append('categories', JSON.stringify(categories));
const response = await fetch('https://prod.visionapi.unsiloed.ai/splitter', {
method: 'POST',
headers: {
'api-key': 'your-api-key'
},
body: formData
});
if (response.ok) {
const result = await response.json();
console.log('Job ID:', result.job_id);
console.log('Status:', result.status);
console.log('Quota Remaining:', result.quota_remaining);
} else {
console.error('Split failed:', response.status, await response.text());
}
```
## Response Examples
```json Split Result (from GET /splitter/{job_id}) theme={null}
{
"success": true,
"message": "Successfully split PDF into 2 files",
"files": [
{
"name": "invoice.pdf",
"fileId": "d079d09f-201c-4420-a50a-b25678a71ae9",
"type": "file",
"path": "invoice.pdf",
"full_path": "https://example-bucket.s3.amazonaws.com/files/ef3ec356-b407-4f9f-ac8f-0dfdef9034c0_invoice.pdf?AWSAccessKeyId=...&Signature=...&Expires=...",
"confidence_score": 0.8
},
{
"name": "contract.pdf",
"fileId": "320616cc-8dfd-4b8a-8474-8e7a42d9e287",
"type": "file",
"path": "contract.pdf",
"full_path": "https://example-bucket.s3.amazonaws.com/files/dfaa5d30-6955-4a69-9c69-7e3c4efd8450_contract.pdf?AWSAccessKeyId=...&Signature=...&Expires=...",
"confidence_score": 0.8
}
]
}
```
```json Error Response - Invalid File theme={null}
{
"detail": "File must be a PDF"
}
```
```json Error Response - Invalid Categories theme={null}
{
"detail": "Categories must be a JSON array"
}
```
```json Error Response - Service Unavailable theme={null}
{
"detail": "Failed to start split job. Please retry."
}
```
# Process Documents in Batches
Source: https://docs.unsiloed.ai/docs/cookbooks/batch-extraction
Extract the same fields from a folder of documents concurrently with Python and the Unsiloed API.
An extraction job spends most of its lifetime waiting for the API to process the document. If we submit one file, wait for it to finish, and only then submit the next, that waiting time accumulates across the whole folder.
This recipe processes several documents at once. A small Python thread pool gives each document the same schema, submits it to `/v2/extract`, polls its job, and collects the result. One failed document returns an error without stopping the rest of the batch.
This recipe builds on the [Extraction quickstart](/docs/document-processing/extraction/quickstart). Read that first if you want a detailed explanation of the request, polling flow, or response fields for one document.
## What We'll Build
A Python script that:
1. Finds every PDF in a `documents/` directory.
2. Processes documents concurrently, using four workers by default.
3. Submits each PDF with one shared extraction schema.
4. Polls each job until it completes or fails.
5. Writes all results and per-file errors to `results.json`.
Our example uses fund fact sheets, but the batch-processing pattern is not specific to financial documents. Replace the PDFs and schema with one document family from your own workflow.
Save this as `batch_extract.py`, place your PDFs in `documents/`, and set `UNSILOED_API_KEY` before running it.
```python batch_extract.py theme={null}
from concurrent.futures import ThreadPoolExecutor
import json
import os
from pathlib import Path
import time
import requests
API_KEY = os.environ["UNSILOED_API_KEY"]
BASE_URL = "https://prod.visionapi.unsiloed.ai"
DOCUMENT_DIR = Path("documents")
OUTPUT_PATH = Path("results.json")
MAX_WORKERS = 4
POLL_SECONDS = 3
MAX_POLLS = 100
SCHEMA = {
"type": "object",
"properties": {
"fund_name": {
"type": "string",
"description": "Full name of the fund",
},
"ticker": {
"type": "string",
"description": "Primary fund ticker",
},
"report_date": {
"type": "string",
"description": "Fact sheet as-of date",
},
"gross_expense_ratio": {
"type": "string",
"description": "Expense ratio with %",
},
},
"required": [
"fund_name",
"ticker",
"report_date",
],
"additionalProperties": False,
}
def submit_document(path):
"""Upload a document and return its job ID."""
with path.open("rb") as document:
response = requests.post(
f"{BASE_URL}/v2/extract",
headers={"api-key": API_KEY},
files={
"pdf_file": (
path.name,
document,
"application/pdf",
)
},
data={
"schema_data": json.dumps(SCHEMA),
"model": "gamma",
"enable_citations": "true",
},
timeout=60,
)
response.raise_for_status()
return response.json()["job_id"]
def wait_for_result(job_id):
"""Poll one job until it completes or fails."""
for _ in range(MAX_POLLS):
response = requests.get(
f"{BASE_URL}/extract/{job_id}",
headers={"api-key": API_KEY},
timeout=30,
)
response.raise_for_status()
job = response.json()
status = str(job.get("status", "")).lower()
if status in ("completed", "review"):
return job["result"]
if status in ("failed", "cancelled"):
message = job.get("error")
if not message:
message = f"job {status}"
raise RuntimeError(message)
time.sleep(POLL_SECONDS)
raise TimeoutError(f"Job {job_id} timed out")
def extract_document(path):
"""Submit a document and return its result."""
job_id = submit_document(path)
result = wait_for_result(job_id)
return {
"file": path.name,
"status": "completed",
"job_id": job_id,
"result": result,
}
def process_document(path):
"""Turn an exception into a per-file error."""
try:
return extract_document(path)
except Exception as error:
return {
"file": path.name,
"status": "failed",
"error": str(error),
}
documents = sorted(
path
for path in DOCUMENT_DIR.iterdir()
if path.suffix.lower() == ".pdf"
)
if not documents:
raise RuntimeError("No PDFs in documents/")
print(
f"Found {len(documents)} documents. "
f"Using {MAX_WORKERS} concurrent workers..."
)
started = time.perf_counter()
with ThreadPoolExecutor(MAX_WORKERS) as pool:
mapped = pool.map(process_document, documents)
results = list(mapped)
output = json.dumps(results, indent=2) + "\n"
OUTPUT_PATH.write_text(output)
completed = sum(
result["status"] == "completed"
for result in results
)
failed = len(results) - completed
print(f"Completed: {completed}")
print(f"Failed: {failed}")
print(f"Wrote {OUTPUT_PATH}")
elapsed = time.perf_counter() - started
print(f"Elapsed: {elapsed:.1f} seconds")
```
## Step 1: Set Up Your Documents and Dependency
We'll create a project directory, add a group of similar PDFs, and install the only Python package the script needs.
### 1.1 Get an Unsiloed API Key
Create an API key in the [Unsiloed dashboard](https://app.unsiloed.ai), then set it in your shell:
```bash theme={null}
export UNSILOED_API_KEY="your-api-key"
```
The script reads the key from the environment, so it never needs to appear in the source file.
### 1.2 Add PDFs to the Input Directory
Create a project directory with a `documents/` subdirectory:
```bash theme={null}
mkdir batch-extraction
cd batch-extraction
mkdir documents
```
Add the PDFs you want to process to `documents/`. The files should belong to the same document family because the script applies one schema to all of them. A batch might contain invoices from different suppliers, bank statements from different institutions, or claim forms from different customers.
For this example, we used these public fund fact sheets:
* [State Street SPDR S\&P 500 ETF Trust](https://www.ssga.com/library-content/products/factsheets/etfs/us/factsheet-us-en-spy.pdf)
* [iShares Core S\&P 500 ETF](https://www.ishares.com/us/literature/fact-sheet/ivv-ishares-core-s-p-500-etf-fund-fact-sheet-en-us.pdf)
* [Oakmark Fund](https://oakmark.com/wp-content/uploads/sites/3/documents/MXFundFactSheet.pdf)
The State Street document below contains the kind of repeated fields we want to collect from every fact sheet: a fund name, ticker, report date, and expense ratio.
The example links point to live issuer documents, so their contents can change. You don't need these exact PDFs to follow the recipe.
### 1.3 Install Requests
Create a virtual environment and install Requests:
```bash theme={null}
python3 -m venv .venv
source .venv/bin/activate
pip install requests
```
On Windows PowerShell, activate the environment with `.venv\Scripts\Activate.ps1`.
## Step 2: Configure the Batch and Its Schema
Create a file named `batch_extract.py` in the project directory. We'll build it in three pieces: imports, batch settings, and the schema for our example documents.
### 2.1 Import the Python Modules
Add these imports at the top of `batch_extract.py`:
```python batch_extract.py theme={null}
from concurrent.futures import ThreadPoolExecutor
import json
import os
from pathlib import Path
import time
import requests
```
Python includes the thread pool, JSON, path, environment, and timing modules. Requests is the only external dependency.
### 2.2 Set the Batch Limits and File Locations
In `batch_extract.py`, add the following configuration immediately below the imports:
```python batch_extract.py theme={null}
API_KEY = os.environ["UNSILOED_API_KEY"]
BASE_URL = "https://prod.visionapi.unsiloed.ai"
DOCUMENT_DIR = Path("documents")
OUTPUT_PATH = Path("results.json")
MAX_WORKERS = 4
POLL_SECONDS = 3
MAX_POLLS = 100
```
This example uses four workers as a conservative starting point so the script behaves predictably across different machines, file sizes, and API plans. Raise or lower `MAX_WORKERS` to suit your workload and account limits.
The polling settings allow each job approximately five minutes to finish. They control how long the script waits for a result, not how many documents Unsiloed can process.
### 2.3 Define the Schema for Your Documents
In `batch_extract.py`, add a `SCHEMA` dictionary immediately below the configuration:
```python batch_extract.py theme={null}
SCHEMA = {
"type": "object",
"properties": {
"fund_name": {
"type": "string",
"description": "Full name of the fund",
},
"ticker": {
"type": "string",
"description": "Primary fund ticker",
},
"report_date": {
"type": "string",
"description": "Fact sheet as-of date",
},
"gross_expense_ratio": {
"type": "string",
"description": "Expense ratio with %",
},
},
"required": [
"fund_name",
"ticker",
"report_date",
],
"additionalProperties": False,
}
```
This schema belongs to our fund-fact-sheet example. Replace the property names, descriptions, and required fields with the data shared by your own documents. If you're processing invoices, for example, you might request `invoice_number`, `vendor_name`, `invoice_date`, and `total_due` instead.
Only require a field when it should appear in every valid document. We keep `gross_expense_ratio` optional because some fact-sheet formats may omit it.
## Step 3: Submit and Poll One Document
Before adding concurrency, we'll define the three small functions that process one PDF from upload to completed result.
### 3.1 Submit a PDF for Extraction
In `batch_extract.py`, add `submit_document()` immediately below `SCHEMA`:
```python batch_extract.py theme={null}
def submit_document(path):
"""Upload a document and return its job ID."""
with path.open("rb") as document:
response = requests.post(
f"{BASE_URL}/v2/extract",
headers={"api-key": API_KEY},
files={
"pdf_file": (
path.name,
document,
"application/pdf",
)
},
data={
"schema_data": json.dumps(SCHEMA),
"model": "gamma",
"enable_citations": "true",
},
timeout=60,
)
response.raise_for_status()
return response.json()["job_id"]
```
The function opens one PDF, sends it with the shared schema, and returns the job ID from the submission response. The `gamma` model is the default, recommended Unsiloed extraction tier. See the [`model` parameter](/docs/api-reference/extraction/extract-data#body-model) for the available tiers.
The `enable_citations` option adds a page number and bounding box to each extracted field, so a reviewer can see where every value came from. Confidence scores come back either way.
### 3.2 Wait for the Job to Finish
In `batch_extract.py`, add `wait_for_result()` immediately below `submit_document()`:
```python batch_extract.py theme={null}
def wait_for_result(job_id):
"""Poll one job until it completes or fails."""
for _ in range(MAX_POLLS):
response = requests.get(
f"{BASE_URL}/extract/{job_id}",
headers={"api-key": API_KEY},
timeout=30,
)
response.raise_for_status()
job = response.json()
status = str(job.get("status", "")).lower()
if status in ("completed", "review"):
return job["result"]
if status in ("failed", "cancelled"):
message = job.get("error")
if not message:
message = f"job {status}"
raise RuntimeError(message)
time.sleep(POLL_SECONDS)
raise TimeoutError(f"Job {job_id} timed out")
```
The function checks the result endpoint every three seconds. It returns the extraction result for a completed job and raises an error for a failed, cancelled, or timed-out job.
### 3.3 Combine Submission and Polling
In `batch_extract.py`, add `extract_document()` immediately below `wait_for_result()`:
```python batch_extract.py theme={null}
def extract_document(path):
"""Submit a document and return its result."""
job_id = submit_document(path)
result = wait_for_result(job_id)
return {
"file": path.name,
"status": "completed",
"job_id": job_id,
"result": result,
}
```
This function gives the thread pool one operation to run for each path. It also adds the source filename and job ID to the result, so we can connect the output to its document later.
## Step 4: Run the Documents as a Batch
The remaining code isolates per-file errors, runs the worker pool, and writes one output file.
### 4.1 Keep One Failure From Stopping the Batch
In `batch_extract.py`, add `process_document()` immediately below `extract_document()`:
```python batch_extract.py theme={null}
def process_document(path):
"""Turn an exception into a per-file error."""
try:
return extract_document(path)
except Exception as error:
return {
"file": path.name,
"status": "failed",
"error": str(error),
}
```
Without this wrapper, an HTTP error or corrupt PDF could raise out of the thread pool and stop the script. Instead, that document becomes a failed result while the other workers continue.
### 4.2 Find the PDFs and Start the Worker Pool
In `batch_extract.py`, add the following code immediately below `process_document()`:
```python batch_extract.py theme={null}
documents = sorted(
path
for path in DOCUMENT_DIR.iterdir()
if path.suffix.lower() == ".pdf"
)
if not documents:
raise RuntimeError("No PDFs in documents/")
print(
f"Found {len(documents)} documents. "
f"Using {MAX_WORKERS} concurrent workers..."
)
started = time.perf_counter()
with ThreadPoolExecutor(MAX_WORKERS) as pool:
mapped = pool.map(process_document, documents)
results = list(mapped)
```
`ThreadPoolExecutor` calls `process_document()` once per PDF. The `MAX_WORKERS` setting determines how many of those calls the script runs concurrently.
Tune the worker count for your workload. More workers can improve throughput, but they also create more simultaneous uploads and polling requests, use more local memory, and may reach your account's rate limits. Without retry logic, the script records a rate-limit response as a failed document.
### 4.3 Save the Results and Print a Summary
At the end of `batch_extract.py`, immediately below the worker-pool block, add:
```python batch_extract.py theme={null}
output = json.dumps(results, indent=2) + "\n"
OUTPUT_PATH.write_text(output)
completed = sum(
result["status"] == "completed"
for result in results
)
failed = len(results) - completed
print(f"Completed: {completed}")
print(f"Failed: {failed}")
print(f"Wrote {OUTPUT_PATH}")
elapsed = time.perf_counter() - started
print(f"Elapsed: {elapsed:.1f} seconds")
```
The script keeps the complete API response for each successful document and the error message for each failed one. The terminal summary gives us a quick check without hiding the per-file details in `results.json`.
## Step 5: Run the Batch and Inspect the Results
Run the completed script from the project directory:
```bash theme={null}
python batch_extract.py
```
A successful run should print output shaped like this:
```text theme={null}
Found 3 documents. Using 4 concurrent workers...
Completed: 3
Failed: 0
Wrote results.json
Elapsed: 33.5 seconds
```
Timing depends on document size, complexity, model load, worker count, and account limits, so your elapsed time should differ.
Open `results.json` to inspect every API result:
```bash theme={null}
python -m json.tool results.json | less
```
Each successful entry keeps the complete extraction response:
```json theme={null}
{
"file": "spy.pdf",
"status": "completed",
"job_id": "6464ef7b-a554-4199-a549-15b6d82d0b49",
"result": {
"ticker": {
"value": "SPY",
"score": {
"grounding_score": 0.995,
"extraction_score": 0.995
},
"citation": {
"page": 1,
"bbox": [502, 40, 550, 71],
"page_width": 612.0,
"page_height": 792.0
}
}
}
}
```
Keeping the scores and citations lets the next stage flag uncertain values and show a reviewer where each value came from. Don't flatten the response to values alone unless the downstream system no longer needs that evidence.
## Where to Take This Next
The pattern stays the same when you replace the sample documents:
* Put one document family in the input directory.
* Describe its shared fields in one schema.
* Tune the worker count for your throughput target, account limits, file sizes, and available memory.
* Keep each file's error separate from the rest of the batch.
This example resubmits every PDF each time it runs. For a scheduled or high-volume production pipeline, add persistent job IDs, retries with backoff, and a queue. Those solve operational problems, but they aren't required to understand or use the core batch-processing pattern.
If the input directory contains unrelated document types, classify and route them before extraction. The [Sort and Extract cookbook](/docs/cookbooks/sort-and-extract) shows how to select a schema based on document type.
Read values, confidence scores, and citations from each completed result.
Route different document types to different schemas before processing them.
# Check Onboarding Documents for Consistency
Source: https://docs.unsiloed.ai/docs/cookbooks/kyc-document-consistency
Classify an identity document and two proofs of address, extract the relevant fields with citations, and return a deterministic PASS, REVIEW, or FAIL decision.
An onboarding flow often receives several documents that should describe the same person: an identity document, a utility bill, and a bank statement. Reading each file is only half the job. We also need to check that the names and addresses agree, confirm that the identity document hasn't expired, and stop when a critical field can't be read confidently.
This recipe classifies three uploaded documents, applies a type-specific extraction schema to each one, and runs deterministic consistency checks over the results. It returns `PASS`, `REVIEW`, or `FAIL`, with confidence scores and source citations preserved for every extracted field.
This workflow checks consistency and basic validity, not authenticity. It doesn't detect forged documents, verify security features, perform face matching, or provide liveness checks. A convincing fake with internally consistent details can pass these rules.
## What We'll Build
A script that:
1. Classifies each file as an identity document, utility bill, bank statement, or other document.
2. Extracts the fields required for that document type with citations enabled.
3. Normalizes names and addresses before comparing them across files.
4. Checks the identity document's expiry date.
5. Returns `FAIL` for a mismatch or expired document, `REVIEW` for uncertain data, and `PASS` only when every check clears.
The three documents follow the same pipeline, while extraction preserves the evidence used by the decision rules:
We'll use a public fictional Luxembourg passport plus two synthetic proof-of-address documents for the same person. The passport number is intentionally obscured, so the verified outcome is `REVIEW`: all cross-document values match, but the workflow refuses to approve a field it can't ground.
Set `UNSILOED_API_KEY` in your environment and save the three sample files in the same directory as the script before running it.
```python check_kyc.py theme={null}
import concurrent.futures
import datetime
import itertools
import json
import mimetypes
import os
import re
import time
import unicodedata
from pathlib import Path
import requests
API_KEY = os.environ["UNSILOED_API_KEY"]
BASE_URL = "https://prod.visionapi.unsiloed.ai"
CONFIDENCE_THRESHOLD = 0.85
FILES = [
Path("passport_specimen.jpg"),
Path("utility_bill.png"),
Path("bank_statement.png"),
]
CATEGORIES = [
{
"name": "Identity Document",
"description": "A passport, national identity card, or driving licence that identifies a person",
},
{
"name": "Utility Bill",
"description": "An electricity, gas, water, internet, or municipal service bill showing a customer and service address",
},
{
"name": "Bank Statement",
"description": "An account statement from a bank showing an account holder, statement period, and transactions",
},
{
"name": "Other",
"description": "A document that does not match any of the other categories",
},
]
SCHEMAS = {
"Identity Document": {
"type": "object",
"properties": {
"full_name": {
"type": "string",
"description": "Full name as printed, combining given names and surname in natural reading order",
},
"document_number": {
"type": "string",
"description": "Passport or identity document number. Return null if the complete number is obscured or illegible. Do not substitute a CAN number",
},
"date_of_birth": {
"type": "string",
"description": "Date of birth normalized to ISO 8601 YYYY-MM-DD",
},
"date_of_expiry": {
"type": "string",
"description": "Date of expiry normalized to ISO 8601 YYYY-MM-DD",
},
"nationality": {
"type": "string",
"description": "Nationality as printed on the identity document",
},
},
"required": ["full_name", "date_of_birth", "date_of_expiry"],
"additionalProperties": False,
},
"Utility Bill": {
"type": "object",
"properties": {
"account_holder_name": {
"type": "string",
"description": "Name of the account holder or customer",
},
"service_address": {
"type": "string",
"description": "Complete service or billing address",
},
"provider": {
"type": "string",
"description": "Utility provider or company name",
},
"account_number": {
"type": "string",
"description": "Utility account number",
},
"amount_due": {
"type": "string",
"description": "Total amount due, including the printed currency symbol",
},
"billing_date": {
"type": "string",
"description": "Billing date normalized to ISO 8601 YYYY-MM-DD",
},
},
"required": ["account_holder_name", "service_address"],
"additionalProperties": False,
},
"Bank Statement": {
"type": "object",
"properties": {
"account_holder_name": {
"type": "string",
"description": "Name of the account holder",
},
"address": {
"type": "string",
"description": "Complete mailing address of the account holder",
},
"bank_name": {
"type": "string",
"description": "Bank or financial institution name",
},
"account_number": {
"type": "string",
"description": "Account number or IBAN",
},
"statement_period": {
"type": "string",
"description": "Statement period exactly as printed",
},
"closing_balance": {
"type": "string",
"description": "Closing or ending balance exactly as printed",
},
},
"required": ["account_holder_name", "address"],
"additionalProperties": False,
},
}
NAME_FIELDS = {
"Identity Document": "full_name",
"Utility Bill": "account_holder_name",
"Bank Statement": "account_holder_name",
}
ADDRESS_FIELDS = {
"Utility Bill": "service_address",
"Bank Statement": "address",
}
CRITICAL_FIELDS = {
"Identity Document": {"full_name", "document_number", "date_of_birth", "date_of_expiry"},
"Utility Bill": {"account_holder_name", "service_address"},
"Bank Statement": {"account_holder_name", "address"},
}
def wait_for(job_type, job_id):
path = "classify" if job_type == "classify" else "extract"
for _ in range(120):
response = requests.get(
f"{BASE_URL}/{path}/{job_id}",
headers={"api-key": API_KEY},
timeout=60,
)
response.raise_for_status()
job = response.json()
if job.get("status") in ("completed", "review"):
return job
if job.get("status") == "failed":
raise RuntimeError(job.get("error") or job.get("message") or f"{job_type} failed")
time.sleep(3)
raise TimeoutError(f"{job_type} job {job_id} did not finish within six minutes")
def upload(path, endpoint, file_field, data):
mime_type = mimetypes.guess_type(path.name)[0] or "application/octet-stream"
with path.open("rb") as document:
response = requests.post(
f"{BASE_URL}/{endpoint}",
headers={"api-key": API_KEY},
files={file_field: (path.name, document, mime_type)},
data=data,
timeout=90,
)
response.raise_for_status()
body = response.json()
if not body.get("job_id"):
raise RuntimeError(f"{endpoint} returned no job_id: {body}")
return body["job_id"]
def confidence(field):
score = field.get("score") or {}
if isinstance(score, (int, float)):
return float(score)
available = [score.get("grounding_score"), score.get("extraction_score")]
available = [float(value) for value in available if value is not None]
return min(available) if available else 0.0
def process_document(path):
classify_id = upload(
path,
"classify",
"pdf_file",
{"categories": json.dumps(CATEGORIES)},
)
classification_job = wait_for("classify", classify_id)
classification = classification_job["result"]
document_type = classification["classification"]
if document_type not in SCHEMAS:
raise RuntimeError(f"{path.name} classified as {document_type}; no extraction schema applies")
extract_id = upload(
path,
"v2/extract",
"pdf_file",
{
"schema_data": json.dumps(SCHEMAS[document_type]),
"schema_name": f"kyc-{document_type.lower().replace(' ', '-')}-v1",
"model": "gamma",
"enable_citations": "true",
},
)
extraction_job = wait_for("extract", extract_id)
fields = {}
for name, field in extraction_job["result"].items():
citation = field.get("citation") or {}
fields[name] = {
"value": field.get("value"),
"confidence": confidence(field),
"page": citation.get("page"),
}
return {
"file": path.name,
"type": document_type,
"classification_confidence": classification["confidence"],
"fields": fields,
"raw": {
"classification": classification_job,
"extraction": extraction_job,
},
}
def normalized_tokens(value):
if not value:
return set()
text = unicodedata.normalize("NFKD", str(value))
text = "".join(character for character in text if not unicodedata.combining(character))
text = re.sub(r"[^a-z0-9]+", " ", text.lower())
return {token for token in text.split() if token}
def token_match(left, right, threshold=0.8):
left_tokens = normalized_tokens(left)
right_tokens = normalized_tokens(right)
if not left_tokens or not right_tokens:
return None
overlap = len(left_tokens & right_tokens) / min(len(left_tokens), len(right_tokens))
return overlap >= threshold
processed = []
errors = []
with concurrent.futures.ThreadPoolExecutor(max_workers=3) as pool:
futures = {pool.submit(process_document, path): path for path in FILES}
for future in concurrent.futures.as_completed(futures):
path = futures[future]
try:
processed.append(future.result())
except Exception as error:
errors.append(f"{path.name}: {error}")
with open("kyc-raw.json", "w") as output:
json.dump({document["file"]: document["raw"] for document in processed}, output, indent=2)
documents = {document["type"]: document for document in processed}
checks = []
named_documents = [doc_type for doc_type in NAME_FIELDS if doc_type in documents]
for left_type, right_type in itertools.combinations(named_documents, 2):
left = documents[left_type]["fields"][NAME_FIELDS[left_type]]["value"]
right = documents[right_type]["fields"][NAME_FIELDS[right_type]]["value"]
checks.append({
"type": "Name",
"documents": [left_type, right_type],
"values": [left, right],
"match": token_match(left, right),
})
addressed_documents = [doc_type for doc_type in ADDRESS_FIELDS if doc_type in documents]
for left_type, right_type in itertools.combinations(addressed_documents, 2):
left = documents[left_type]["fields"][ADDRESS_FIELDS[left_type]]["value"]
right = documents[right_type]["fields"][ADDRESS_FIELDS[right_type]]["value"]
checks.append({
"type": "Address",
"documents": [left_type, right_type],
"values": [left, right],
"match": token_match(left, right, threshold=0.7),
})
flags = list(errors)
for document_type, document in documents.items():
if document["classification_confidence"] < CONFIDENCE_THRESHOLD:
flags.append(f"{document['file']}: low classification confidence")
for field_name in CRITICAL_FIELDS[document_type]:
field = document["fields"].get(field_name, {})
if field.get("value") in (None, "", "None") or field.get("confidence", 0) == 0:
flags.append(f"{document_type}.{field_name}: unreadable or not grounded")
elif field["confidence"] < CONFIDENCE_THRESHOLD:
flags.append(
f"{document_type}.{field_name}: confidence {field['confidence']:.3f} is below {CONFIDENCE_THRESHOLD}"
)
expired = False
identity = documents.get("Identity Document")
if identity:
expiry = identity["fields"].get("date_of_expiry", {}).get("value")
try:
expired = datetime.date.fromisoformat(expiry) < datetime.date.today()
except (TypeError, ValueError):
flags.append("Identity Document.date_of_expiry: missing or not valid ISO 8601")
mismatch = any(check["match"] is False for check in checks)
inconclusive = any(check["match"] is None for check in checks)
if inconclusive:
flags.append("At least one cross-document comparison was inconclusive")
if expired or mismatch:
decision = "FAIL"
reason = "identity document expired" if expired else "name or address mismatch"
elif flags:
decision = "REVIEW"
reason = "uncertain or unreadable data needs manual review"
else:
decision = "PASS"
reason = "all documents are consistent and clearly read"
result = {
"decision": decision,
"reason": reason,
"checks": checks,
"flags": flags,
"documents": {
document_type: {
"file": document["file"],
"classification_confidence": document["classification_confidence"],
"fields": document["fields"],
}
for document_type, document in documents.items()
},
}
with open("kyc-result.json", "w") as output:
json.dump(result, output, indent=2)
print(json.dumps({
"decision": result["decision"],
"reason": result["reason"],
"checks": result["checks"],
"flags": result["flags"],
}, indent=2))
```
Save this as `check-kyc.mjs`. It requires Node.js 18 or newer for the global `fetch`, `FormData`, and `Blob` APIs.
```javascript check-kyc.mjs theme={null}
import fs from "node:fs";
import path from "node:path";
const API_KEY = process.env.UNSILOED_API_KEY;
const BASE_URL = "https://prod.visionapi.unsiloed.ai";
const CONFIDENCE_THRESHOLD = 0.85;
if (!API_KEY) throw new Error("Set UNSILOED_API_KEY before running the script");
const FILES = ["passport_specimen.jpg", "utility_bill.png", "bank_statement.png"];
const CATEGORIES = [
{
name: "Identity Document",
description: "A passport, national identity card, or driving licence that identifies a person",
},
{
name: "Utility Bill",
description: "An electricity, gas, water, internet, or municipal service bill showing a customer and service address",
},
{
name: "Bank Statement",
description: "An account statement from a bank showing an account holder, statement period, and transactions",
},
{ name: "Other", description: "A document that does not match any of the other categories" },
];
const SCHEMAS = {
"Identity Document": {
type: "object",
properties: {
full_name: {
type: "string",
description: "Full name as printed, combining given names and surname in natural reading order",
},
document_number: {
type: "string",
description: "Passport or identity document number. Return null if the complete number is obscured or illegible. Do not substitute a CAN number",
},
date_of_birth: {
type: "string",
description: "Date of birth normalized to ISO 8601 YYYY-MM-DD",
},
date_of_expiry: {
type: "string",
description: "Date of expiry normalized to ISO 8601 YYYY-MM-DD",
},
nationality: { type: "string", description: "Nationality as printed on the identity document" },
},
required: ["full_name", "date_of_birth", "date_of_expiry"],
additionalProperties: false,
},
"Utility Bill": {
type: "object",
properties: {
account_holder_name: { type: "string", description: "Name of the account holder or customer" },
service_address: { type: "string", description: "Complete service or billing address" },
provider: { type: "string", description: "Utility provider or company name" },
account_number: { type: "string", description: "Utility account number" },
amount_due: { type: "string", description: "Total amount due, including the printed currency symbol" },
billing_date: { type: "string", description: "Billing date normalized to ISO 8601 YYYY-MM-DD" },
},
required: ["account_holder_name", "service_address"],
additionalProperties: false,
},
"Bank Statement": {
type: "object",
properties: {
account_holder_name: { type: "string", description: "Name of the account holder" },
address: { type: "string", description: "Complete mailing address of the account holder" },
bank_name: { type: "string", description: "Bank or financial institution name" },
account_number: { type: "string", description: "Account number or IBAN" },
statement_period: { type: "string", description: "Statement period exactly as printed" },
closing_balance: { type: "string", description: "Closing or ending balance exactly as printed" },
},
required: ["account_holder_name", "address"],
additionalProperties: false,
},
};
const NAME_FIELDS = {
"Identity Document": "full_name",
"Utility Bill": "account_holder_name",
"Bank Statement": "account_holder_name",
};
const ADDRESS_FIELDS = { "Utility Bill": "service_address", "Bank Statement": "address" };
const CRITICAL_FIELDS = {
"Identity Document": ["full_name", "document_number", "date_of_birth", "date_of_expiry"],
"Utility Bill": ["account_holder_name", "service_address"],
"Bank Statement": ["account_holder_name", "address"],
};
const sleep = (milliseconds) => new Promise((resolve) => setTimeout(resolve, milliseconds));
async function waitFor(jobType, jobId) {
const endpoint = jobType === "classify" ? "classify" : "extract";
for (let attempt = 0; attempt < 120; attempt++) {
const response = await fetch(`${BASE_URL}/${endpoint}/${jobId}`, {
headers: { "api-key": API_KEY },
});
if (!response.ok) throw new Error(`${jobType} poll failed: HTTP ${response.status} ${await response.text()}`);
const job = await response.json();
if (["completed", "review"].includes(job.status)) return job;
if (job.status === "failed") throw new Error(job.error || job.message || `${jobType} failed`);
await sleep(3000);
}
throw new Error(`${jobType} job ${jobId} did not finish within six minutes`);
}
function mimeType(fileName) {
const extension = path.extname(fileName).toLowerCase();
if (extension === ".jpg" || extension === ".jpeg") return "image/jpeg";
if (extension === ".png") return "image/png";
if (extension === ".pdf") return "application/pdf";
return "application/octet-stream";
}
async function upload(fileName, endpoint, fileField, data) {
const form = new FormData();
form.append(
fileField,
new Blob([fs.readFileSync(fileName)], { type: mimeType(fileName) }),
path.basename(fileName),
);
for (const [key, value] of Object.entries(data)) form.append(key, value);
const response = await fetch(`${BASE_URL}/${endpoint}`, {
method: "POST",
headers: { "api-key": API_KEY },
body: form,
});
if (!response.ok) throw new Error(`${endpoint} failed: HTTP ${response.status} ${await response.text()}`);
const body = await response.json();
if (!body.job_id) throw new Error(`${endpoint} returned no job_id: ${JSON.stringify(body)}`);
return body.job_id;
}
function confidence(field) {
const score = field.score || {};
if (typeof score === "number") return score;
const available = [score.grounding_score, score.extraction_score]
.filter((value) => value != null)
.map(Number);
return available.length ? Math.min(...available) : 0;
}
async function processDocument(fileName) {
const classifyId = await upload(fileName, "classify", "pdf_file", {
categories: JSON.stringify(CATEGORIES),
});
const classificationJob = await waitFor("classify", classifyId);
const classification = classificationJob.result;
const documentType = classification.classification;
if (!SCHEMAS[documentType]) {
throw new Error(`${fileName} classified as ${documentType}; no extraction schema applies`);
}
const extractId = await upload(fileName, "v2/extract", "pdf_file", {
schema_data: JSON.stringify(SCHEMAS[documentType]),
schema_name: `kyc-${documentType.toLowerCase().replaceAll(" ", "-")}-v1`,
model: "gamma",
enable_citations: "true",
});
const extractionJob = await waitFor("extract", extractId);
const fields = Object.fromEntries(
Object.entries(extractionJob.result).map(([name, field]) => [name, {
value: field.value,
confidence: confidence(field),
page: field.citation?.page ?? null,
}]),
);
return {
file: fileName,
type: documentType,
classification_confidence: classification.confidence,
fields,
raw: { classification: classificationJob, extraction: extractionJob },
};
}
function normalizedTokens(value) {
if (!value) return new Set();
const normalized = String(value)
.normalize("NFKD")
.replace(/\p{Diacritic}/gu, "")
.toLowerCase()
.replace(/[^a-z0-9]+/g, " ")
.trim();
return new Set(normalized ? normalized.split(/\s+/) : []);
}
function tokenMatch(left, right, threshold = 0.8) {
const leftTokens = normalizedTokens(left);
const rightTokens = normalizedTokens(right);
if (!leftTokens.size || !rightTokens.size) return null;
const overlap = [...leftTokens].filter((token) => rightTokens.has(token)).length;
return overlap / Math.min(leftTokens.size, rightTokens.size) >= threshold;
}
function pairs(values) {
const output = [];
for (let left = 0; left < values.length; left++) {
for (let right = left + 1; right < values.length; right++) {
output.push([values[left], values[right]]);
}
}
return output;
}
const processed = await Promise.all(FILES.map(async (fileName) => {
try {
return await processDocument(fileName);
} catch (error) {
return { file: fileName, error: error.message };
}
}));
const completed = processed.filter((document) => !document.error);
fs.writeFileSync(
"kyc-raw.json",
JSON.stringify(Object.fromEntries(completed.map((document) => [document.file, document.raw])), null, 2),
);
const documents = Object.fromEntries(completed.map((document) => [document.type, document]));
const checks = [];
const namedDocuments = Object.keys(NAME_FIELDS).filter((documentType) => documents[documentType]);
for (const [leftType, rightType] of pairs(namedDocuments)) {
const left = documents[leftType].fields[NAME_FIELDS[leftType]].value;
const right = documents[rightType].fields[NAME_FIELDS[rightType]].value;
checks.push({
type: "Name",
documents: [leftType, rightType],
values: [left, right],
match: tokenMatch(left, right),
});
}
const addressedDocuments = Object.keys(ADDRESS_FIELDS).filter((documentType) => documents[documentType]);
for (const [leftType, rightType] of pairs(addressedDocuments)) {
const left = documents[leftType].fields[ADDRESS_FIELDS[leftType]].value;
const right = documents[rightType].fields[ADDRESS_FIELDS[rightType]].value;
checks.push({
type: "Address",
documents: [leftType, rightType],
values: [left, right],
match: tokenMatch(left, right, 0.7),
});
}
const flags = processed.filter((document) => document.error).map((document) => `${document.file}: ${document.error}`);
for (const [documentType, document] of Object.entries(documents)) {
if (document.classification_confidence < CONFIDENCE_THRESHOLD) {
flags.push(`${document.file}: low classification confidence`);
}
for (const fieldName of CRITICAL_FIELDS[documentType]) {
const field = document.fields[fieldName] || {};
if (field.value == null || field.value === "" || field.value === "None" || !field.confidence) {
flags.push(`${documentType}.${fieldName}: unreadable or not grounded`);
} else if (field.confidence < CONFIDENCE_THRESHOLD) {
flags.push(
`${documentType}.${fieldName}: confidence ${field.confidence.toFixed(3)} is below ${CONFIDENCE_THRESHOLD}`,
);
}
}
}
let expired = false;
const expiry = documents["Identity Document"]?.fields.date_of_expiry?.value;
if (expiry) {
const expiryDate = new Date(`${expiry}T00:00:00Z`);
if (Number.isNaN(expiryDate.getTime())) {
flags.push("Identity Document.date_of_expiry: missing or not valid ISO 8601");
} else {
expired = expiryDate < new Date();
}
} else if (documents["Identity Document"]) {
flags.push("Identity Document.date_of_expiry: missing or not valid ISO 8601");
}
const mismatch = checks.some((check) => check.match === false);
const inconclusive = checks.some((check) => check.match == null);
if (inconclusive) flags.push("At least one cross-document comparison was inconclusive");
let decision;
let reason;
if (expired || mismatch) {
decision = "FAIL";
reason = expired ? "identity document expired" : "name or address mismatch";
} else if (flags.length) {
decision = "REVIEW";
reason = "uncertain or unreadable data needs manual review";
} else {
decision = "PASS";
reason = "all documents are consistent and clearly read";
}
const result = {
decision,
reason,
checks,
flags,
documents: Object.fromEntries(Object.entries(documents).map(([documentType, document]) => [
documentType,
{
file: document.file,
classification_confidence: document.classification_confidence,
fields: document.fields,
},
])),
};
fs.writeFileSync("kyc-result.json", JSON.stringify(result, null, 2));
console.log(JSON.stringify({ decision, reason, checks, flags }, null, 2));
```
## Step 1: Set Up the Sample Documents
Before writing code, we'll download the tested sample set and configure the runtime.
### 1.1 Get an Unsiloed API Key
Get an API key from the [Unsiloed dashboard](https://app.unsiloed.ai), then export it as an environment variable:
```bash theme={null}
export UNSILOED_API_KEY="your-api-key"
```
### 1.2 Download the Fictional Onboarding Set
Create a new directory, then download the three files from the Unsiloed cookbook repository:
```bash theme={null}
curl -LO https://raw.githubusercontent.com/Unsiloed-AI/cookbook/b1d9c38fa5125eb148e14849d4c26d9845b31a97/kyc-app/samples/passport_specimen.jpg
curl -LO https://raw.githubusercontent.com/Unsiloed-AI/cookbook/b1d9c38fa5125eb148e14849d4c26d9845b31a97/kyc-app/samples/utility_bill.png
curl -LO https://raw.githubusercontent.com/Unsiloed-AI/cookbook/b1d9c38fa5125eb148e14849d4c26d9845b31a97/kyc-app/samples/bank_statement.png
```
The passport is a fictional Luxembourg government specimen for Ketty Maus. Some secure fields, including the complete passport number, are intentionally obscured. The utility bill and bank statement are synthetic documents created for this example with the same name and address.
### 1.3 Install the Python Dependency
The JavaScript version uses only Node.js built-ins. For Python, install `requests`:
```bash theme={null}
pip install requests
```
## Step 2: Classify Each Document Before Extraction
An intake folder doesn't reliably tell us which file is an ID or proof of address. We classify first, then use the result to select the right extraction schema.
Define mutually exclusive categories, including an `Other` fallback:
```json theme={null}
[
{
"name": "Identity Document",
"description": "A passport, national identity card, or driving licence that identifies a person"
},
{
"name": "Utility Bill",
"description": "A utility bill showing a customer and service address"
},
{
"name": "Bank Statement",
"description": "A bank statement showing an account holder, period, and transactions"
},
{
"name": "Other",
"description": "A document that does not match any other category"
}
]
```
The scripts submit all three files concurrently. In a verified run, the API classified every sample correctly with confidence above `0.9999999`.
## Step 3: Extract Fields With a Schema per Type
Each document answers different questions. A passport has a date of birth and expiry date; a utility bill has a service address; a bank statement has an account holder and statement period. A single catch-all schema would ask documents for fields they can't contain.
The scripts keep one schema per classification result and enable citations on every extraction:
```text theme={null}
Identity Document -> name, document number, birth date, expiry date, nationality
Utility Bill -> account holder, service address, provider, account, amount, billing date
Bank Statement -> account holder, address, bank, account, period, closing balance
```
Date descriptions request ISO 8601 output. The passport prints `16 02 2036`; extraction returns `2036-02-16`, which Python and JavaScript can compare without a list of locale-specific date formats.
We leave `document_number` out of the schema's `required` array. The sample's complete passport number isn't visible, so requiring it could pressure the extraction toward a guess. It remains a critical field in the decision rules, where a `null` value correctly triggers review.
## Step 4: Compare Confidence, Names, Addresses, and Expiry
The decision engine uses only deterministic code. The model reads each document, but it doesn't decide whether the person passes onboarding.
### 4.1 Gate Critical Fields by Confidence
Each extracted field has a grounding score and an extraction score. The scripts use the lower available value:
```text theme={null}
confidence = min(grounding_score, extraction_score)
```
This prevents a value with strong extraction confidence but weak source grounding from passing automatically. A critical field becomes a review item when it is missing, ungrounded, or below the example threshold of `0.85`.
Start with `0.85` for this example, then calibrate the threshold against documents your team has already reviewed. It isn't a universal compliance threshold.
### 4.2 Normalize Before Comparing
Names and addresses rarely use identical punctuation and spacing. The scripts lowercase the values, remove punctuation and diacritics, split them into tokens, and compare token overlap.
For the sample, these two addresses normalize to the same token set:
```text theme={null}
12, Rue de la Gare
L-1611 Luxembourg
12, Rue de la Gare, L-1611 Luxembourg
```
This is intentionally conservative. Production identity matching may need locale-specific address parsing, transliteration, aliases, and approved name-change handling.
### 4.3 Check the Expiry Date in Code
The ISO date lets us compare the identity document against today's date directly. An expired identity document returns `FAIL`; a missing or malformed expiry date returns `REVIEW`.
## Step 5: Return a Review Decision With Evidence
Run the completed script for your runtime:
Save the Python script from the full-script accordion as `check_kyc.py`, then run:
```bash theme={null}
python check_kyc.py
```
Save the JavaScript script from the full-script accordion as `check-kyc.mjs`, then run:
```bash theme={null}
node check-kyc.mjs
```
Both implementations write two files:
* `kyc-raw.json` preserves the complete classification and extraction responses, including job IDs and citations.
* `kyc-result.json` contains the normalized decision payload for the onboarding system.
### Sample Output
The live API classified every sample correctly and extracted all visible critical values. Every name comparison passed, and the utility-bill address matched the bank-statement address. The complete passport number is obscured, so the result stops for review:
```json theme={null}
{
"decision": "REVIEW",
"reason": "uncertain or unreadable data needs manual review",
"checks": [
{
"type": "Name",
"documents": ["Identity Document", "Utility Bill"],
"values": ["Ketty Maus", "Ketty Maus"],
"match": true
},
{
"type": "Name",
"documents": ["Identity Document", "Bank Statement"],
"values": ["Ketty Maus", "Ketty Maus"],
"match": true
},
{
"type": "Name",
"documents": ["Utility Bill", "Bank Statement"],
"values": ["Ketty Maus", "Ketty Maus"],
"match": true
},
{
"type": "Address",
"documents": ["Utility Bill", "Bank Statement"],
"values": [
"12, Rue de la Gare\nL-1611 Luxembourg",
"12, Rue de la Gare\nL-1611 Luxembourg"
],
"match": true
}
],
"flags": [
"Identity Document.document_number: unreadable or not grounded"
]
}
```
The passport-number extraction itself carries the evidence for that decision:
```json theme={null}
{
"value": null,
"score": {
"grounding_score": 0.0,
"extraction_score": 0.0
},
"citation": null
}
```
The workflow doesn't turn missing evidence into a guessed value. A reviewer can request a clearer image or a different identity document before the onboarding process continues.
## Where to Take This Next
For a production onboarding service, add your organization's policy checks around this core pipeline.
Useful extensions include:
* Checking whether proof-of-address documents are recent enough for your policy.
* Comparing a selfie against the identity-document portrait with an approved identity provider.
* Sending review items to an evidence viewer that draws each citation box on the source document.
* Recording reviewer corrections separately from the raw API response.
* Encrypting documents and deleting stored artifacts according to your retention policy.
Define document categories and route each file to the right workflow.
Add or refine fields for the identity documents your onboarding policy accepts.
Inspect value, confidence, citation, and job metadata fields.
# Cookbooks
Source: https://docs.unsiloed.ai/docs/cookbooks/overview
End-to-end recipes that combine parsing, extraction, classification, and splitting to solve real document-processing problems.
The [API reference](/docs/api-reference/parser/parse-document) and the [Document Processing](/docs/document-processing/parsing/parsing) guides cover each endpoint on its own. Cookbooks go the other way: each one starts from a problem a team has, then walks the full path from raw file to a result you can ship, often chaining several operations together.
Every recipe is a complete, copy-pasteable walkthrough built around real documents you can download and run against.
## Recipes
Each recipe solves one document-processing problem from start to finish.
Apply one extraction schema to a folder of similar documents and process several files concurrently.
Classify an identity document and two proofs of address, cross-check their fields, and route unreadable values to review.
Coming soon: extract line items and totals from an invoice and route uncertain values to a human.
Split a PDF holding several document types into one file per type, then run a type-specific extraction schema over each.
Find personal data with extraction, locate every occurrence with parsing, and black it out to produce a clean redacted PDF.
Coming soon: pull the same structured fields out of documents written in over 50 languages.
Coming soon: pull clauses, parties, and tracked changes out of a legal agreement.
Want a recipe for a use case that isn't here yet? Email [support@unsiloed.ai](mailto:support@unsiloed.ai) and tell us what you're building.
# Redact Sensitive Data From a Document
Source: https://docs.unsiloed.ai/docs/cookbooks/redact-sensitive-data
Find personal data in a PDF with extraction, locate every occurrence with parsing, and black it out to produce a clean redacted PDF.
Before a document leaves your hands, the sensitive parts often have to go: names, account numbers, email addresses, phone numbers. Doing it by hand is slow, and a single missed occurrence leaks the very data you meant to protect. Worse, the obvious shortcut, drawing a black box over the text in a PDF viewer, doesn't actually remove anything. The text still sits underneath, ready to be copied straight back out. Governments and law firms have leaked data exactly this way.
This recipe redacts properly. It uses extraction to work out what the sensitive values are, parsing to find every place they appear on the page, and then removes them, producing a `redacted.pdf` you can safely share. You don't have to know the values in advance; the extractor reads them off the document for you.
This recipe chains two endpoints covered on their own in the [Extract quickstart](/docs/document-processing/extraction/quickstart) and the [parsing guide](/docs/document-processing/parsing/parsing). Read those first if you want the detail on either call in isolation.
## What We'll Build
A script that:
1. Sends a document to `/v2/extract` with a schema describing the sensitive fields, and reads back their values.
2. Sends the same document to `/parse` to get every word on the page with its bounding box.
3. Matches each sensitive value against the parsed words to find every occurrence, not just the first.
4. Blacks out each match and writes a `redacted.pdf` with the underlying text removed.
We'll run it against a one-page account statement that repeats the same personal details several times. Grab the full script from the dropdown below if you'd rather skip the walkthrough.
Set `UNSILOED_API_KEY` in your environment and save the document as `document.pdf` in the same directory before running.
```python redact.py theme={null}
import json
import os
import re
import time
import fitz # PyMuPDF
import requests
API_KEY = os.environ["UNSILOED_API_KEY"]
BASE_URL = "https://prod.visionapi.unsiloed.ai"
# The kinds of information we want gone. Each description tells the extractor
# what to look for, so adapt these fields to your own document.
SCHEMA = {
"type": "object",
"properties": {
"full_name": {"type": "string", "description": "The full name of the account holder"},
"account_number": {"type": "string", "description": "The bank account number"},
"email": {"type": "string", "description": "The account holder's email address"},
"phone": {"type": "string", "description": "The account holder's phone number"},
},
"required": ["full_name", "account_number", "email", "phone"],
"additionalProperties": False,
}
def wait_for(url):
for _ in range(90): # roughly 6 minutes at 4 seconds per poll
job = requests.get(url, headers={"api-key": API_KEY}).json()
status = job.get("status")
if status in ("completed", "review", "Succeeded"):
return job
if status in ("failed", "Failed", "Cancelled"):
# The extractor reports its reason in "error"; the parser uses "message".
raise RuntimeError(job.get("error") or job.get("message") or f"job {status}")
time.sleep(4)
raise TimeoutError(f"Job at {url} did not finish in time")
# Step 1: ask the extractor what the sensitive values are.
with open("document.pdf", "rb") as f:
resp = requests.post(
f"{BASE_URL}/v2/extract",
headers={"api-key": API_KEY},
files={"pdf_file": ("document.pdf", f, "application/pdf")},
data={"schema_data": json.dumps(SCHEMA), "model": "gamma"},
)
resp.raise_for_status()
job = resp.json()
result = wait_for(f"{BASE_URL}/extract/{job['job_id']}")["result"]
sensitive = []
for field in SCHEMA["properties"]:
value = result[field]["value"]
if value is None:
print(f"Field '{field}' not found in the document, skipping.")
continue
sensitive.append(value)
print("Will redact:", sensitive)
# Step 2: parse the document to get every word with a page-relative box.
with open("document.pdf", "rb") as f:
resp = requests.post(
f"{BASE_URL}/parse",
headers={"api-key": API_KEY},
files={"file": ("document.pdf", f, "application/pdf")},
)
resp.raise_for_status()
job = resp.json()
parsed = wait_for(f"{BASE_URL}/parse/{job['job_id']}")
def normalize(text):
return re.sub(r"[^a-z0-9]", "", str(text).lower())
words = []
for chunk in parsed["chunks"]:
for seg in chunk["segments"]:
sx, sy = seg["bbox"]["left"], seg["bbox"]["top"]
for word in seg.get("ocr") or []:
norm = normalize(word["text"])
if not norm:
continue
b = word["bbox"]
words.append({
"norm": norm,
"page": seg["page_number"],
"page_width": seg["page_width"],
"x": sx + b["left"],
"width": b["width"],
# y + height lands on the text baseline reliably, even when the
# reported height is not, so we anchor redaction boxes to it.
"baseline": sy + b["top"] + b["height"],
})
# With no words there is nothing to match, and the script would "succeed"
# by writing an unredacted copy. Fail loudly instead.
if not words:
raise RuntimeError("Parse returned no OCR words; refusing to write an unredacted PDF")
# Step 3: find every occurrence of each value by normalized concatenation, so a
# value matches however OCR split it (hyphens, "@") or wrapped across lines.
def find_runs(value):
target = normalize(value)
runs = []
if len(target) < 3:
return runs
for i in range(len(words)):
acc = ""
for j in range(i, min(i + 10, len(words))):
if words[j]["page"] != words[i]["page"]:
break
acc += words[j]["norm"]
if acc == target:
runs.append(words[i:j + 1])
break
if not target.startswith(acc):
break
return runs
# Step 4: turn each run into one box per line, anchored to the baseline.
BAND_ABOVE, BAND_BELOW, PAD = 19, 6, 2
def line_boxes(run):
lines, current = [], [run[0]]
for word in run[1:]:
same_line = abs(word["baseline"] - current[-1]["baseline"]) <= 12
if same_line and word["x"] >= current[-1]["x"] - 2:
current.append(word)
else:
lines.append(current)
current = [word]
lines.append(current)
boxes = []
for line in lines:
baseline = max(w["baseline"] for w in line)
boxes.append((
min(w["x"] for w in line) - PAD,
baseline - BAND_ABOVE,
max(w["x"] + w["width"] for w in line) + PAD,
baseline + BAND_BELOW,
))
return boxes
doc = fitz.open("document.pdf")
boxes_by_page = {}
for value in sensitive:
for run in find_runs(value):
page_index = run[0]["page"] - 1
# Parse boxes are in render pixels; scale them back to PDF points.
scale = run[0]["page_width"] / doc[page_index].rect.width
for x0, y0, x1, y1 in line_boxes(run):
rect = fitz.Rect(x0 / scale, y0 / scale, x1 / scale, y1 / scale)
boxes_by_page.setdefault(page_index, []).append(rect)
# apply_redactions deletes any text that touches the redaction rect, even
# by a fraction of a point, so we delete with a rect shrunk clear of the
# neighboring lines, then draw the full-height box as plain ink afterwards.
for page_index, rects in boxes_by_page.items():
page = doc[page_index]
for rect in rects:
page.add_redact_annot(rect + (0, 2, 0, -2), fill=(0, 0, 0))
page.apply_redactions()
for rect in rects:
page.draw_rect(rect, fill=(0, 0, 0), color=None)
doc.save("redacted.pdf")
print("Wrote redacted.pdf")
```
Save this as `redact.mjs` or set `"type": "module"` in your `package.json`. Requires Node.js 18 or newer for the global `fetch`, `FormData`, and `Blob`.
```javascript redact.mjs theme={null}
import { readFileSync, writeFileSync } from "node:fs";
import { createCanvas } from "@napi-rs/canvas";
import { PDFDocument } from "pdf-lib";
import * as pdfjsLib from "pdfjs-dist/legacy/build/pdf.mjs";
const API_KEY = process.env.UNSILOED_API_KEY;
const BASE_URL = "https://prod.visionapi.unsiloed.ai";
// The kinds of information we want gone. Each description tells the extractor
// what to look for, so adapt these fields to your own document.
const SCHEMA = {
type: "object",
properties: {
full_name: { type: "string", description: "The full name of the account holder" },
account_number: { type: "string", description: "The bank account number" },
email: { type: "string", description: "The account holder's email address" },
phone: { type: "string", description: "The account holder's phone number" },
},
required: ["full_name", "account_number", "email", "phone"],
additionalProperties: false,
};
async function waitFor(url) {
for (let i = 0; i < 90; i++) { // roughly 6 minutes at 4 seconds per poll
const job = await (await fetch(url, { headers: { "api-key": API_KEY } })).json();
if (["completed", "review", "Succeeded"].includes(job.status)) return job;
// The extractor reports its reason in "error"; the parser uses "message".
if (["failed", "Failed", "Cancelled"].includes(job.status)) throw new Error(job.error || job.message || `job ${job.status}`);
await new Promise((r) => setTimeout(r, 4000));
}
throw new Error(`Job at ${url} did not finish in time`);
}
const bytes = readFileSync("document.pdf");
// Step 1: ask the extractor what the sensitive values are.
const extractForm = new FormData();
extractForm.append("pdf_file", new Blob([bytes], { type: "application/pdf" }), "document.pdf");
extractForm.append("schema_data", JSON.stringify(SCHEMA));
extractForm.append("model", "gamma");
const extractRes = await fetch(`${BASE_URL}/v2/extract`, { method: "POST", headers: { "api-key": API_KEY }, body: extractForm });
if (!extractRes.ok) throw new Error(`extract submit failed: HTTP ${extractRes.status} ${await extractRes.text()}`);
const extractJob = await extractRes.json();
const result = (await waitFor(`${BASE_URL}/extract/${extractJob.job_id}`)).result;
const sensitive = [];
for (const field of Object.keys(SCHEMA.properties)) {
const value = result[field].value;
if (value == null) {
console.log(`Field '${field}' not found in the document, skipping.`);
continue;
}
sensitive.push(value);
}
console.log("Will redact:", sensitive);
// Step 2: parse to get every word with a page-relative box.
const parseForm = new FormData();
parseForm.append("file", new Blob([bytes], { type: "application/pdf" }), "document.pdf");
const parseRes = await fetch(`${BASE_URL}/parse`, { method: "POST", headers: { "api-key": API_KEY }, body: parseForm });
if (!parseRes.ok) throw new Error(`parse submit failed: HTTP ${parseRes.status} ${await parseRes.text()}`);
const parseJob = await parseRes.json();
const parsed = await waitFor(`${BASE_URL}/parse/${parseJob.job_id}`);
const normalize = (text) => String(text).toLowerCase().replace(/[^a-z0-9]/g, "");
const words = [];
for (const chunk of parsed.chunks) {
for (const seg of chunk.segments) {
const { left: sx, top: sy } = seg.bbox;
for (const word of seg.ocr || []) {
const norm = normalize(word.text);
if (!norm) continue;
const b = word.bbox;
words.push({
norm,
page: seg.page_number,
pageWidth: seg.page_width,
x: sx + b.left,
width: b.width,
// y + height lands on the baseline reliably; the height alone does not.
baseline: sy + b.top + b.height,
});
}
}
}
// With no words there is nothing to match, and the script would "succeed"
// by writing an unredacted copy. Fail loudly instead.
if (words.length === 0) throw new Error("Parse returned no OCR words; refusing to write an unredacted PDF");
// Step 3: find every occurrence by normalized concatenation, so a value matches
// however OCR split it (hyphens, "@") or wrapped across lines.
function findRuns(value) {
const target = normalize(value);
const runs = [];
if (target.length < 3) return runs;
for (let i = 0; i < words.length; i++) {
let acc = "";
for (let j = i; j < words.length && words[j].page === words[i].page && j - i < 10; j++) {
acc += words[j].norm;
if (acc === target) { runs.push(words.slice(i, j + 1)); break; }
if (!target.startsWith(acc)) break;
}
}
return runs;
}
// Step 4: turn each run into one box per line, anchored to the baseline.
const BAND_ABOVE = 19, BAND_BELOW = 6, PAD = 2;
function lineBoxes(run) {
const lines = [];
let current = [run[0]];
for (const word of run.slice(1)) {
const sameLine = Math.abs(word.baseline - current[current.length - 1].baseline) <= 12;
if (sameLine && word.x >= current[current.length - 1].x - 2) current.push(word);
else { lines.push(current); current = [word]; }
}
lines.push(current);
return lines.map((line) => {
const baseline = Math.max(...line.map((w) => w.baseline));
return [
Math.min(...line.map((w) => w.x)) - PAD,
baseline - BAND_ABOVE,
Math.max(...line.map((w) => w.x + w.width)) + PAD,
baseline + BAND_BELOW,
];
});
}
const boxesByPage = new Map();
for (const value of sensitive) {
for (const run of findRuns(value)) {
const page = run[0].page;
if (!boxesByPage.has(page)) boxesByPage.set(page, []);
boxesByPage.get(page).push(...lineBoxes(run));
}
}
// Step 5: render each page to an image, paint the boxes, and rebuild the PDF.
// Flattening to an image removes the text for good; a drawn box alone would not.
const pageWidths = new Map(words.map((w) => [w.page, w.pageWidth]));
const fonts = "./node_modules/pdfjs-dist/standard_fonts/";
const pdf = await pdfjsLib.getDocument({ data: new Uint8Array(bytes), standardFontDataUrl: fonts }).promise;
const out = await PDFDocument.create();
for (let pageNo = 1; pageNo <= pdf.numPages; pageNo++) {
const page = await pdf.getPage(pageNo);
const base = page.getViewport({ scale: 1.0 }); // page size in PDF points
// Render at whatever scale matches parse's pixel grid for this page.
const scale = (pageWidths.get(pageNo) ?? base.width * 2) / base.width;
const viewport = page.getViewport({ scale });
const canvas = createCanvas(viewport.width, viewport.height);
const ctx = canvas.getContext("2d");
await page.render({ canvasContext: ctx, viewport, canvas }).promise;
ctx.fillStyle = "#000";
for (const [x0, y0, x1, y1] of boxesByPage.get(pageNo) || []) {
ctx.fillRect(x0, y0, x1 - x0, y1 - y0);
}
const png = await out.embedPng(canvas.toBuffer("image/png"));
const outPage = out.addPage([base.width, base.height]);
outPage.drawImage(png, { x: 0, y: 0, width: base.width, height: base.height });
}
writeFileSync("redacted.pdf", await out.save());
console.log("Wrote redacted.pdf");
```
## Step 1: Set Up Your Environment
Before writing any code, we need three things: an API key, a document to redact, and a few libraries for the chosen language.
### 1.1 Get an Unsiloed AI API Key
To get API access, [sign up on Unsiloed AI](https://cal.com/aman-mishra-p0ry57/15min). Export your key as an environment variable so it stays out of source control:
```bash theme={null}
export UNSILOED_API_KEY="your-api-key"
```
### 1.2 Pick a Document
The walkthrough assumes a PDF saved as `document.pdf` in your working directory. To follow along with the exact output shown below, download our [sample account statement](https://raw.githubusercontent.com/Unsiloed-AI/cookbook/9c80a90e0315a33c9b8a68d8b3355199771b598f/sample-documents/sample-statement.pdf): a one-page letter that repeats the account holder's name, account number, and email so we can prove every occurrence gets removed, not just the first.
### 1.3 Install Dependencies
The API calls are plain HTTP, but writing the redacted file means rendering the PDF, so each language needs a couple of libraries.
You need Python 3.8 or newer. Install `requests` for the API calls and `PyMuPDF` for the redaction:
```bash theme={null}
pip install requests PyMuPDF
```
You need Node.js 18 or newer for the global `fetch`, `FormData`, and `Blob`. Install the PDF rendering and writing libraries:
```bash theme={null}
npm install pdfjs-dist @napi-rs/canvas pdf-lib
```
`pdfjs-dist` renders each page, `@napi-rs/canvas` gives it a canvas to draw on (with prebuilt binaries, so there are no system dependencies), and `pdf-lib` assembles the redacted pages back into a PDF.
## Step 2: Find the Sensitive Values With Extraction
We start with extraction rather than a list of search terms because we don't want to hard-code the data we're protecting. A schema describes the kinds of information that are sensitive, and the extractor reads the actual values off the document. Point it at a different statement and it finds that person's details instead.
### 2.1 Describe the Sensitive Fields
Each field is a name and a description. The description tells the extractor what to look for, so the clearer it is, the more reliable the result. We're after four pieces of personal data.
Create a file called `redact.py` with the configuration and schema:
```python redact.py theme={null}
import json
import os
import re
import time
import fitz # PyMuPDF
import requests
API_KEY = os.environ["UNSILOED_API_KEY"]
BASE_URL = "https://prod.visionapi.unsiloed.ai"
SCHEMA = {
"type": "object",
"properties": {
"full_name": {"type": "string", "description": "The full name of the account holder"},
"account_number": {"type": "string", "description": "The bank account number"},
"email": {"type": "string", "description": "The account holder's email address"},
"phone": {"type": "string", "description": "The account holder's phone number"},
},
"required": ["full_name", "account_number", "email", "phone"],
"additionalProperties": False,
}
```
Create a file called `redact.mjs` with the configuration and schema:
```javascript redact.mjs theme={null}
import { readFileSync, writeFileSync } from "node:fs";
import { createCanvas } from "@napi-rs/canvas";
import { PDFDocument } from "pdf-lib";
import * as pdfjsLib from "pdfjs-dist/legacy/build/pdf.mjs";
const API_KEY = process.env.UNSILOED_API_KEY;
const BASE_URL = "https://prod.visionapi.unsiloed.ai";
const SCHEMA = {
type: "object",
properties: {
full_name: { type: "string", description: "The full name of the account holder" },
account_number: { type: "string", description: "The bank account number" },
email: { type: "string", description: "The account holder's email address" },
phone: { type: "string", description: "The account holder's phone number" },
},
required: ["full_name", "account_number", "email", "phone"],
additionalProperties: false,
};
```
### 2.2 Submit the Document and Read the Values
Both endpoints in this recipe run asynchronously: we submit a job, get a `job_id`, and poll until it's done. Since we do that twice, it's worth a small helper. The extractor accepts `completed` as its done state; the parser uses `Succeeded`, so the helper checks for both. The failure side is just as inconsistent: a job can end up `failed`, `Failed`, or `Cancelled`, and the parser reports its reason in a `message` field where the extractor uses `error`. The helper covers all of those too, because otherwise a cancelled job would poll until the timeout and surface as a misleading "did not finish in time".
Add the polling helper and the extract call:
```python redact.py theme={null}
def wait_for(url):
for _ in range(90): # roughly 6 minutes at 4 seconds per poll
job = requests.get(url, headers={"api-key": API_KEY}).json()
status = job.get("status")
if status in ("completed", "review", "Succeeded"):
return job
if status in ("failed", "Failed", "Cancelled"):
# The extractor reports its reason in "error"; the parser uses "message".
raise RuntimeError(job.get("error") or job.get("message") or f"job {status}")
time.sleep(4)
raise TimeoutError(f"Job at {url} did not finish in time")
with open("document.pdf", "rb") as f:
resp = requests.post(
f"{BASE_URL}/v2/extract",
headers={"api-key": API_KEY},
files={"pdf_file": ("document.pdf", f, "application/pdf")},
data={"schema_data": json.dumps(SCHEMA), "model": "gamma"},
)
resp.raise_for_status()
job = resp.json()
result = wait_for(f"{BASE_URL}/extract/{job['job_id']}")["result"]
sensitive = []
for field in SCHEMA["properties"]:
value = result[field]["value"]
if value is None:
print(f"Field '{field}' not found in the document, skipping.")
continue
sensitive.append(value)
print("Will redact:", sensitive)
```
Add the polling helper and the extract call:
```javascript redact.mjs theme={null}
async function waitFor(url) {
for (let i = 0; i < 90; i++) { // roughly 6 minutes at 4 seconds per poll
const job = await (await fetch(url, { headers: { "api-key": API_KEY } })).json();
if (["completed", "review", "Succeeded"].includes(job.status)) return job;
// The extractor reports its reason in "error"; the parser uses "message".
if (["failed", "Failed", "Cancelled"].includes(job.status)) throw new Error(job.error || job.message || `job ${job.status}`);
await new Promise((r) => setTimeout(r, 4000));
}
throw new Error(`Job at ${url} did not finish in time`);
}
const bytes = readFileSync("document.pdf");
const extractForm = new FormData();
extractForm.append("pdf_file", new Blob([bytes], { type: "application/pdf" }), "document.pdf");
extractForm.append("schema_data", JSON.stringify(SCHEMA));
extractForm.append("model", "gamma");
const extractRes = await fetch(`${BASE_URL}/v2/extract`, { method: "POST", headers: { "api-key": API_KEY }, body: extractForm });
if (!extractRes.ok) throw new Error(`extract submit failed: HTTP ${extractRes.status} ${await extractRes.text()}`);
const extractJob = await extractRes.json();
const result = (await waitFor(`${BASE_URL}/extract/${extractJob.job_id}`)).result;
const sensitive = [];
for (const field of Object.keys(SCHEMA.properties)) {
const value = result[field].value;
if (value == null) {
console.log(`Field '${field}' not found in the document, skipping.`);
continue;
}
sensitive.push(value);
}
console.log("Will redact:", sensitive);
```
Each entry in the extractor's `result` is an object with a `value`, so we pull out the values into a flat `sensitive` list. For the sample document, that prints:
```text theme={null}
Will redact: ['Jonathan Hale', '4471-8829-0007', 'jhale@example.com', '(415) 555-0142']
```
These are the strings we now need to find and remove wherever they appear.
Leave the extractor's `detect_pii` parameter at its default of `false` for this recipe. That flag exists to block extraction from documents that contain PII: with it enabled, the endpoint returns HTTP 200 with `job_id: null` and `status: "pii_blocked"` instead of starting a job, and the polling step fails. Extracting PII is the point of this recipe, so the gate has to stay off.
## Step 3: Locate Every Occurrence With Parsing
Extraction told us what is sensitive. It does not tell us where each value sits on the page, or how many times it appears. For that we parse the document. Parsing returns the page broken into segments, and within each segment an `ocr` array giving every word along with its bounding box. That word-level detail is what lets us draw a box around each occurrence.
### 3.1 Parse the Document
Send the same file to `/parse`. Note the form field is `file` here, not `pdf_file`: the parser and the extractor use different field names. We leave `ocr_strategy` at its default `auto_detection`: it returns word-level `ocr` boxes for digital and scanned PDFs alike, and in our testing `force_ocr` degrades the geometry to line-level boxes, which makes redaction miss values.
Add the parse call:
```python redact.py theme={null}
with open("document.pdf", "rb") as f:
resp = requests.post(
f"{BASE_URL}/parse",
headers={"api-key": API_KEY},
files={"file": ("document.pdf", f, "application/pdf")},
)
resp.raise_for_status()
job = resp.json()
parsed = wait_for(f"{BASE_URL}/parse/{job['job_id']}")
```
Add the parse call:
```javascript redact.mjs theme={null}
const parseForm = new FormData();
parseForm.append("file", new Blob([bytes], { type: "application/pdf" }), "document.pdf");
const parseRes = await fetch(`${BASE_URL}/parse`, { method: "POST", headers: { "api-key": API_KEY }, body: parseForm });
if (!parseRes.ok) throw new Error(`parse submit failed: HTTP ${parseRes.status} ${await parseRes.text()}`);
const parseJob = await parseRes.json();
const parsed = await waitFor(`${BASE_URL}/parse/${parseJob.job_id}`);
```
### 3.2 Build a Flat List of Words
A word's `ocr` box is given relative to its segment, so its position on the page is the segment's top-left corner plus the word's own offset. We flatten every word into one list, recording its normalized text, its page, and a box.
Two details matter for drawing clean boxes later:
* **The reported word height is unreliable.** Some boxes come back only a couple of pixels tall. The bottom edge, `top + height`, sits on the text baseline consistently, so we store that baseline and reconstruct a full-height box from it in Step 4.
* **Parse boxes are in render pixels,** measured against the rendered page that `page_width` and `page_height` describe. We keep each word's `page_width` so we can compute the exact pixels-to-points ratio when we redact.
Add the normalizer and the word-collection loop:
```python redact.py theme={null}
def normalize(text):
return re.sub(r"[^a-z0-9]", "", str(text).lower())
words = []
for chunk in parsed["chunks"]:
for seg in chunk["segments"]:
sx, sy = seg["bbox"]["left"], seg["bbox"]["top"]
for word in seg.get("ocr") or []:
norm = normalize(word["text"])
if not norm:
continue
b = word["bbox"]
words.append({
"norm": norm,
"page": seg["page_number"],
"page_width": seg["page_width"],
"x": sx + b["left"],
"width": b["width"],
"baseline": sy + b["top"] + b["height"],
})
# With no words there is nothing to match, and the script would "succeed"
# by writing an unredacted copy. Fail loudly instead.
if not words:
raise RuntimeError("Parse returned no OCR words; refusing to write an unredacted PDF")
```
Add the normalizer and the word-collection loop:
```javascript redact.mjs theme={null}
const normalize = (text) => String(text).toLowerCase().replace(/[^a-z0-9]/g, "");
const words = [];
for (const chunk of parsed.chunks) {
for (const seg of chunk.segments) {
const { left: sx, top: sy } = seg.bbox;
for (const word of seg.ocr || []) {
const norm = normalize(word.text);
if (!norm) continue;
const b = word.bbox;
words.push({
norm,
page: seg.page_number,
pageWidth: seg.page_width,
x: sx + b.left,
width: b.width,
baseline: sy + b.top + b.height,
});
}
}
}
// With no words there is nothing to match, and the script would "succeed"
// by writing an unredacted copy. Fail loudly instead.
if (words.length === 0) throw new Error("Parse returned no OCR words; refusing to write an unredacted PDF");
```
We normalize text down to lowercase letters and digits, dropping spaces and punctuation. That's what makes the matching in the next step robust.
The guard at the end matters: every later step quietly does nothing when `words` is empty, so without it the script would still write a `redacted.pdf` with nothing redacted. A redaction script should fail loudly rather than produce a clean-looking file that leaks everything.
## Step 4: Match Values to Their Locations
Now we find each sensitive value in the word list. A naive equality check breaks down quickly: `4471-8829-0007` might arrive as one OCR token or several, an email keeps its `@`, and a phone number can wrap across a line break. So instead of matching word by word, we match by **normalized concatenation**. Starting at each word, we glue the normalized words together one at a time until the running string equals the normalized target, and stop early the moment it can no longer lead to a match.
### 4.1 Find Every Run of Words That Spells a Value
Add the matcher:
```python redact.py theme={null}
def find_runs(value):
target = normalize(value)
runs = []
if len(target) < 3:
return runs
for i in range(len(words)):
acc = ""
for j in range(i, min(i + 10, len(words))):
if words[j]["page"] != words[i]["page"]:
break
acc += words[j]["norm"]
if acc == target:
runs.append(words[i:j + 1])
break
if not target.startswith(acc):
break
return runs
```
Add the matcher:
```javascript redact.mjs theme={null}
function findRuns(value) {
const target = normalize(value);
const runs = [];
if (target.length < 3) return runs;
for (let i = 0; i < words.length; i++) {
let acc = "";
for (let j = i; j < words.length && words[j].page === words[i].page && j - i < 10; j++) {
acc += words[j].norm;
if (acc === target) { runs.push(words.slice(i, j + 1)); break; }
if (!target.startsWith(acc)) break;
}
}
return runs;
}
```
Each value can return several runs, one per occurrence. We skip targets shorter than three characters so a stray initial can't trigger a flood of matches.
## Step 5: Black Out Every Match
A run is a list of words. To redact it we need rectangles, and we need them to behave at two awkward edges: a value can wrap across a line, and individual word heights are unreliable. We handle both by grouping each run into lines and anchoring every box to the baseline we stored earlier.
### 5.1 Turn a Run Into One Box per Line
We walk the run's words, starting a new line whenever the baseline jumps or the text steps back to the left margin. For each line we take the horizontal extent from the words and a fixed band above the baseline (plus a little below for descenders).
Add the box builder:
```python redact.py theme={null}
BAND_ABOVE, BAND_BELOW, PAD = 19, 6, 2
def line_boxes(run):
lines, current = [], [run[0]]
for word in run[1:]:
same_line = abs(word["baseline"] - current[-1]["baseline"]) <= 12
if same_line and word["x"] >= current[-1]["x"] - 2:
current.append(word)
else:
lines.append(current)
current = [word]
lines.append(current)
boxes = []
for line in lines:
baseline = max(w["baseline"] for w in line)
boxes.append((
min(w["x"] for w in line) - PAD,
baseline - BAND_ABOVE,
max(w["x"] + w["width"] for w in line) + PAD,
baseline + BAND_BELOW,
))
return boxes
```
Add the box builder:
```javascript redact.mjs theme={null}
const BAND_ABOVE = 19, BAND_BELOW = 6, PAD = 2;
function lineBoxes(run) {
const lines = [];
let current = [run[0]];
for (const word of run.slice(1)) {
const sameLine = Math.abs(word.baseline - current[current.length - 1].baseline) <= 12;
if (sameLine && word.x >= current[current.length - 1].x - 2) current.push(word);
else { lines.push(current); current = [word]; }
}
lines.push(current);
return lines.map((line) => {
const baseline = Math.max(...line.map((w) => w.baseline));
return [
Math.min(...line.map((w) => w.x)) - PAD,
baseline - BAND_ABOVE,
Math.max(...line.map((w) => w.x + w.width)) + PAD,
baseline + BAND_BELOW,
];
});
}
```
### 5.2 Remove the Text
This is where the two languages diverge, because covering text and removing it are not the same thing. A black rectangle painted over a PDF leaves the original characters in the file, selectable and searchable underneath. To redact for real you either delete the content or destroy it.
* **Python** uses PyMuPDF's redaction annotations. `apply_redactions` deletes the underlying text and image data inside each box, then fills it black. The output stays a normal PDF, and everything outside the boxes is left untouched and still selectable.
* **JavaScript** renders each page to an image, paints the boxes onto the pixels, and rebuilds the PDF from those images. Once a page is an image, there is no text layer left to recover.
Add the redaction loop and save. We scale each box from parse's render pixels back to PDF points using the ratio of `page_width` to the page's width in points. Deletion and drawing use different rectangles on purpose: `apply_redactions` removes any text that touches the redaction rect, even by a fraction of a point, so we shrink the deletion rect two points clear of the neighboring lines (otherwise a box can silently delete words from the line above or below), then draw the full-height black box as plain ink:
```python redact.py theme={null}
doc = fitz.open("document.pdf")
boxes_by_page = {}
for value in sensitive:
for run in find_runs(value):
page_index = run[0]["page"] - 1
scale = run[0]["page_width"] / doc[page_index].rect.width
for x0, y0, x1, y1 in line_boxes(run):
rect = fitz.Rect(x0 / scale, y0 / scale, x1 / scale, y1 / scale)
boxes_by_page.setdefault(page_index, []).append(rect)
for page_index, rects in boxes_by_page.items():
page = doc[page_index]
for rect in rects:
page.add_redact_annot(rect + (0, 2, 0, -2), fill=(0, 0, 0))
page.apply_redactions()
for rect in rects:
page.draw_rect(rect, fill=(0, 0, 0), color=None)
doc.save("redacted.pdf")
print("Wrote redacted.pdf")
```
Run it:
```bash theme={null}
python redact.py
```
Group the boxes by page, then render, paint, and reassemble. We derive each page's render scale from the `page_width` that parse reported, so the canvas matches parse's pixel grid exactly. When placing each image we use the page's size in points, so the rebuilt page keeps its original dimensions:
```javascript redact.mjs theme={null}
const boxesByPage = new Map();
for (const value of sensitive) {
for (const run of findRuns(value)) {
const page = run[0].page;
if (!boxesByPage.has(page)) boxesByPage.set(page, []);
boxesByPage.get(page).push(...lineBoxes(run));
}
}
const pageWidths = new Map(words.map((w) => [w.page, w.pageWidth]));
const fonts = "./node_modules/pdfjs-dist/standard_fonts/";
const pdf = await pdfjsLib.getDocument({ data: new Uint8Array(bytes), standardFontDataUrl: fonts }).promise;
const out = await PDFDocument.create();
for (let pageNo = 1; pageNo <= pdf.numPages; pageNo++) {
const page = await pdf.getPage(pageNo);
const base = page.getViewport({ scale: 1.0 });
const scale = (pageWidths.get(pageNo) ?? base.width * 2) / base.width;
const viewport = page.getViewport({ scale });
const canvas = createCanvas(viewport.width, viewport.height);
const ctx = canvas.getContext("2d");
await page.render({ canvasContext: ctx, viewport, canvas }).promise;
ctx.fillStyle = "#000";
for (const [x0, y0, x1, y1] of boxesByPage.get(pageNo) || []) {
ctx.fillRect(x0, y0, x1 - x0, y1 - y0);
}
const png = await out.embedPng(canvas.toBuffer("image/png"));
const outPage = out.addPage([base.width, base.height]);
outPage.drawImage(png, { x: 0, y: 0, width: base.width, height: base.height });
}
writeFileSync("redacted.pdf", await out.save());
console.log("Wrote redacted.pdf");
```
Run it:
```bash theme={null}
node redact.mjs
```
## Step 6: Confirm the Text Is Really Gone
The point of redaction is that the data can't be recovered, so it's worth checking rather than trusting the visual. Open `redacted.pdf` and try to select the blacked-out text: nothing should be there. You can confirm the same thing in code by pulling the text layer back out of the result.
With Python and PyMuPDF:
```python theme={null}
import fitz
text = " ".join(page.get_text() for page in fitz.open("redacted.pdf"))
for value in ["Jonathan Hale", "4471-8829-0007", "jhale@example.com"]:
print(value, "->", "REMOVED" if value not in text else "STILL PRESENT")
```
For the sample document this prints `REMOVED` for every value, while the parts we didn't target, like the closing balance and the dates, are still present and selectable:
```text theme={null}
Jonathan Hale -> REMOVED
4471-8829-0007 -> REMOVED
jhale@example.com -> REMOVED
```
The JavaScript output is image-only, so it has no text layer at all; selecting anywhere on the page returns nothing.
## Where to Take This Next
The schema is the lever here. Add fields for any other data you need gone, such as a date of birth, a postal address, or a national ID number, and the same pipeline finds and removes it.
A few directions to take it further:
* **Redact signatures.** Parsing labels handwritten signatures with a `Signature` segment type, but only when you submit the parse job with `layout_analysis=advanced_layout_detection`; the default layout analysis we use in this recipe never returns it. With that parameter set, you can black out each signature segment's box directly, without an extraction step, since the segment already carries its location.
* **Handle scanned documents.** Because the locations come from the parser's OCR rather than the PDF text layer, the same script works on scans and photos, not just digital PDFs.
* **Review before you ship.** For a human-in-the-loop workflow, draw the boxes in a bright color first and have a reviewer confirm them, then switch to black once the set is approved.
How schemas drive extraction, including nested objects and arrays.
The full parse response, segment types, and word-level OCR boxes.
Every segment type the parser can return, including `Signature`.
Browse the full request and response specs for `/v2/extract` and `/parse`.
# Sort and Extract a Mixed Document Pile
Source: https://docs.unsiloed.ai/docs/cookbooks/sort-and-extract
Split a PDF that holds several different documents into one file per type, then run a type-specific extraction schema over each.
A scanned batch or a shared inbox rarely holds one clean document. It holds an invoice, a receipt, and a purchase order merged into a single PDF, in no particular order. You can't run one extraction schema over the whole thing, because the fields you want from an invoice aren't the fields you want from a receipt.
This recipe handles that in two stages. First we split the pile into one file per document type, which also tells us what each file is. Then we run a different extraction schema over each file, chosen by its type. The result is a tidy object keyed by document type, with the right fields pulled from each.
This recipe chains two endpoints covered on their own in the [Splitting quickstart](/docs/document-processing/splitting/quickstart) and the [Extract quickstart](/docs/document-processing/extraction/quickstart). Read those first if you want the detail on either call in isolation.
## What We'll Build
A script that:
1. Sends a multi-document PDF to `/splitter` along with the document types we expect to find.
2. Gets back one PDF per type, each labelled with its category and hosted at a URL.
3. Picks a matching extraction schema for each file and sends it to `/v2/extract` by URL.
4. Collects the extracted fields into a single object, keyed by document type.
We'll run it against a three-page sample that holds an invoice, a receipt, and a purchase order. Grab the full script from the dropdown below if you'd rather skip the walkthrough.
Set `UNSILOED_API_KEY` in your environment and save the document pile as `pile.pdf` in the same directory before running.
```python sort_and_extract.py theme={null}
import json
import os
import time
import requests
API_KEY = os.environ["UNSILOED_API_KEY"]
BASE_URL = "https://prod.visionapi.unsiloed.ai"
# The document types we expect in the pile. The description guides the split.
CATEGORIES = [
{"name": "Invoice", "description": "A bill issued by a seller requesting payment for goods or services"},
{"name": "Receipt", "description": "Proof of a completed purchase from a store or point of sale"},
{"name": "Purchase Order", "description": "A buyer-issued document authorizing a purchase from a vendor"},
]
# One extraction schema per type. The keys match the category names above.
SCHEMAS = {
"Invoice": {
"type": "object",
"properties": {
"vendor_name": {"type": "string", "description": "Company issuing the invoice (the seller)"},
"invoice_number": {"type": "string", "description": "Invoice identifier"},
"total_due": {"type": "number", "description": "Total amount due in US dollars"},
},
"required": ["vendor_name", "invoice_number", "total_due"],
"additionalProperties": False,
},
"Receipt": {
"type": "object",
"properties": {
"store_name": {"type": "string", "description": "Name of the store"},
"transaction_id": {"type": "string", "description": "Transaction or receipt ID"},
"total": {"type": "number", "description": "Total amount paid in US dollars"},
},
"required": ["store_name", "transaction_id", "total"],
"additionalProperties": False,
},
"Purchase Order": {
"type": "object",
"properties": {
"po_number": {"type": "string", "description": "Purchase order number"},
"vendor_name": {"type": "string", "description": "Vendor the order is placed with"},
"required_by": {"type": "string", "description": "Date the order is required by"},
},
"required": ["po_number", "vendor_name", "required_by"],
"additionalProperties": False,
},
}
def wait_for(url):
"""Poll a job URL until it finishes, then return the completed response."""
for _ in range(90): # roughly 6 minutes at 4 seconds per poll
job = requests.get(url, headers={"api-key": API_KEY}).json()
if job.get("status") in ("completed", "review"):
return job
if job.get("status") == "failed":
raise RuntimeError(job.get("error", "job failed"))
time.sleep(4)
raise TimeoutError(f"Job at {url} did not finish in time")
# Stage 1: split the pile into one file per document type.
with open("pile.pdf", "rb") as f:
resp = requests.post(
f"{BASE_URL}/splitter",
headers={"api-key": API_KEY},
files={"file": ("pile.pdf", f, "application/pdf")},
data={"categories": json.dumps(CATEGORIES)},
)
resp.raise_for_status()
job = resp.json()
sections = wait_for(f"{BASE_URL}/splitter/{job['job_id']}")["result"]["files"]
print(f"Split into {len(sections)} documents: {[s['name'] for s in sections]}")
# Stage 2: extract type-specific fields from each section, by URL.
extract_jobs = {}
for section in sections:
category = os.path.splitext(section["name"])[0]
schema = SCHEMAS.get(category)
if schema is None:
print(f"No schema for '{category}', skipping.")
continue
resp = requests.post(
f"{BASE_URL}/v2/extract",
headers={"api-key": API_KEY},
data={
"file_url": section["full_path"],
"schema_data": json.dumps(schema),
"model": "gamma",
},
)
resp.raise_for_status()
extract_jobs[category] = resp.json()["job_id"]
results = {}
for category, job_id in extract_jobs.items():
done = wait_for(f"{BASE_URL}/extract/{job_id}")
results[category] = {field: cell["value"] for field, cell in done["result"].items()}
print(json.dumps(results, indent=2))
```
Save this as `sort_and_extract.mjs` or set `"type": "module"` in your `package.json`. Requires Node.js 18 or newer for the global `fetch`, `FormData`, and `Blob`.
```javascript sort_and_extract.mjs theme={null}
import fs from "node:fs";
const API_KEY = process.env.UNSILOED_API_KEY;
const BASE_URL = "https://prod.visionapi.unsiloed.ai";
// The document types we expect in the pile. The description guides the split.
const CATEGORIES = [
{ name: "Invoice", description: "A bill issued by a seller requesting payment for goods or services" },
{ name: "Receipt", description: "Proof of a completed purchase from a store or point of sale" },
{ name: "Purchase Order", description: "A buyer-issued document authorizing a purchase from a vendor" },
];
// One extraction schema per type. The keys match the category names above.
const SCHEMAS = {
Invoice: {
type: "object",
properties: {
vendor_name: { type: "string", description: "Company issuing the invoice (the seller)" },
invoice_number: { type: "string", description: "Invoice identifier" },
total_due: { type: "number", description: "Total amount due in US dollars" },
},
required: ["vendor_name", "invoice_number", "total_due"],
additionalProperties: false,
},
Receipt: {
type: "object",
properties: {
store_name: { type: "string", description: "Name of the store" },
transaction_id: { type: "string", description: "Transaction or receipt ID" },
total: { type: "number", description: "Total amount paid in US dollars" },
},
required: ["store_name", "transaction_id", "total"],
additionalProperties: false,
},
"Purchase Order": {
type: "object",
properties: {
po_number: { type: "string", description: "Purchase order number" },
vendor_name: { type: "string", description: "Vendor the order is placed with" },
required_by: { type: "string", description: "Date the order is required by" },
},
required: ["po_number", "vendor_name", "required_by"],
additionalProperties: false,
},
};
async function waitFor(url) {
// Poll a job URL until it finishes, then return the completed response.
for (let i = 0; i < 90; i++) { // roughly 6 minutes at 4 seconds per poll
const job = await (await fetch(url, { headers: { "api-key": API_KEY } })).json();
if (["completed", "review"].includes(job.status)) return job;
if (job.status === "failed") throw new Error(job.error || "job failed");
await new Promise((r) => setTimeout(r, 4000));
}
throw new Error(`Job at ${url} did not finish in time`);
}
// Stage 1: split the pile into one file per document type.
const splitForm = new FormData();
splitForm.append("file", new Blob([fs.readFileSync("pile.pdf")]), "pile.pdf");
splitForm.append("categories", JSON.stringify(CATEGORIES));
const splitRes = await fetch(`${BASE_URL}/splitter`, {
method: "POST",
headers: { "api-key": API_KEY },
body: splitForm,
});
if (!splitRes.ok) throw new Error(`split submit failed: HTTP ${splitRes.status} ${await splitRes.text()}`);
const splitJob = await splitRes.json();
const sections = (await waitFor(`${BASE_URL}/splitter/${splitJob.job_id}`)).result.files;
console.log(`Split into ${sections.length} documents: ${sections.map((s) => s.name).join(", ")}`);
// Stage 2: extract type-specific fields from each section, by URL.
const extractJobs = {};
for (const section of sections) {
const category = section.name.replace(/\.pdf$/, "");
const schema = SCHEMAS[category];
if (!schema) {
console.log(`No schema for '${category}', skipping.`);
continue;
}
const extractForm = new FormData();
extractForm.append("file_url", section.full_path);
extractForm.append("schema_data", JSON.stringify(schema));
extractForm.append("model", "gamma");
const extractRes = await fetch(`${BASE_URL}/v2/extract`, {
method: "POST",
headers: { "api-key": API_KEY },
body: extractForm,
});
if (!extractRes.ok) throw new Error(`extract submit failed: HTTP ${extractRes.status} ${await extractRes.text()}`);
extractJobs[category] = (await extractRes.json()).job_id;
}
const results = {};
for (const [category, jobId] of Object.entries(extractJobs)) {
const done = await waitFor(`${BASE_URL}/extract/${jobId}`);
results[category] = Object.fromEntries(
Object.entries(done.result).map(([field, cell]) => [field, cell.value]),
);
}
console.log(JSON.stringify(results, null, 2));
```
## Step 1: Set Up Your Environment
Before writing any code, we need three things: an API key, a document pile, and the runtime for our chosen language.
### 1.1 Get an Unsiloed AI API Key
To get API access, [sign up on Unsiloed AI](https://cal.com/aman-mishra-p0ry57/15min). Export your key as an environment variable so it stays out of source control:
```bash theme={null}
export UNSILOED_API_KEY="your-api-key"
```
### 1.2 Pick a Document Pile
The walkthrough assumes a PDF saved as `pile.pdf` in your working directory. To follow along with the exact output shown below, download our [sample document pile](https://raw.githubusercontent.com/Unsiloed-AI/cookbook/9c80a90e0315a33c9b8a68d8b3355199771b598f/sample-documents/sample-split.pdf): a three-page PDF that holds an invoice, a store receipt, and a purchase order, one per page. The three documents are unrelated and in mixed order, which is exactly the case this recipe is built for.
### 1.3 Install Dependencies
You need Python 3.8 or newer. Install the `requests` package:
```bash theme={null}
pip install requests
```
You need Node.js 18 or newer for the global `fetch`, `FormData`, and `Blob`. No external packages needed.
## Step 2: Split the Pile by Document Type
The splitter takes the whole PDF plus a list of the document types we expect, and returns one new PDF per type. Each returned file is labelled with the category it matched and hosted at a URL we can hand straight to the extractor in the next step, so there's no need to save anything to disk in between.
### 2.1 Describe the Document Types
A category is a `name` and a `description`. The description does the work: it tells the splitter how to recognize each type, so the clearer it is, the more reliable the split. We're looking for three types.
Create a file called `sort_and_extract.py` with the configuration and categories:
```python sort_and_extract.py theme={null}
import json
import os
import time
import requests
API_KEY = os.environ["UNSILOED_API_KEY"]
BASE_URL = "https://prod.visionapi.unsiloed.ai"
CATEGORIES = [
{"name": "Invoice", "description": "A bill issued by a seller requesting payment for goods or services"},
{"name": "Receipt", "description": "Proof of a completed purchase from a store or point of sale"},
{"name": "Purchase Order", "description": "A buyer-issued document authorizing a purchase from a vendor"},
]
```
Create a file called `sort_and_extract.mjs` with the configuration and categories:
```javascript sort_and_extract.mjs theme={null}
import fs from "node:fs";
const API_KEY = process.env.UNSILOED_API_KEY;
const BASE_URL = "https://prod.visionapi.unsiloed.ai";
const CATEGORIES = [
{ name: "Invoice", description: "A bill issued by a seller requesting payment for goods or services" },
{ name: "Receipt", description: "Proof of a completed purchase from a store or point of sale" },
{ name: "Purchase Order", description: "A buyer-issued document authorizing a purchase from a vendor" },
];
```
### 2.2 Submit the Pile and Wait for the Split
Both endpoints in this recipe run asynchronously: we submit a job, get a `job_id`, and poll until it's done. Since we do that twice, it's worth a small helper. Then we post the pile to `/splitter` with the categories as a JSON string under the `categories` field.
Add the polling helper and the split call:
```python sort_and_extract.py theme={null}
def wait_for(url):
for _ in range(90): # roughly 6 minutes at 4 seconds per poll
job = requests.get(url, headers={"api-key": API_KEY}).json()
if job.get("status") in ("completed", "review"):
return job
if job.get("status") == "failed":
raise RuntimeError(job.get("error", "job failed"))
time.sleep(4)
raise TimeoutError(f"Job at {url} did not finish in time")
with open("pile.pdf", "rb") as f:
resp = requests.post(
f"{BASE_URL}/splitter",
headers={"api-key": API_KEY},
files={"file": ("pile.pdf", f, "application/pdf")},
data={"categories": json.dumps(CATEGORIES)},
)
resp.raise_for_status()
job = resp.json()
sections = wait_for(f"{BASE_URL}/splitter/{job['job_id']}")["result"]["files"]
print(f"Split into {len(sections)} documents: {[s['name'] for s in sections]}")
```
Note the form field is `file` here, not `pdf_file`. The splitter and the extractor use different field names.
Add the polling helper and the split call:
```javascript sort_and_extract.mjs theme={null}
async function waitFor(url) {
for (let i = 0; i < 90; i++) { // roughly 6 minutes at 4 seconds per poll
const job = await (await fetch(url, { headers: { "api-key": API_KEY } })).json();
if (["completed", "review"].includes(job.status)) return job;
if (job.status === "failed") throw new Error(job.error || "job failed");
await new Promise((r) => setTimeout(r, 4000));
}
throw new Error(`Job at ${url} did not finish in time`);
}
const splitForm = new FormData();
splitForm.append("file", new Blob([fs.readFileSync("pile.pdf")]), "pile.pdf");
splitForm.append("categories", JSON.stringify(CATEGORIES));
const splitRes = await fetch(`${BASE_URL}/splitter`, {
method: "POST",
headers: { "api-key": API_KEY },
body: splitForm,
});
if (!splitRes.ok) throw new Error(`split submit failed: HTTP ${splitRes.status} ${await splitRes.text()}`);
const splitJob = await splitRes.json();
const sections = (await waitFor(`${BASE_URL}/splitter/${splitJob.job_id}`)).result.files;
console.log(`Split into ${sections.length} documents: ${sections.map((s) => s.name).join(", ")}`);
```
Note the form field is `file` here, not `pdf_file`. The splitter and the extractor use different field names.
Each entry in `sections` describes one split-out document. The two fields we use are `name` (the category it matched, with a `.pdf` suffix, such as `Invoice.pdf`) and `full_path` (a URL to the new single-document PDF). Those URLs are [signed S3 links that expire roughly an hour after the response](/docs/document-processing/splitting/response-format#file-object), so we use them right away in the next step rather than storing them. If a link does lapse, re-issue `GET /splitter/{job_id}` to get fresh ones.
## Step 3: Extract the Right Fields From Each Document
Now we loop over the split files. For each one, we look up the schema that matches its type and send the file's URL to `/v2/extract`. Passing `file_url` means the extractor pulls the document straight from the split result, so we never download and re-upload it.
### 3.1 Define a Schema per Type
Each document type wants different fields. We keep one schema per category in a dictionary, keyed by the same names we gave the splitter, so we can look up the right schema by the file's category.
Add the schema map below the categories:
```python sort_and_extract.py theme={null}
SCHEMAS = {
"Invoice": {
"type": "object",
"properties": {
"vendor_name": {"type": "string", "description": "Company issuing the invoice (the seller)"},
"invoice_number": {"type": "string", "description": "Invoice identifier"},
"total_due": {"type": "number", "description": "Total amount due in US dollars"},
},
"required": ["vendor_name", "invoice_number", "total_due"],
"additionalProperties": False,
},
"Receipt": {
"type": "object",
"properties": {
"store_name": {"type": "string", "description": "Name of the store"},
"transaction_id": {"type": "string", "description": "Transaction or receipt ID"},
"total": {"type": "number", "description": "Total amount paid in US dollars"},
},
"required": ["store_name", "transaction_id", "total"],
"additionalProperties": False,
},
"Purchase Order": {
"type": "object",
"properties": {
"po_number": {"type": "string", "description": "Purchase order number"},
"vendor_name": {"type": "string", "description": "Vendor the order is placed with"},
"required_by": {"type": "string", "description": "Date the order is required by"},
},
"required": ["po_number", "vendor_name", "required_by"],
"additionalProperties": False,
},
}
```
Add the schema map below the categories:
```javascript sort_and_extract.mjs theme={null}
const SCHEMAS = {
Invoice: {
type: "object",
properties: {
vendor_name: { type: "string", description: "Company issuing the invoice (the seller)" },
invoice_number: { type: "string", description: "Invoice identifier" },
total_due: { type: "number", description: "Total amount due in US dollars" },
},
required: ["vendor_name", "invoice_number", "total_due"],
additionalProperties: false,
},
Receipt: {
type: "object",
properties: {
store_name: { type: "string", description: "Name of the store" },
transaction_id: { type: "string", description: "Transaction or receipt ID" },
total: { type: "number", description: "Total amount paid in US dollars" },
},
required: ["store_name", "transaction_id", "total"],
additionalProperties: false,
},
"Purchase Order": {
type: "object",
properties: {
po_number: { type: "string", description: "Purchase order number" },
vendor_name: { type: "string", description: "Vendor the order is placed with" },
required_by: { type: "string", description: "Date the order is required by" },
},
required: ["po_number", "vendor_name", "required_by"],
additionalProperties: false,
},
};
```
### 3.2 Extract Each Section by URL
Loop over the split files, look up the schema for each category, and submit an extract job for each one. We strip the `.pdf` suffix from the file name to get the category key, and skip any type we don't have a schema for. Submitting every job before polling any of them matters here: the extractor fetches each `file_url` at submission time, so the signed links get consumed while they're fresh instead of waiting behind another job's polling loop, and the jobs run in parallel on the server rather than one at a time.
Add the extraction loop:
```python sort_and_extract.py theme={null}
extract_jobs = {}
for section in sections:
category = os.path.splitext(section["name"])[0]
schema = SCHEMAS.get(category)
if schema is None:
print(f"No schema for '{category}', skipping.")
continue
resp = requests.post(
f"{BASE_URL}/v2/extract",
headers={"api-key": API_KEY},
data={
"file_url": section["full_path"],
"schema_data": json.dumps(schema),
"model": "gamma",
},
)
resp.raise_for_status()
extract_jobs[category] = resp.json()["job_id"]
results = {}
for category, job_id in extract_jobs.items():
done = wait_for(f"{BASE_URL}/extract/{job_id}")
results[category] = {field: cell["value"] for field, cell in done["result"].items()}
```
Because we extract by URL, there's no file upload here, so everything goes in the `data` form fields instead of `files`.
Add the extraction loop:
```javascript sort_and_extract.mjs theme={null}
const extractJobs = {};
for (const section of sections) {
const category = section.name.replace(/\.pdf$/, "");
const schema = SCHEMAS[category];
if (!schema) {
console.log(`No schema for '${category}', skipping.`);
continue;
}
const extractForm = new FormData();
extractForm.append("file_url", section.full_path);
extractForm.append("schema_data", JSON.stringify(schema));
extractForm.append("model", "gamma");
const extractRes = await fetch(`${BASE_URL}/v2/extract`, {
method: "POST",
headers: { "api-key": API_KEY },
body: extractForm,
});
if (!extractRes.ok) throw new Error(`extract submit failed: HTTP ${extractRes.status} ${await extractRes.text()}`);
extractJobs[category] = (await extractRes.json()).job_id;
}
const results = {};
for (const [category, jobId] of Object.entries(extractJobs)) {
const done = await waitFor(`${BASE_URL}/extract/${jobId}`);
results[category] = Object.fromEntries(
Object.entries(done.result).map(([field, cell]) => [field, cell.value]),
);
}
```
## Step 4: Print the Combined Result
The `results` object now holds the extracted fields for every document in the pile, keyed by type. Print it as formatted JSON.
Add the final line and run the script:
```python sort_and_extract.py theme={null}
print(json.dumps(results, indent=2))
```
```bash theme={null}
python sort_and_extract.py
```
Add the final line and run the script:
```javascript sort_and_extract.mjs theme={null}
console.log(JSON.stringify(results, null, 2));
```
```bash theme={null}
node sort_and_extract.mjs
```
### Sample Output
Run against the [sample pile](https://raw.githubusercontent.com/Unsiloed-AI/cookbook/9c80a90e0315a33c9b8a68d8b3355199771b598f/sample-documents/sample-split.pdf), the script splits the three pages into three labelled documents, then extracts the fields each type calls for:
```text theme={null}
Split into 3 documents: ['Invoice.pdf', 'Receipt.pdf', 'Purchase Order.pdf']
```
```json theme={null}
{
"Invoice": {
"vendor_name": "Greenfield Print & Bindery",
"invoice_number": "GPB-20418",
"total_due": 2150.0
},
"Receipt": {
"store_name": "Cooper's Office Supply",
"transaction_id": "7741-208-4490",
"total": 87.61
},
"Purchase Order": {
"po_number": "PO-2026-0813",
"vendor_name": "Westbrook Furniture Co.",
"required_by": "June 5, 2026"
}
}
```
One PDF went in; three correctly typed records came out, each carrying only the fields that make sense for its document type. A receipt never gets asked for an invoice number, and the purchase order's vendor is pulled from the right place on the page.
## Where to Take This Next
The pile in this recipe had one page per document, but the splitter handles multi-page documents too, grouping consecutive pages that belong together. To process a folder of mixed PDFs, wrap the whole script in a loop over your files and merge each run's `results` into one collection.
How the splitter decides document boundaries, and the full response reference.
Label a single document by type without splitting it.
Build richer per-type schemas, including nested objects and line-item tables.
Browse the full request and response specs for `/splitter` and `/v2/extract`.
# Extract a Table from an Image or PDF in Microsoft Excel
Source: https://docs.unsiloed.ai/docs/cookbooks/spreadsheets/excel-extract-table
Build an Excel task-pane add-in that sends an image or PDF to Unsiloed and writes the first detected table to a worksheet.
This guide uses Excel for the web, so you can follow it on macOS without
installing the Microsoft Office desktop apps.
In this guide, we'll build an **Unsiloed Table Extractor** task pane for Excel.
You choose an image or PDF, enter an Unsiloed API key, and click **Extract
table**. The add-in writes the first detected table into a new worksheet.
The finished add-in turns an invoice table into editable Excel cells:
The project has four parts: a local HTTPS server, an HTML task pane, its
JavaScript behavior, and the XML manifest that connects the page to Excel.
## What You Need
Before you start, gather:
* Node.js 22.12 or higher
* A Microsoft account that can upload custom add-ins in [Excel for the web](https://excel.cloud.microsoft)
* An Unsiloed API key from the [Unsiloed dashboard](https://app.unsiloed.ai)
* The [sample invoice PDF](/docs/images/excel/sample-invoice.pdf)
The account requirement matters because an organization administrator can
disable custom add-in uploads. This guide uses a personal Microsoft account and
Chrome on macOS.
## What We'll Build
The integration follows four operations:
1. Office.js confirms that the page is running inside Excel.
2. The task pane uploads the selected document to Unsiloed and polls the parse job.
3. JavaScript converts the returned table HTML into rows and columns.
4. Office.js writes the resulting array to a new worksheet in one operation.
Build each part in the guide below, or copy the complete project first.
If you want to run the integration without following each explanation, run:
```bash theme={null}
mkdir unsiloed-excel-addin
cd unsiloed-excel-addin
npm init -y
npm install --save-dev vite office-addin-dev-certs
```
Then create the following four files. You can also give this section to a coding
agent and ask it to create the project.
**`vite.config.mjs`**
```javascript vite.config.mjs theme={null}
import { defineConfig } from "vite";
import devCerts from "office-addin-dev-certs";
export default defineConfig(async ({ command }) => {
if (command === "serve") {
const https = await devCerts.getHttpsServerOptions();
return {
server: {
host: "localhost",
port: 3000,
strictPort: true,
https,
},
};
}
return {};
});
```
**`index.html`**
```html index.html theme={null}
Unsiloed Table Extractor
Turn a document table into cells.
Choose an image or PDF. Unsiloed writes the first table to a new worksheet.
```
**`taskpane.js`**
```javascript taskpane.js theme={null}
const API_URL = "https://prod.visionapi.unsiloed.ai";
const form = document.querySelector("#form");
const keyInput = document.querySelector("#api-key");
const fileInput = document.querySelector("#file");
const button = document.querySelector("#extract");
const statusMessage = document.querySelector("#status");
let excelReady = false;
let isBusy = false;
Office.onReady((info) => {
excelReady = info.host === Office.HostType.Excel;
updateButton();
if (!excelReady) {
showStatus("Open this page as an Excel add-in.", "error");
}
});
keyInput.addEventListener("input", updateButton);
fileInput.addEventListener("change", updateButton);
form.addEventListener("submit", async (event) => {
event.preventDefault();
const file = fileInput.files[0];
const apiKey = keyInput.value.trim();
if (!file || !apiKey || !excelReady || isBusy) {
return;
}
isBusy = true;
updateButton();
showStatus("Sending the document to Unsiloed…");
try {
const job = await parseDocument(file, apiKey);
const rows = tableRows(job);
showStatus("Writing the table into Excel…");
const sheetName = await writeRows(rows);
showStatus(`Added ${rows.length} rows to ${sheetName}.`, "success");
} catch (error) {
const message =
error instanceof Error
? error.message
: "The table could not be extracted.";
showStatus(message, "error");
} finally {
isBusy = false;
updateButton();
}
});
async function parseDocument(file, apiKey) {
const formData = new FormData();
formData.append("file", file, file.name);
const response = await apiFetch(`${API_URL}/parse`, apiKey, {
method: "POST",
body: formData,
});
const submission = await response.json();
if (!submission.job_id) {
throw new Error("Unsiloed returned no job ID.");
}
return waitForParse(submission.job_id, apiKey);
}
async function waitForParse(jobId, apiKey) {
const timeout = AbortSignal.timeout(5 * 60 * 1000);
try {
while (true) {
await wait(2500, timeout);
const result = await apiFetch(
`${API_URL}/parse/${encodeURIComponent(jobId)}`,
apiKey,
{ signal: timeout },
);
const job = await result.json();
if (job.status === "Succeeded") {
return job;
}
if (job.status === "Failed" || job.status === "Cancelled") {
throw new Error(job.message || `Parsing ${job.status.toLowerCase()}.`);
}
}
} catch (error) {
if (timeout.aborted) {
throw new Error("Parsing took longer than five minutes.");
}
throw error;
}
}
async function apiFetch(url, apiKey, options = {}) {
const response = await fetch(url, {
...options,
headers: {
...options.headers,
"api-key": apiKey,
},
});
if (response.ok) {
return response;
}
const text = await response.text();
let message = text;
try {
const body = JSON.parse(text);
message = body.error?.message || body.message || body.detail || text;
} catch {
// Keep the plain-text response when the API does not return JSON.
}
throw new Error(message || `Unsiloed returned HTTP ${response.status}.`);
}
function wait(milliseconds, signal) {
return new Promise((resolve, reject) => {
if (signal.aborted) {
reject(signal.reason);
return;
}
const finish = () => {
signal.removeEventListener("abort", cancel);
resolve();
};
const cancel = () => {
clearTimeout(timer);
reject(signal.reason);
};
const timer = setTimeout(finish, milliseconds);
signal.addEventListener("abort", cancel, { once: true });
});
}
function tableRows(job) {
const segments = (job.chunks || []).flatMap(
(chunk) => chunk.segments || [],
);
const table = segments.find(
(segment) => segment.segment_type === "Table",
);
if (!table?.html) {
throw new Error("Unsiloed did not find a table.");
}
return htmlTableRows(table.html);
}
function htmlTableRows(html) {
const document = new DOMParser().parseFromString(html, "text/html");
const rowElements = [...document.querySelectorAll("tr")];
const grid = [];
const pending = new Map();
for (const [rowIndex, rowElement] of rowElements.entries()) {
const row = [];
let columnIndex = 0;
const consumePending = () => {
while (pending.has(`${rowIndex}:${columnIndex}`)) {
row[columnIndex] = pending.get(`${rowIndex}:${columnIndex}`);
columnIndex += 1;
}
};
consumePending();
for (const cell of rowElement.querySelectorAll(
":scope > th, :scope > td",
)) {
consumePending();
for (const lineBreak of cell.querySelectorAll("br")) {
lineBreak.replaceWith("\n");
}
const value = safeCellText(cell.textContent || "");
const colspan = positiveSpan(cell.getAttribute("colspan"));
const rowspan = positiveSpan(cell.getAttribute("rowspan"));
for (let x = 0; x < colspan; x += 1) {
row[columnIndex + x] = x === 0 ? value : "";
for (let y = 1; y < rowspan; y += 1) {
pending.set(`${rowIndex + y}:${columnIndex + x}`, "");
}
}
columnIndex += colspan;
}
consumePending();
grid.push(row);
}
const width = Math.max(0, ...grid.map((row) => row.length));
if (width === 0) {
throw new Error("The detected table contained no cells.");
}
return grid.map((row) =>
Array.from({ length: width }, (_, index) => row[index] ?? ""),
);
}
function positiveSpan(rawValue) {
const value = Number.parseInt(rawValue || "1", 10);
return Number.isFinite(value) && value > 0 ? value : 1;
}
function safeCellText(value) {
const text = value.replace(/\u00a0/g, " ").trim();
// A leading apostrophe prevents extracted text from becoming an Excel formula.
return /^[=+\-@]/.test(text) ? `'${text}` : text;
}
async function writeRows(rows) {
return Excel.run(async (context) => {
const sheets = context.workbook.worksheets;
sheets.load("items/name");
await context.sync();
const existingNames = new Set(
sheets.items.map((sheet) => sheet.name.toLowerCase()),
);
let sheetName = "Extracted";
let suffix = 2;
while (existingNames.has(sheetName.toLowerCase())) {
sheetName = `Extracted ${suffix}`;
suffix += 1;
}
const sheet = sheets.add(sheetName);
const range = sheet.getRangeByIndexes(
0,
0,
rows.length,
rows[0].length,
);
range.numberFormat = rows.map((row) => row.map(() => "@"));
range.values = rows;
range.format.autofitColumns();
range.format.autofitRows();
sheet.activate();
range.select();
await context.sync();
return sheetName;
});
}
function updateButton() {
const missingInput = !keyInput.value.trim() || !fileInput.files.length;
button.disabled = isBusy || !excelReady || missingInput;
}
function showStatus(message, type = "") {
statusMessage.textContent = message;
statusMessage.className = type;
}
```
**`manifest.xml`**
```xml manifest.xml theme={null}
e8d967ef-5e99-4e86-88fd-57efb72c6ae81.0.0.0Unsiloed AIen-USReadWriteDocument
```
Continue at [Step 5](#step-5-start-the-add-in) to run and load the completed
add-in.
## Step 1: Create the Project
Open Terminal and run:
```bash theme={null}
mkdir unsiloed-excel-addin
cd unsiloed-excel-addin
npm init -y
npm install --save-dev vite office-addin-dev-certs
```
Vite serves the task-pane page while we develop it. The certificate helper
creates a local certificate that Office trusts, because Office add-ins must use
HTTPS even during local development.
## Step 2: Serve the Add-in over HTTPS
Create `vite.config.mjs` in the `unsiloed-excel-addin` directory:
```javascript vite.config.mjs theme={null}
import { defineConfig } from "vite";
import devCerts from "office-addin-dev-certs";
export default defineConfig(async ({ command }) => {
if (command === "serve") {
const https = await devCerts.getHttpsServerOptions();
return {
server: {
host: "localhost",
port: 3000,
strictPort: true,
https,
},
};
}
return {};
});
```
The `command === "serve"` check limits certificate setup to the development
server. A production build doesn't need the local certificate. `strictPort`
keeps the URL predictable: if port 3000 is busy, Vite reports the conflict
instead of choosing a different URL that no longer matches the manifest.
## Step 3: Build the Task Pane
The task pane separates presentation from behavior. We'll first assemble
`index.html`, then build `taskpane.js` in the same order that data moves
through the integration.
### 3.1 Create the HTML Page
Create `index.html` beside `vite.config.mjs`:
```html index.html theme={null}
Unsiloed Table Extractor
Turn a document table into cells.
Choose an image or PDF. Unsiloed writes the first table to a new worksheet.
```
The first script loads Office.js from Microsoft's content delivery network. It
supplies the `Office` and `Excel` objects used later. The final script loads our
code as a JavaScript module after the form exists in the page.
The API key uses a password input, but the browser still holds its value in
memory. This recipe doesn't put the key in the workbook or browser storage.
### 3.2 Add Task-Pane Styles
Inside the `` element of `index.html`, add this `
```
An Office task pane is a narrow browser window. The fixed content width and
full-width controls keep the form usable when the reader resizes the pane. The
`error` and `success` classes let the JavaScript report progress without
using browser alerts.
### 3.3 Connect the Interface to Excel
Create `taskpane.js` beside `index.html` and add:
```javascript taskpane.js theme={null}
const API_URL = "https://prod.visionapi.unsiloed.ai";
const form = document.querySelector("#form");
const keyInput = document.querySelector("#api-key");
const fileInput = document.querySelector("#file");
const button = document.querySelector("#extract");
const statusMessage = document.querySelector("#status");
let excelReady = false;
let isBusy = false;
```
The constants keep stable references to the controls. The two booleans track
whether Office.js has connected to Excel and whether an extraction is already
running.
Below those declarations in `taskpane.js`, add the Office readiness and input
listeners:
```javascript taskpane.js theme={null}
Office.onReady((info) => {
excelReady = info.host === Office.HostType.Excel;
updateButton();
if (!excelReady) {
showStatus("Open this page as an Excel add-in.", "error");
}
});
keyInput.addEventListener("input", updateButton);
fileInput.addEventListener("change", updateButton);
```
A task pane is a web page, but the Excel APIs aren't ready when the browser first
parses it. `Office.onReady` runs after Office connects the page to its host. We
enable extraction only when that host is Excel and both form fields have values.
### 3.4 Coordinate the Extraction
Continue `taskpane.js` with the form's submit handler:
```javascript taskpane.js theme={null}
form.addEventListener("submit", async (event) => {
event.preventDefault();
const file = fileInput.files[0];
const apiKey = keyInput.value.trim();
if (!file || !apiKey || !excelReady || isBusy) {
return;
}
isBusy = true;
updateButton();
showStatus("Sending the document to Unsiloed…");
try {
const job = await parseDocument(file, apiKey);
const rows = tableRows(job);
showStatus("Writing the table into Excel…");
const sheetName = await writeRows(rows);
showStatus(`Added ${rows.length} rows to ${sheetName}.`, "success");
} catch (error) {
const message =
error instanceof Error
? error.message
: "The table could not be extracted.";
showStatus(message, "error");
} finally {
isBusy = false;
updateButton();
}
});
```
This function contains orchestration rather than API details. Reading the
`try` block from top to bottom shows the application flow: parse the file,
convert a table to rows, and write those rows to Excel. Each operation lives in
a focused helper function below.
The `isBusy` guard prevents a second submission while the first job is running.
The `finally` block clears it after either success or failure, so the reader can
retry without reloading the add-in.
### 3.5 Send and Poll the Parse Job
Add the function that uploads the file:
```javascript taskpane.js theme={null}
async function parseDocument(file, apiKey) {
const formData = new FormData();
formData.append("file", file, file.name);
const response = await apiFetch(`${API_URL}/parse`, apiKey, {
method: "POST",
body: formData,
});
const submission = await response.json();
if (!submission.job_id) {
throw new Error("Unsiloed returned no job ID.");
}
return waitForParse(submission.job_id, apiKey);
}
```
The `FormData` object creates the `multipart/form-data` request expected by
`POST /parse`. Unsiloed returns a job ID rather than the completed document,
because parsing can take longer than a single HTTP request should remain open.
The final line hands that ID to a separate polling function. Below
`parseDocument`, add the bounded polling loop:
```javascript taskpane.js theme={null}
async function waitForParse(jobId, apiKey) {
const timeout = AbortSignal.timeout(5 * 60 * 1000);
try {
while (true) {
await wait(2500, timeout);
const result = await apiFetch(
`${API_URL}/parse/${encodeURIComponent(jobId)}`,
apiKey,
{ signal: timeout },
);
const job = await result.json();
if (job.status === "Succeeded") {
return job;
}
if (job.status === "Failed" || job.status === "Cancelled") {
throw new Error(job.message || `Parsing ${job.status.toLowerCase()}.`);
}
}
} catch (error) {
if (timeout.aborted) {
throw new Error("Parsing took longer than five minutes.");
}
throw error;
}
}
```
The loop checks `GET /parse/{job_id}` every 2.5 seconds. It returns the job as
soon as Unsiloed reports `Succeeded`, surfaces terminal failures immediately,
and passes one abort signal to both the delays and network requests. That signal
ends the entire polling phase after five minutes, including a stalled request.
Continue `taskpane.js` with the authenticated request and abortable delay
helpers:
```javascript taskpane.js theme={null}
async function apiFetch(url, apiKey, options = {}) {
const response = await fetch(url, {
...options,
headers: {
...options.headers,
"api-key": apiKey,
},
});
if (response.ok) {
return response;
}
const text = await response.text();
let message = text;
try {
const body = JSON.parse(text);
message = body.error?.message || body.message || body.detail || text;
} catch {
// Keep the plain-text response when the API does not return JSON.
}
throw new Error(message || `Unsiloed returned HTTP ${response.status}.`);
}
function wait(milliseconds, signal) {
return new Promise((resolve, reject) => {
if (signal.aborted) {
reject(signal.reason);
return;
}
const finish = () => {
signal.removeEventListener("abort", cancel);
resolve();
};
const cancel = () => {
clearTimeout(timer);
reject(signal.reason);
};
const timer = setTimeout(finish, milliseconds);
signal.addEventListener("abort", cancel, { once: true });
});
}
```
Every Unsiloed request needs the `api-key` header. Centralizing that header in
`apiFetch` avoids repeating authentication code in the upload and polling
functions. The helper also turns unsuccessful responses into JavaScript errors,
which the submit handler already knows how to display. The delay helper listens
to the same abort signal as `fetch`, so the deadline covers both waiting and
network activity.
### 3.6 Convert Table HTML into a Grid
Add the table conversion functions:
```javascript taskpane.js theme={null}
function tableRows(job) {
const segments = (job.chunks || []).flatMap(
(chunk) => chunk.segments || [],
);
const table = segments.find(
(segment) => segment.segment_type === "Table",
);
if (!table?.html) {
throw new Error("Unsiloed did not find a table.");
}
return htmlTableRows(table.html);
}
function htmlTableRows(html) {
const document = new DOMParser().parseFromString(html, "text/html");
const rowElements = [...document.querySelectorAll("tr")];
const grid = [];
const pending = new Map();
for (const [rowIndex, rowElement] of rowElements.entries()) {
const row = [];
let columnIndex = 0;
const consumePending = () => {
while (pending.has(`${rowIndex}:${columnIndex}`)) {
row[columnIndex] = pending.get(`${rowIndex}:${columnIndex}`);
columnIndex += 1;
}
};
consumePending();
for (const cell of rowElement.querySelectorAll(
":scope > th, :scope > td",
)) {
consumePending();
for (const lineBreak of cell.querySelectorAll("br")) {
lineBreak.replaceWith("\n");
}
const value = safeCellText(cell.textContent || "");
const colspan = positiveSpan(cell.getAttribute("colspan"));
const rowspan = positiveSpan(cell.getAttribute("rowspan"));
for (let x = 0; x < colspan; x += 1) {
row[columnIndex + x] = x === 0 ? value : "";
for (let y = 1; y < rowspan; y += 1) {
pending.set(`${rowIndex + y}:${columnIndex + x}`, "");
}
}
columnIndex += colspan;
}
consumePending();
grid.push(row);
}
const width = Math.max(0, ...grid.map((row) => row.length));
if (width === 0) {
throw new Error("The detected table contained no cells.");
}
return grid.map((row) =>
Array.from({ length: width }, (_, index) => row[index] ?? ""),
);
}
function positiveSpan(rawValue) {
const value = Number.parseInt(rawValue || "1", 10);
return Number.isFinite(value) && value > 0 ? value : 1;
}
function safeCellText(value) {
const text = value.replace(/\u00a0/g, " ").trim();
// A leading apostrophe prevents extracted text from becoming an Excel formula.
return /^[=+\-@]/.test(text) ? `'${text}` : text;
}
```
The parse response organizes content as chunks containing segments. We flatten
those segments, select the first `Table`, and use the browser's `DOMParser`
to read its HTML rows and cells.
The `pending` map reserves positions covered by `rowspan`, while the nested
loops expand `colspan` into empty placeholders. Replacing ` ` elements with
newlines preserves multiline descriptions in a single Excel cell. The final
mapping pads every row to the same width.
The `safeCellText` function prefixes values beginning with `=`, `+`, `-`, or
`@` with an apostrophe. Excel then treats untrusted document content as text
instead of a formula.
### 3.7 Write the Grid and Finish the Interface
Continue `taskpane.js` with:
```javascript taskpane.js theme={null}
async function writeRows(rows) {
return Excel.run(async (context) => {
const sheets = context.workbook.worksheets;
sheets.load("items/name");
await context.sync();
const existingNames = new Set(
sheets.items.map((sheet) => sheet.name.toLowerCase()),
);
let sheetName = "Extracted";
let suffix = 2;
while (existingNames.has(sheetName.toLowerCase())) {
sheetName = `Extracted ${suffix}`;
suffix += 1;
}
const sheet = sheets.add(sheetName);
const range = sheet.getRangeByIndexes(
0,
0,
rows.length,
rows[0].length,
);
range.numberFormat = rows.map((row) => row.map(() => "@"));
range.values = rows;
range.format.autofitColumns();
range.format.autofitRows();
sheet.activate();
range.select();
await context.sync();
return sheetName;
});
}
```
Office.js uses a queued object model. `sheets.load("items/name")` requests the
existing worksheet names, and the first `context.sync()` retrieves them. We
use those names to avoid overwriting a previous extraction.
Setting every destination cell's number format to `@` makes Excel preserve
values such as `00124`, `$86.00`, and `02 July 2026` as extracted text.
Assigning `rows` to `range.values` then queues the complete table as a single
write. The second `context.sync()` sends the write, formatting, activation, and
selection operations to Excel together.
Finish `taskpane.js` with the interface helpers:
```javascript taskpane.js theme={null}
function updateButton() {
const missingInput = !keyInput.value.trim() || !fileInput.files.length;
button.disabled = isBusy || !excelReady || missingInput;
}
function showStatus(message, type = "") {
statusMessage.textContent = message;
statusMessage.className = type;
}
```
The button stays disabled when Excel isn't ready, either input is missing, or
an extraction is already running. At this point, `taskpane.js` should match the
complete version at the top of the guide.
## Step 4: Describe the Add-in to Excel
Excel doesn't discover a local web page by itself. The manifest identifies the
add-in, declares where it can run, points Excel to the task pane, and requests
workbook access.
### 4.1 Add the Identity and Display Details
Create `manifest.xml` beside the other project files and add:
```xml manifest.xml theme={null}
e8d967ef-5e99-4e86-88fd-57efb72c6ae81.0.0.0Unsiloed AIen-US
```
The first line declares an XML file encoded as UTF-8. The two namespace URLs
identify Microsoft's Office manifest vocabulary and the XML Schema Instance
vocabulary; Excel uses the latter to interpret `xsi:type="TaskPaneApp"`. The
URLs are identifiers, not pages the add-in downloads at runtime.
The remaining elements describe the add-in:
* `Id` is a stable UUID that distinguishes this add-in from others. Keep this
value for the recipe, or generate a new UUID with `uuidgen` when adapting it.
* `Version` tracks releases of the add-in rather than the Office.js version.
* `ProviderName` and `DefaultLocale` identify the publisher and fallback language.
* `DisplayName` and `Description` appear in Excel's add-in interface.
* `IconUrl` and `SupportUrl` provide the metadata expected by Microsoft's
distribution validator.
The root element remains open because we'll add Excel-specific settings next.
`IconUrl` and `SupportUrl` are formally optional for local sideloading, but
Microsoft's distribution validator expects them. They do not affect the
localhost connection.
### 4.2 Grant Excel Access and Load the Page
Append the rest of the manifest:
```xml manifest.xml theme={null}
ReadWriteDocument
```
The `Workbook` host restricts this add-in to Excel. `SourceLocation` must
match the HTTPS host and port in `vite.config.mjs`; it tells Excel which page
to place in the task pane. `ReadWriteDocument` allows Office.js to create the
worksheet and write the extracted values.
The manifest contains no API credentials. Excel uses it only to register the
add-in and decide what workbook access to grant.
## Step 5: Start the Add-in
From `unsiloed-excel-addin`, run:
```bash theme={null}
npx vite
```
The first run may ask for your macOS password to trust the localhost
certificate. Keep this Terminal window open.
Open this page in Chrome and confirm that it doesn't show a certificate warning:
```text theme={null}
https://localhost:3000/index.html
```
The page says **Open this page as an Excel add-in** when viewed directly. That is
expected because Office.js isn't connected to a workbook yet.
When Excel loads the localhost task pane, Chrome may ask whether Excel can
connect to devices on your local network. Click **Allow**. Chrome treats a
public site loading a loopback URL as local network access.
## Step 6: Load the Add-in in Excel
Open [Excel for the web](https://excel.cloud.microsoft), sign in, and create a
blank workbook.
Microsoft uses different labels for the custom add-in menu across account
types. If **Manage My Add-ins** isn't present, look for **More Settings** or
**Advanced**. The [Microsoft sideloading guide](https://learn.microsoft.com/en-us/office/dev/add-ins/testing/sideload-office-add-ins-for-testing)
documents the current routes. If none of these options appears, your
organization may have disabled custom add-ins; use a personal account or ask
your administrator to enable sideloading.
On the **Home** tab, click **Add-ins** (1).
In the panel that opens, click **More Add-ins** (2).
Open **Manage My Add-ins** and choose **Upload My Add-in** (3).
Click **Browse** (4), select `unsiloed-excel-addin/manifest.xml`, and then
click **Upload** (5). The Upload button becomes available after you select
the file.
The **Unsiloed Table Extractor** pane should open on the right side of the
workbook (6).
Excel for the web stores this sideloaded manifest in the current browser
profile. Upload it again if you clear the browser data or switch browsers.
## Step 7: Extract a Table
Use the task pane to process the sample document:
1. Enter your Unsiloed API key.
2. Click **Choose File** and select the downloaded `sample-invoice.pdf` file.
3. Click **Extract table**.
The numbered controls in the task pane follow the same order:
Unsiloed parses the document, and Excel creates an `Extracted` worksheet (4).
The sample produces 10 rows and four columns. The header row should contain
`DESCRIPTION`, `QTY`, `UNIT PRICE`, and `AMOUNT`; the final row should contain
`Total due` and `$11,557.22`. Running the add-in again creates `Extracted 2`,
preserving the earlier result.
The code formats the destination range as text before assigning values, which
preserves currency symbols, leading zeros, dates, percentages, and table
labels. Use [schema-based extraction](/docs/document-processing/extraction/extraction)
when you need typed fields in fixed columns.
## Troubleshoot the Add-in
Confirm that `npx vite` is still running and that
`https://localhost:3000/index.html` opens in the same browser without a
certificate warning. In Chrome's site settings for `excel.cloud.microsoft`,
set **Local network access** to **Allow**, reload the workbook, and upload
`manifest.xml` again.
Open **Home → Add-ins** and select **Unsiloed Table Extractor**. If it isn't
listed, use **Upload My Add-in** to load the manifest again.
Confirm that the table has visible rows or aligned columns, then try a
larger or sharper image. The recipe stops instead of writing ordinary text
segments when the parse result contains no `Table` segment.
Re-enter the Unsiloed API key. The add-in doesn't save the key between task
pane sessions.
## Limitations and Next Steps
This example has a few limitations:
* It writes the first `Table` segment. For documents with several tables, add a
preview or selector before calling `writeRows`.
* It writes every cell as text. Use typed schema extraction when downstream
formulas need numbers or dates.
* It processes one local file at a time and stops polling after five minutes.
* It keeps the API key in task-pane memory. That works for a personal local
tool, but a distributed add-in should use per-user authentication or a
backend that keeps shared credentials out of browser code.
The parsing and element-type references below are the best places to extend
the example without changing its Office.js foundation.
Review the asynchronous parse workflow and response structure.
See the segment types returned by parsing, including `Table`.
# Extract Chart Data with a Custom Function in Google Sheets
Source: https://docs.unsiloed.ai/docs/cookbooks/spreadsheets/google-sheets-custom-function
Create an UNSILOED_CHART custom function that turns a chart image into clean, numeric rows and columns in Google Sheets.
Charts often contain values that aren't available as selectable text. A Google
Sheets custom function can send a chart image to Unsiloed and return its
categories and series as editable cells.
In this guide, we'll create `UNSILOED_CHART`. The function downloads a chart
image, asks Unsiloed for a chart-ready Markdown table, and converts the table to
numeric spreadsheet values. The finished formula looks like this:
```excel theme={null}
=UNSILOED_CHART(A2, B2)
```
The function accepts an image URL in `A2`. The value in `B2` triggers another
status check when parsing takes longer than one formula execution. An optional
third argument returns a numeric x-axis for scatter charts. In the example
below, Unsiloed extracts 20 quarters from a dual-axis chart, including revenue
in millions of pounds and margin as a percentage. We then rebuild it as an
editable combination chart.
Custom functions can return values but can't insert or modify a Google Sheets
chart. After the data spills into the sheet, we'll use **Insert → Chart** to
create an editable chart from it.
## What We'll Build
The custom function performs five operations:
1. It downloads a public chart image and uploads the bytes to `/parse`.
2. It asks the page analyzer to end its response with one Markdown data table.
3. It stores the returned job ID so later calculations can reuse it.
4. It checks the job once during each calculation.
5. It converts the table into validated, chart-ready rows and numeric series.
Google stops a custom function after 30 seconds, and Apps Script can't interrupt
an outstanding URL fetch. If the job is still running, the formula returns a
status message. Change the refresh value later to check the saved job again. For
slower or business-critical workflows, use the menu-driven cookbook instead.
Paste this into `Code.gs`, replacing its contents. The script reads the API key
from the script property configured in [Step 3](#step-3-store-your-api-key).
```javascript Code.gs theme={null}
const UNSILOED_BASE_URL = "https://prod.visionapi.unsiloed.ai";
const UNSILOED_JOB_MAX_AGE_MS = 24 * 60 * 60 * 1000;
const UNSILOED_MAX_CACHED_JOBS = 50;
const UNSILOED_JOB_VERSION = "chart-v6";
const UNSILOED_JOB_PREFIX = "UNSILOED_JOB_";
const UNSILOED_CHART_PROMPT = [
"Extract every visible chart data point.",
"Preserve category labels, series names, and units exactly.",
"End with exactly one Markdown table.",
"Put category or x-axis values in the first column",
"and numeric series encoded by chart marks in the remaining columns.",
"Do not add columns that only repeat plotted values as text, rankings, annotations,",
"or derived comparisons.",
"Every data row must contain the plotted value when it is visible.",
"Include visible units in series headers.",
"Keep values on multiple axes in their original units.",
"For visible error bars, add numeric x error and y error columns.",
"Use an empty cell only for a genuinely missing value.",
"Use a period as the decimal separator.",
"Do not put units or grouping separators in numeric cells."
].join(" ");
/**
* Extracts chart data from a public image URL.
*
* @param {string} imageUrl A public PNG, JPEG, or TIFF URL.
* @param {*} refresh Change this value to check an unfinished job again.
* @param {boolean} numericX Use TRUE to return a numeric first column.
* @return {Array>} Chart-ready data.
* @customfunction
*/
function UNSILOED_CHART(imageUrl, refresh, numericX) {
void refresh;
if (Array.isArray(imageUrl)) imageUrl = imageUrl[0][0];
imageUrl = String(imageUrl || "").trim();
if (!/^https?:\/\//i.test(imageUrl)) {
throw new Error("Pass a public HTTP image URL.");
}
const apiKey = unsiloedApiKey();
const jobId = getOrCreateChartJob(imageUrl, apiKey);
let job;
try {
job = unsiloedRequest(
"/parse/" + encodeURIComponent(jobId),
apiKey
);
} catch (error) {
if (error.httpStatus === 404) {
const cleared = clearCachedChartJob(imageUrl, apiKey, jobId);
const action = cleared
? "The saved job was cleared. Change the refresh value to submit it again."
: "Change the refresh value to try again.";
throw new Error("The saved parse job no longer exists. " + action);
}
throw error;
}
if (job.status === "Succeeded") {
return firstChartRows(job, numericX === true);
}
if (job.status === "Failed" || job.status === "Cancelled") {
const cleared = clearCachedChartJob(imageUrl, apiKey, jobId);
const action = cleared
? "The saved job was cleared. Change the refresh value to submit it again."
: "Change the refresh value to try again.";
throw new Error(
(job.message || "Parsing " + job.status.toLowerCase() + ".") + " " + action
);
}
return [["Still processing. Change the refresh value to check again."]];
}
function unsiloedApiKey() {
const apiKey = PropertiesService.getScriptProperties()
.getProperty("UNSILOED_API_KEY");
if (!apiKey) throw new Error("Add UNSILOED_API_KEY to the script properties.");
return apiKey;
}
function getOrCreateChartJob(imageUrl, apiKey) {
const properties = PropertiesService.getScriptProperties();
const propertyName = jobPropertyName(imageUrl, apiKey);
const cached = readCachedChartJob(properties, propertyName);
if (cached) return cached.jobId;
const lock = LockService.getScriptLock();
if (!lock.tryLock(5000)) throw new Error("Try the formula again.");
try {
const checkedAgain = readCachedChartJob(properties, propertyName);
if (checkedAgain) return checkedAgain.jobId;
pruneChartJobs(properties);
const image = fetchChartImage(imageUrl);
const started = unsiloedRequest("/parse", apiKey, {
method: "post",
payload: chartParsePayload(image)
});
if (typeof started.job_id !== "string" || !started.job_id) {
throw new Error("Unsiloed returned no job ID.");
}
properties.setProperty(propertyName, JSON.stringify({
jobId: started.job_id,
createdAt: Date.now()
}));
return started.job_id;
} finally {
lock.releaseLock();
}
}
function readCachedChartJob(properties, propertyName) {
const saved = properties.getProperty(propertyName);
if (!saved) return null;
try {
const state = JSON.parse(saved);
const age = Date.now() - state.createdAt;
if (
typeof state.jobId === "string" &&
state.jobId &&
Number.isFinite(state.createdAt) &&
age >= 0 &&
age < UNSILOED_JOB_MAX_AGE_MS
) {
return state;
}
} catch (error) {
// Invalid entries are removed during the next locked cleanup.
}
return null;
}
function clearCachedChartJob(imageUrl, apiKey, jobId) {
const properties = PropertiesService.getScriptProperties();
const propertyName = jobPropertyName(imageUrl, apiKey);
const lock = LockService.getScriptLock();
if (!lock.tryLock(5000)) return false;
try {
const saved = properties.getProperty(propertyName);
if (!saved) return true;
try {
const state = JSON.parse(saved);
if (state.jobId !== jobId) return false;
} catch (error) {
// The matching property is unusable and can be removed.
}
properties.deleteProperty(propertyName);
return true;
} finally {
lock.releaseLock();
}
}
function pruneChartJobs(properties) {
const all = properties.getProperties();
const now = Date.now();
const valid = [];
Object.keys(all).forEach(key => {
if (key.indexOf(UNSILOED_JOB_PREFIX) !== 0) return;
try {
const state = JSON.parse(all[key]);
const age = now - state.createdAt;
if (
typeof state.jobId === "string" &&
state.jobId &&
Number.isFinite(state.createdAt) &&
age >= 0 &&
age < UNSILOED_JOB_MAX_AGE_MS
) {
valid.push({ key: key, createdAt: state.createdAt });
}
} catch (error) {
// Invalid entries are excluded from the retained set.
}
});
valid.sort((a, b) => b.createdAt - a.createdAt);
const retained = new Set(
valid.slice(0, Math.max(UNSILOED_MAX_CACHED_JOBS - 1, 0))
.map(entry => entry.key)
);
Object.keys(all).forEach(key => {
if (
key.indexOf(UNSILOED_JOB_PREFIX) === 0 &&
!retained.has(key)
) {
properties.deleteProperty(key);
}
});
}
function chartParsePayload(image) {
const segmentAnalysis = {
Page: {
markdown: "VLM",
model_id: "nova",
vlm: UNSILOED_CHART_PROMPT
}
};
const outputFields = {
html: false,
markdown: true,
ocr: false,
image: false,
content: false,
bbox: false,
confidence: false,
embed: false,
chart_data: false
};
return {
file: image,
use_high_resolution: "true",
layout_analysis: "page_by_page",
segment_analysis: JSON.stringify(segmentAnalysis),
response_profile: "custom",
output_fields: JSON.stringify(outputFields)
};
}
function fetchChartImage(imageUrl) {
const response = UrlFetchApp.fetch(imageUrl, {
followRedirects: true,
muteHttpExceptions: true
});
const status = response.getResponseCode();
if (status < 200 || status >= 300) {
throw new Error("The image URL returned HTTP " + status + ".");
}
const image = response.getBlob();
if (!/^image\//i.test(image.getContentType())) {
throw new Error("The URL did not return an image.");
}
const path = imageUrl.split(/[?#]/)[0];
const fileName = path.substring(path.lastIndexOf("/") + 1) || "chart.png";
return image.setName(fileName);
}
function jobPropertyName(imageUrl, apiKey) {
const digest = Utilities.computeDigest(
Utilities.DigestAlgorithm.SHA_256,
UNSILOED_JOB_VERSION + "\n" + imageUrl + "\n" + apiKey,
Utilities.Charset.UTF_8
);
const hex = digest.map(byte =>
(byte + 256).toString(16).slice(-2)
).join("");
return UNSILOED_JOB_PREFIX + hex;
}
function unsiloedRequest(path, apiKey, options) {
const requestOptions = options || {};
const response = UrlFetchApp.fetch(UNSILOED_BASE_URL + path, {
muteHttpExceptions: true,
...requestOptions,
headers: {
...(requestOptions.headers || {}),
"api-key": apiKey
}
});
const text = response.getContentText();
const status = response.getResponseCode();
if (status >= 200 && status < 300) return JSON.parse(text);
let message = text;
try {
const body = JSON.parse(text);
message = body.error?.message || body.message || body.detail || text;
} catch (error) {
// Keep the plain-text response.
}
const requestError = new Error(
message || "Unsiloed returned HTTP " + status + "."
);
requestError.httpStatus = status;
throw requestError;
}
function firstChartRows(job, numericX) {
const tables = (job.chunks || [])
.flatMap(chunk => chunk.segments || [])
.flatMap(segment => markdownTables(segment.markdown || ""));
const candidates = [];
let conversionError = null;
tables.forEach(rows => {
if (rows.length < 2 || rows[0].length < 2) return;
try {
const converted = rows.map((row, rowIndex) => row.map(
(cell, columnIndex) =>
rowIndex === 0 || (columnIndex === 0 && !numericX)
? cell
: chartValue(cell)
));
if (numericX) {
converted.slice(1).forEach((row, rowIndex) => {
if (typeof row[0] !== "number") {
throw new Error(
'X-axis value in row ' + (rowIndex + 2) +
' is not numeric: "' + row[0] + '".'
);
}
});
}
candidates.push(chartReadyRows(converted));
} catch (error) {
conversionError = error;
}
});
if (candidates.length > 1) {
throw new Error("Unsiloed returned more than one chart-ready table.");
}
if (candidates.length === 1) return candidates[0];
if (conversionError) throw conversionError;
throw new Error("Unsiloed did not return a chart data table.");
}
function markdownTables(markdown) {
const lines = markdown.split(/\r?\n/);
const tables = [];
for (let index = 0; index < lines.length - 1; index += 1) {
const header = splitMarkdownRow(lines[index]);
const separator = splitMarkdownRow(lines[index + 1]);
if (
header.length < 2 ||
separator.length !== header.length ||
!separator.every(cell => /^:?-+:?$/.test(cell))
) {
continue;
}
const width = header.length;
const rows = [header];
index += 2;
while (index < lines.length && isMarkdownRow(lines[index])) {
const row = splitMarkdownRow(lines[index]);
if (row.length > width) {
throw new Error("A Markdown data row has more cells than its header.");
}
rows.push(row.concat(Array(width - row.length).fill("")));
index += 1;
}
if (rows.length > 1) tables.push(rows);
index -= 1;
}
return tables;
}
function isMarkdownRow(line) {
return splitMarkdownRow(line).length > 1;
}
function splitMarkdownRow(line) {
let body = String(line || "").trim();
if (body.charAt(0) === "|") body = body.slice(1);
if (endsWithUnescapedPipe(body)) body = body.slice(0, -1);
const cells = [];
let cell = "";
for (let index = 0; index < body.length; index += 1) {
const character = body.charAt(index);
const next = body.charAt(index + 1);
if (character === "\\" && next && isMarkdownEscapable(next)) {
cell += next;
index += 1;
} else if (character === "|") {
cells.push(cleanMarkdownCell(cell));
cell = "";
} else {
cell += character;
}
}
cells.push(cleanMarkdownCell(cell));
return cells;
}
function endsWithUnescapedPipe(text) {
if (text.charAt(text.length - 1) !== "|") return false;
let backslashes = 0;
for (
let index = text.length - 2;
index >= 0 && text.charAt(index) === "\\";
index -= 1
) {
backslashes += 1;
}
return backslashes % 2 === 0;
}
function isMarkdownEscapable(character) {
const code = character.charCodeAt(0);
return (
(code >= 33 && code <= 47) ||
(code >= 58 && code <= 64) ||
(code >= 91 && code <= 96) ||
(code >= 123 && code <= 126)
);
}
function cleanMarkdownCell(cell) {
return cell.trim()
.replace(/^\*\*(.*)\*\*$/, "$1")
.replace(/`([^`]*)`/g, "$1");
}
function chartReadyRows(rows) {
const columns = [0];
for (let column = 1; column < rows[0].length; column += 1) {
const values = rows.slice(1).map(row => row[column]);
const populated = values.filter(value => value !== "");
const numeric = populated.filter(value => typeof value === "number");
if (populated.length === 0) {
throw new Error(
'Column "' + (rows[0][column] || column + 1) + '" has no values.'
);
}
if (numeric.length !== populated.length) {
const invalid = populated.find(value => typeof value !== "number");
throw new Error(
'Column "' + (rows[0][column] || column + 1) +
'" contains a nonnumeric value: "' + invalid + '".'
);
}
columns.push(column);
}
if (columns.length === 1) {
throw new Error("Unsiloed returned no numeric chart series.");
}
return rows.map(row => columns.map(column => row[column]));
}
function chartValue(cell) {
const value = cell.trim();
if (!value) return "";
const normalized = value.replace(/\u2212/g, "-");
const looksNumeric = /^[+-]?(?:\d[\d.,]*|\.\d+)(?:e[+-]?\d+)?$/i
.test(normalized);
if (looksNumeric && normalized.indexOf(",") !== -1) {
throw new Error('Numeric cell "' + value + '" contains a comma.');
}
if (!/^[+-]?(?:\d+\.?\d*|\.\d+)(?:e[+-]?\d+)?$/i.test(normalized)) {
return value;
}
const number = Number(normalized);
if (!Number.isFinite(number)) {
throw new Error('Numeric cell "' + value + '" is outside the supported range.');
}
if (Number.isInteger(number) && !Number.isSafeInteger(number)) {
throw new Error('Integer cell "' + value + '" cannot be represented exactly.');
}
return number;
}
```
## Requirements for the Custom Function
Before you start, gather:
* A Google account and a spreadsheet you can edit
* An Unsiloed API key from the [Unsiloed dashboard](https://app.unsiloed.ai)
* A public PNG, JPEG, or TIFF URL that returns the image bytes
A private Google Drive sharing link doesn't work because the custom function
can't use an interactive download page. Use a direct public URL or a presigned
cloud-storage URL.
This guide uses the synthetic D11 dual-axis chart from the Unsiloed
[ChartParse-Bench dataset](https://huggingface.co/datasets/Unsiloed/chart-parse-bench),
published under CC BY 4.0. The dataset accompanies the [chart-parsing benchmark
article](https://www.unsiloed.ai/blog/unsiloed-achieves-sota-chart-parsing):
```text theme={null}
https://huggingface.co/datasets/Unsiloed/chart-parse-bench/resolve/ba7316eb4c398963163b8f2d21d97142db1d8ce6/images/D11.png
```
## Step 1: Open the Apps Script Editor
The custom function lives in the Apps Script project attached to the
spreadsheet.
In Google Sheets, choose **Extensions → Apps Script**.
Google creates a bound script project and opens `Code.gs` in a new tab. Delete
the empty `myFunction` that Google adds to a new project.
## Step 2: Build the Custom Function
We'll build the same script in twelve focused sections. Keep each addition in
`Code.gs` in the order shown below. If you copied the completed script from the
accordion, use this section to understand its request, cache, and
table-conversion logic.
### 2.1 Configure Chart Extraction
At the top of `Code.gs`, add the API URL, cache limits, extraction prompt, and
API key helper:
```javascript Code.gs theme={null}
const UNSILOED_BASE_URL = "https://prod.visionapi.unsiloed.ai";
const UNSILOED_JOB_MAX_AGE_MS = 24 * 60 * 60 * 1000;
const UNSILOED_MAX_CACHED_JOBS = 50;
const UNSILOED_JOB_VERSION = "chart-v6";
const UNSILOED_JOB_PREFIX = "UNSILOED_JOB_";
const UNSILOED_CHART_PROMPT = [
"Extract every visible chart data point.",
"Preserve category labels, series names, and units exactly.",
"End with exactly one Markdown table.",
"Put category or x-axis values in the first column",
"and numeric series encoded by chart marks in the remaining columns.",
"Do not add columns that only repeat plotted values as text, rankings, annotations,",
"or derived comparisons.",
"Every data row must contain the plotted value when it is visible.",
"Include visible units in series headers.",
"Keep values on multiple axes in their original units.",
"For visible error bars, add numeric x error and y error columns.",
"Use an empty cell only for a genuinely missing value.",
"Use a period as the decimal separator.",
"Do not put units or grouping separators in numeric cells."
].join(" ");
function unsiloedApiKey() {
const apiKey = PropertiesService.getScriptProperties()
.getProperty("UNSILOED_API_KEY");
if (!apiKey) throw new Error("Add UNSILOED_API_KEY to the script properties.");
return apiKey;
}
```
The prompt asks for one rectangular table with canonical numeric cells. It also
keeps multi-axis values in their original units and preserves visible error bars
as separate columns. The job version and API-key fingerprint become part of the
cache key. We'll add `UNSILOED_API_KEY` in
[Step 3](#step-3-store-your-api-key).
### 2.2 Check the Parse Job
Below `unsiloedApiKey`, add the function that Sheets calls:
```javascript Code.gs theme={null}
/**
* Extracts chart data from a public image URL.
*
* @param {string} imageUrl A public PNG, JPEG, or TIFF URL.
* @param {*} refresh Change this value to check an unfinished job again.
* @param {boolean} numericX Use TRUE to return a numeric first column.
* @return {Array>} Chart-ready data.
* @customfunction
*/
function UNSILOED_CHART(imageUrl, refresh, numericX) {
void refresh;
if (Array.isArray(imageUrl)) imageUrl = imageUrl[0][0];
imageUrl = String(imageUrl || "").trim();
if (!/^https?:\/\//i.test(imageUrl)) {
throw new Error("Pass a public HTTP image URL.");
}
const apiKey = unsiloedApiKey();
const jobId = getOrCreateChartJob(imageUrl, apiKey);
let job;
try {
job = unsiloedRequest(
"/parse/" + encodeURIComponent(jobId),
apiKey
);
} catch (error) {
if (error.httpStatus === 404) {
const cleared = clearCachedChartJob(imageUrl, apiKey, jobId);
const action = cleared
? "The saved job was cleared. Change the refresh value to submit it again."
: "Change the refresh value to try again.";
throw new Error("The saved parse job no longer exists. " + action);
}
throw error;
}
if (job.status === "Succeeded") {
return firstChartRows(job, numericX === true);
}
if (job.status === "Failed" || job.status === "Cancelled") {
const cleared = clearCachedChartJob(imageUrl, apiKey, jobId);
const action = cleared
? "The saved job was cleared. Change the refresh value to submit it again."
: "Change the refresh value to try again.";
throw new Error(
(job.message || "Parsing " + job.status.toLowerCase() + ".") + " " + action
);
}
return [["Still processing. Change the refresh value to check again."]];
}
```
The `@customfunction` tag adds `UNSILOED_CHART` to Sheets autocomplete. Each
calculation performs one status request. If the job is still running, changing
`refresh` starts another calculation that checks the saved job ID.
### 2.3 Reuse or Create a Parse Job
Below `UNSILOED_CHART`, add the helper that reads the cache before taking a lock,
then checks it again before submitting an image:
```javascript Code.gs theme={null}
function getOrCreateChartJob(imageUrl, apiKey) {
const properties = PropertiesService.getScriptProperties();
const propertyName = jobPropertyName(imageUrl, apiKey);
const cached = readCachedChartJob(properties, propertyName);
if (cached) return cached.jobId;
const lock = LockService.getScriptLock();
if (!lock.tryLock(5000)) throw new Error("Try the formula again.");
try {
const checkedAgain = readCachedChartJob(properties, propertyName);
if (checkedAgain) return checkedAgain.jobId;
pruneChartJobs(properties);
const image = fetchChartImage(imageUrl);
const started = unsiloedRequest("/parse", apiKey, {
method: "post",
payload: chartParsePayload(image)
});
if (typeof started.job_id !== "string" || !started.job_id) {
throw new Error("Unsiloed returned no job ID.");
}
properties.setProperty(propertyName, JSON.stringify({
jobId: started.job_id,
createdAt: Date.now()
}));
return started.job_id;
} finally {
lock.releaseLock();
}
}
```
Reading before locking lets an existing job return immediately. The second read
prevents two simultaneous calculations from creating ordinary duplicate jobs.
A remote job accepted immediately before Apps Script terminates may still lose
its ID because the parse API doesn't expose an idempotency key.
### 2.4 Validate and Clear Cached Jobs
Below `getOrCreateChartJob`, add helpers that reject malformed cache entries and
remove a failed or missing job without deleting a newer replacement:
```javascript Code.gs theme={null}
function readCachedChartJob(properties, propertyName) {
const saved = properties.getProperty(propertyName);
if (!saved) return null;
try {
const state = JSON.parse(saved);
const age = Date.now() - state.createdAt;
if (
typeof state.jobId === "string" &&
state.jobId &&
Number.isFinite(state.createdAt) &&
age >= 0 &&
age < UNSILOED_JOB_MAX_AGE_MS
) {
return state;
}
} catch (error) {
// Invalid entries are removed during the next locked cleanup.
}
return null;
}
function clearCachedChartJob(imageUrl, apiKey, jobId) {
const properties = PropertiesService.getScriptProperties();
const propertyName = jobPropertyName(imageUrl, apiKey);
const lock = LockService.getScriptLock();
if (!lock.tryLock(5000)) return false;
try {
const saved = properties.getProperty(propertyName);
if (!saved) return true;
try {
const state = JSON.parse(saved);
if (state.jobId !== jobId) return false;
} catch (error) {
// The matching property is unusable and can be removed.
}
properties.deleteProperty(propertyName);
return true;
} finally {
lock.releaseLock();
}
}
```
The job ID comparison matters when concurrent calculations overlap. It prevents
an old failed response from removing a newer job saved under the same key.
### 2.5 Bound the Job Cache
Below `clearCachedChartJob`, add cleanup and cache-key helpers:
```javascript Code.gs theme={null}
function pruneChartJobs(properties) {
const all = properties.getProperties();
const now = Date.now();
const valid = [];
Object.keys(all).forEach(key => {
if (key.indexOf(UNSILOED_JOB_PREFIX) !== 0) return;
try {
const state = JSON.parse(all[key]);
const age = now - state.createdAt;
if (
typeof state.jobId === "string" &&
state.jobId &&
Number.isFinite(state.createdAt) &&
age >= 0 &&
age < UNSILOED_JOB_MAX_AGE_MS
) {
valid.push({ key: key, createdAt: state.createdAt });
}
} catch (error) {
// Invalid entries are excluded from the retained set.
}
});
valid.sort((a, b) => b.createdAt - a.createdAt);
const retained = new Set(
valid.slice(0, Math.max(UNSILOED_MAX_CACHED_JOBS - 1, 0))
.map(entry => entry.key)
);
Object.keys(all).forEach(key => {
if (
key.indexOf(UNSILOED_JOB_PREFIX) === 0 &&
!retained.has(key)
) {
properties.deleteProperty(key);
}
});
}
function jobPropertyName(imageUrl, apiKey) {
const digest = Utilities.computeDigest(
Utilities.DigestAlgorithm.SHA_256,
UNSILOED_JOB_VERSION + "\n" + imageUrl + "\n" + apiKey,
Utilities.Charset.UTF_8
);
const hex = digest.map(byte =>
(byte + 256).toString(16).slice(-2)
).join("");
return UNSILOED_JOB_PREFIX + hex;
}
```
Cleanup retains at most 50 current jobs and removes expired, malformed, and
older entries. Hashing the API key into the property name starts a new cache when
the spreadsheet changes Unsiloed accounts without exposing the key.
### 2.6 Build the Chart Parse Request
Below `jobPropertyName`, add the payload that tells Unsiloed how to analyze the
image and which response fields to return:
```javascript Code.gs theme={null}
function chartParsePayload(image) {
const segmentAnalysis = {
Page: {
markdown: "VLM",
model_id: "nova",
vlm: UNSILOED_CHART_PROMPT
}
};
const outputFields = {
html: false,
markdown: true,
ocr: false,
image: false,
content: false,
bbox: false,
confidence: false,
embed: false,
chart_data: false
};
return {
file: image,
use_high_resolution: "true",
layout_analysis: "page_by_page",
segment_analysis: JSON.stringify(segmentAnalysis),
response_profile: "custom",
output_fields: JSON.stringify(outputFields)
};
}
```
The page analyzer sees the complete standalone chart, including its labels and
axes. Because the image is one chart, the whole page is the chart, and the prompt
controls the shape of the table that comes back. The custom response profile
returns the Markdown table the sheet needs and omits fields the formula doesn't
use.
The request turns off `chart_data` because the prompt already produces the
table. For charts embedded in multi-page documents, send `extract_charts=true`
instead and read the structured `chart_data` field on each chart segment. See
the [parse job response reference](/docs/api-reference/parser/get-parse-job-status)
for that field.
### 2.7 Download the Source Image
Below `chartParsePayload`, add the image downloader:
```javascript Code.gs theme={null}
function fetchChartImage(imageUrl) {
const response = UrlFetchApp.fetch(imageUrl, {
followRedirects: true,
muteHttpExceptions: true
});
const status = response.getResponseCode();
if (status < 200 || status >= 300) {
throw new Error("The image URL returned HTTP " + status + ".");
}
const image = response.getBlob();
if (!/^image\//i.test(image.getContentType())) {
throw new Error("The URL did not return an image.");
}
const path = imageUrl.split(/[?#]/)[0];
const fileName = path.substring(path.lastIndexOf("/") + 1) || "chart.png";
return image.setName(fileName);
}
```
Downloading in Apps Script ensures Unsiloed receives image bytes and the correct
content type. Apps Script doesn't expose a per-request fetch timeout, so a slow
download or API request can still reach the custom function's runtime limit.
### 2.8 Send Authenticated API Requests
Below `fetchChartImage`, add the shared request helper:
```javascript Code.gs theme={null}
function unsiloedRequest(path, apiKey, options) {
const requestOptions = options || {};
const response = UrlFetchApp.fetch(UNSILOED_BASE_URL + path, {
muteHttpExceptions: true,
...requestOptions,
headers: {
...(requestOptions.headers || {}),
"api-key": apiKey
}
});
const text = response.getContentText();
const status = response.getResponseCode();
if (status >= 200 && status < 300) return JSON.parse(text);
let message = text;
try {
const body = JSON.parse(text);
message = body.error?.message || body.message || body.detail || text;
} catch (error) {
// Keep the plain-text response.
}
const requestError = new Error(
message || "Unsiloed returned HTTP " + status + "."
);
requestError.httpStatus = status;
throw requestError;
}
```
The helper adds the API key and surfaces API error messages in the formula cell.
It also preserves the HTTP status so the main function can clear a missing job
without treating temporary server errors as permanent.
### 2.9 Select the Chart Data Table
Below `unsiloedRequest`, add the helper that validates every extracted table:
```javascript Code.gs theme={null}
function firstChartRows(job, numericX) {
const tables = (job.chunks || [])
.flatMap(chunk => chunk.segments || [])
.flatMap(segment => markdownTables(segment.markdown || ""));
const candidates = [];
let conversionError = null;
tables.forEach(rows => {
if (rows.length < 2 || rows[0].length < 2) return;
try {
const converted = rows.map((row, rowIndex) => row.map(
(cell, columnIndex) =>
rowIndex === 0 || (columnIndex === 0 && !numericX)
? cell
: chartValue(cell)
));
if (numericX) {
converted.slice(1).forEach((row, rowIndex) => {
if (typeof row[0] !== "number") {
throw new Error(
'X-axis value in row ' + (rowIndex + 2) +
' is not numeric: "' + row[0] + '".'
);
}
});
}
candidates.push(chartReadyRows(converted));
} catch (error) {
conversionError = error;
}
});
if (candidates.length > 1) {
throw new Error("Unsiloed returned more than one chart-ready table.");
}
if (candidates.length === 1) return candidates[0];
if (conversionError) throw conversionError;
throw new Error("Unsiloed did not return a chart data table.");
}
```
The function accepts a table only after numeric-series validation. It preserves
the first column as text by default, protecting numeric-looking category labels.
Pass `TRUE` as the optional third formula argument when a scatter chart needs a
numeric x-axis. In that mode, a missing or nonnumeric x value produces an error.
### 2.10 Parse Markdown Tables
Below `firstChartRows`, add the tolerant Markdown table parser:
```javascript Code.gs theme={null}
function markdownTables(markdown) {
const lines = markdown.split(/\r?\n/);
const tables = [];
for (let index = 0; index < lines.length - 1; index += 1) {
const header = splitMarkdownRow(lines[index]);
const separator = splitMarkdownRow(lines[index + 1]);
if (
header.length < 2 ||
separator.length !== header.length ||
!separator.every(cell => /^:?-+:?$/.test(cell))
) {
continue;
}
const width = header.length;
const rows = [header];
index += 2;
while (index < lines.length && isMarkdownRow(lines[index])) {
const row = splitMarkdownRow(lines[index]);
if (row.length > width) {
throw new Error("A Markdown data row has more cells than its header.");
}
rows.push(row.concat(Array(width - row.length).fill("")));
index += 1;
}
if (rows.length > 1) tables.push(rows);
index -= 1;
}
return tables;
}
function isMarkdownRow(line) {
return splitMarkdownRow(line).length > 1;
}
```
The parser accepts tables with or without outside pipes, verifies that the
header and delimiter widths match, and pads short data rows into a rectangular
array. It rejects wider rows because an extra pipe could otherwise shift or
hide a value.
### 2.11 Split Markdown Cells Safely
Below `isMarkdownRow`, add the row parser and backslash helpers:
```javascript Code.gs theme={null}
function splitMarkdownRow(line) {
let body = String(line || "").trim();
if (body.charAt(0) === "|") body = body.slice(1);
if (endsWithUnescapedPipe(body)) body = body.slice(0, -1);
const cells = [];
let cell = "";
for (let index = 0; index < body.length; index += 1) {
const character = body.charAt(index);
const next = body.charAt(index + 1);
if (character === "\\" && next && isMarkdownEscapable(next)) {
cell += next;
index += 1;
} else if (character === "|") {
cells.push(cleanMarkdownCell(cell));
cell = "";
} else {
cell += character;
}
}
cells.push(cleanMarkdownCell(cell));
return cells;
}
function endsWithUnescapedPipe(text) {
if (text.charAt(text.length - 1) !== "|") return false;
let backslashes = 0;
for (
let index = text.length - 2;
index >= 0 && text.charAt(index) === "\\";
index -= 1
) {
backslashes += 1;
}
return backslashes % 2 === 0;
}
function isMarkdownEscapable(character) {
const code = character.charCodeAt(0);
return (
(code >= 33 && code <= 47) ||
(code >= 58 && code <= 64) ||
(code >= 91 && code <= 96) ||
(code >= 123 && code <= 126)
);
}
function cleanMarkdownCell(cell) {
return cell.trim()
.replace(/^\*\*(.*)\*\*$/, "$1")
.replace(/`([^`]*)`/g, "$1");
}
```
Only punctuation consumes a Markdown backslash escape. Ordinary text such as
`C:\temp` keeps its literal backslash, while `\|` remains a pipe inside its
cell.
### 2.12 Validate Series and Numbers
At the bottom of `Code.gs`, add the numeric-series and cell-value validators:
```javascript Code.gs theme={null}
function chartReadyRows(rows) {
const columns = [0];
for (let column = 1; column < rows[0].length; column += 1) {
const values = rows.slice(1).map(row => row[column]);
const populated = values.filter(value => value !== "");
const numeric = populated.filter(value => typeof value === "number");
if (populated.length === 0) {
throw new Error(
'Column "' + (rows[0][column] || column + 1) + '" has no values.'
);
}
if (numeric.length !== populated.length) {
const invalid = populated.find(value => typeof value !== "number");
throw new Error(
'Column "' + (rows[0][column] || column + 1) +
'" contains a nonnumeric value: "' + invalid + '".'
);
}
columns.push(column);
}
if (columns.length === 1) {
throw new Error("Unsiloed returned no numeric chart series.");
}
return rows.map(row => columns.map(column => row[column]));
}
function chartValue(cell) {
const value = cell.trim();
if (!value) return "";
const normalized = value.replace(/\u2212/g, "-");
const looksNumeric = /^[+-]?(?:\d[\d.,]*|\.\d+)(?:e[+-]?\d+)?$/i
.test(normalized);
if (looksNumeric && normalized.indexOf(",") !== -1) {
throw new Error('Numeric cell "' + value + '" contains a comma.');
}
if (!/^[+-]?(?:\d+\.?\d*|\.\d+)(?:e[+-]?\d+)?$/i.test(normalized)) {
return value;
}
const number = Number(normalized);
if (!Number.isFinite(number)) {
throw new Error('Numeric cell "' + value + '" is outside the supported range.');
}
if (Number.isInteger(number) && !Number.isSafeInteger(number)) {
throw new Error('Integer cell "' + value + '" cannot be represented exactly.');
}
return number;
}
```
Every returned series must contain numeric values. Empty, text-only, and mixed
numeric/text columns fail with the column name and offending value instead of
disappearing silently. Equal numeric series remain separate, and numeric
overflow or unsafe integers produce an error rather than a corrupted
spreadsheet value.
When `Code.gs` contains all twelve sections, save the project and confirm that
Apps Script reports no syntax error. The finished script should use the
`chart-v6` cache version and accept `imageUrl`, `refresh`, and `numericX` as
the function's three arguments.
## Step 3: Store Your API Key
Keep the API key out of spreadsheet cells and source code by storing it as a
script property.
In the Apps Script editor, open **Project Settings** from the left sidebar.
Scroll to **Script Properties**, then click **Add script property**.
Set the property name to `UNSILOED_API_KEY`, paste your API key into its value
field, and click **Save script properties**.
Spreadsheet editors can view its bound Apps Script project and script
properties. Use a dedicated API key, restrict edit access to trusted
collaborators, and rotate the key if the spreadsheet is shared unexpectedly.
Treat presigned image URLs as credentials too, and give them short expiration
times.
## Step 4: Extract the Chart Data
We'll keep the image URL and refresh value separate from the formula so a status
check doesn't require editing the formula text.
Return to the spreadsheet and add these labels and values:
| Cell | Value |
| - | - |
| `A1` | `Chart image URL` |
| `A2` | `https://huggingface.co/datasets/Unsiloed/chart-parse-bench/resolve/ba7316eb4c398963163b8f2d21d97142db1d8ce6/images/D11.png` |
| `B1` | `Refresh` |
| `B2` | `1` |
The red box shows the input range. The URL appears truncated in the cell,
but Sheets retains its complete value.
Select `A4` and enter:
```excel theme={null}
=UNSILOED_CHART(A2, B2)
```
For a scatter chart that needs numeric x values, use:
```excel theme={null}
=UNSILOED_CHART(A2, B2, TRUE)
```
Sheets displays `Loading...` while the function submits or checks the parse
job. Keep the cells below and to the right of `A4` empty so the result has
room to expand.
The first calculation may return this status instead of data:
```text theme={null}
Still processing. Change the refresh value to check again.
```
Wait about 20 seconds, then increase `B2` from `1` to `2`. Sheets
recalculates the function and checks the saved job instead of submitting
the image again. Repeat after another short wait if the job is still
processing.
When the job succeeds, the formula expands into the chart's categories and
series.
The result contains one row for each of the 20 quarters. Unsiloed keeps the
`revenue (£m)` and `margin (%)` units in the headers while estimating the plotted
values from the chart's pixels. For example, the extracted quarter 7 margin is
`16.2`, while the synthetic source data is approximately `16.15`. Verify values
against source data before using them for consequential decisions.
The margin values use a 0–100 percentage scale, so `13.9` means 13.9%. Don't
apply Sheets' percentage format, which would display that value as 1,390%.
Each image URL reuses a saved job for up to 24 hours, subject to the 50-entry
cache limit. To deliberately submit a fresh job, change
`UNSILOED_JOB_VERSION`, save `Code.gs`, and change `B2`. If you instead delete
its `UNSILOED_JOB_...` script property, also change `B2` to recalculate the
formula.
## Step 5: Recreate the Chart in Google Sheets
The extracted cells become the data source for the native chart.
Select the extracted category and value columns, including their headers. Then
choose **Insert → Chart**.
In the Chart editor, choose **Clustered Column - Line on Secondary Axis**. This
uses columns for revenue and a line for margin, matching the source chart's two
scales. Enable **Use row 4 as headers** and **Use column A as labels** if Sheets
doesn't detect them automatically.
The recreated chart is editable and linked to the extracted cells. It preserves
the extracted categories, estimated values, series types, units, and two axes.
Google Sheets applies its own fonts, spacing, and colors. The source image's
gray highlighted interval is a visual annotation rather than a data series, so
add that styling manually if you need it. Treat the result as a faithful data
reconstruction, not a pixel-for-pixel copy of the source image.
## Troubleshoot the Custom Function
The first argument is empty, a filename, or a non-HTTP value. Put a complete
public or presigned image URL in the referenced cell.
The URL returned an HTML download page instead of image bytes. Use a direct
PNG, JPEG, or TIFF URL.
The bound script has no API key. Add the property in [Step
3](#step-3-store-your-api-key). Script properties aren't copied with a
spreadsheet, so add it again in the copied sheet's script project.
Wait about 20 seconds, then change the refresh value. The next calculation
resumes the existing job. Apps Script stops custom functions after 30
seconds, so a slow parse can require more than one check.
Another calculation briefly held the script lock. Change the refresh value
after a few seconds. This reuses the existing job when one is available.
Temporary HTTP errors usually resolve on a later refresh. For an account or
credit error, check the Unsiloed dashboard before retrying.
Clear the cells below and to the right of the formula. Sheets doesn't let a
spilled array overwrite existing values.
Some spreadsheet locales use semicolons between arguments. Enter
`=UNSILOED_CHART(A2; B2)` instead.
Chart extraction reconstructs values from pixels. Small labels, overlapping
marks, unlabeled ticks, and compressed images can reduce fidelity. Use the
highest-resolution source available and verify decision-critical values
against the original image.
## Choose Between a Formula and a Menu Command
Use this custom function when one public chart image should produce one dataset
at the formula location. Use the [menu-driven table extraction
guide](/docs/cookbooks/spreadsheets/google-sheets-extract-table) when you need to read
an image stored inside a cell, create a new worksheet, or process several source
images in one run.
See how Unsiloed evaluates chart accuracy and point density.
Run a menu command that reads an image and writes a new worksheet.
Review the parse workflow and the document structures it returns.
Connect a spreadsheet to Unsiloed and compare workflow patterns.
# Extract a Table from an Image in Google Sheets
Source: https://docs.unsiloed.ai/docs/cookbooks/spreadsheets/google-sheets-extract-table
Extract a table from an image in Google Sheets and write it to a second sheet with Apps Script and the Unsiloed API.
Everything runs inside Google Sheets. After adding one Apps Script file and an
API key, you can extract tables with a spreadsheet menu command.
## Why Extract Tables in the Sheet
A table might arrive as a screenshot, a photo, or an image embedded in a report.
Retyping it is slow and introduces errors. In this guide, we'll add a menu command
that extracts the table from a selected image and writes it to a second sheet.
The script uses [`/parse`](/docs/document-processing/parsing/parsing) to identify the
document structure, including tables. It doesn't need a predefined schema for
each table layout.
## What We'll Build
A spreadsheet-bound Apps Script that:
1. Reads the image out of the selected cell.
2. Sends it to `/parse` and waits for the job to finish.
3. Finds the first table in the result and turns its HTML into rows.
4. Writes those rows to a new results sheet.
Build the script in [Step 2](#step-2-write-the-script), or copy the completed
version below.
Paste this into `Code.gs`, replacing its contents. The script reads the API key
from a script property configured in [Step 3](#step-3-store-your-api-key).
```javascript Code.gs theme={null}
const BASE_URL = "https://prod.visionapi.unsiloed.ai";
function unsiloedApiKey() {
const apiKey = PropertiesService.getScriptProperties()
.getProperty("UNSILOED_API_KEY");
if (!apiKey) throw new Error("Add UNSILOED_API_KEY to the script properties.");
return apiKey;
}
function onOpen() {
SpreadsheetApp.getUi().createMenu("Unsiloed")
.addItem("Extract Table From Image", "extractTable")
.addToUi();
}
function extractTable() {
const image = SpreadsheetApp.getActiveSheet().getActiveCell().getValue();
if (!image || image.valueType !== SpreadsheetApp.ValueType.IMAGE) {
throw new Error("Select a cell holding an image placed in the cell, not over the grid.");
}
const rows = tableRows(parseImage(imageBlob(image)));
const spreadsheet = SpreadsheetApp.getActiveSpreadsheet();
const sheetName = spreadsheet.getSheetByName("Extracted")
? "Extracted " + Date.now()
: "Extracted";
const sheet = spreadsheet.insertSheet(sheetName);
sheet.getRange(1, 1, rows.length, rows[0].length).setValues(rows);
}
// The cell holds a link to the image, not the bytes themselves.
function imageBlob(image) {
const blob = UrlFetchApp.fetch(image.getContentUrl()).getBlob();
const extensions = {
"image/png": "png",
"image/jpeg": "jpg",
"image/tiff": "tiff"
};
const extension = extensions[blob.getContentType()];
if (!extension) throw new Error("Use a PNG, JPEG, or TIFF image.");
return blob.setName("table." + extension);
}
function parseImage(blob) {
const headers = { "api-key": unsiloedApiKey() };
const started = UrlFetchApp.fetch(BASE_URL + "/parse", {
method: "post",
headers,
payload: { file: blob }
});
const jobId = JSON.parse(started.getContentText()).job_id;
for (let attempt = 0; attempt < 100; attempt++) {
Utilities.sleep(3000);
const response = UrlFetchApp.fetch(BASE_URL + "/parse/" + jobId, {
headers
});
const job = JSON.parse(response.getContentText());
if (job.status === "Succeeded") return job;
if (job.status === "Failed" || job.status === "Cancelled") {
throw new Error(job.message || "Parsing " + job.status.toLowerCase() + ".");
}
}
throw new Error("Parsing took too long.");
}
// No schema needed: /parse returns whatever table the image happens to hold.
function tableRows(job) {
const table = job.chunks
.flatMap(chunk => chunk.segments)
.find(segment => segment.segment_type === "Table");
if (!table) throw new Error("No table found in that image.");
if (!table.html) throw new Error("The detected table contained no HTML.");
const htmlRows = table.html.match(/
/gi);
if (!htmlRows) throw new Error("The detected table contained no rows.");
const rows = htmlRows.map(tr =>
(tr.match(//gi) || []).map(safeCellText));
if (!rows.some(row => row.length)) {
throw new Error("The detected table contained no cells.");
}
// setValues needs every row the same length, but spanning cells can make rows short.
const width = Math.max(...rows.map(row => row.length));
return rows.map(row => row.concat(Array(width - row.length).fill("")));
}
function safeCellText(cell) {
const text = cell
.replace(/ /gi, "\n")
.replace(/<[^>]+>/g, "")
.replace(/ /gi, " ")
.replace(/"/gi, '"')
.replace(/'|'/gi, "'")
.replace(/</gi, "<")
.replace(/>/gi, ">")
.replace(/(\d+);/g, (_, code) => String.fromCodePoint(Number(code)))
.replace(/([0-9a-f]+);/gi, (_, code) => String.fromCodePoint(parseInt(code, 16)))
.replace(/&/gi, "&")
.trim();
// Prevent document text from becoming a spreadsheet formula.
return /^[=+\-@]/.test(text) ? "'" + text : text;
}
```
## Requirements for Extracting Tables in Google Sheets
Before you start, gather:
* A Google account and a spreadsheet you can edit.
* An Unsiloed API key from the [dashboard](https://app.unsiloed.ai).
* A PNG, JPEG, or TIFF image containing a table. This guide uses a
[fund performance page](https://www.unsiloed.ai/docs/images/fund-performance-original.png)
so you can follow along with the same numbers.
## Step 1: Put the Image in a Cell
Google Sheets can hold an image two ways. **Insert image in cell** makes the image
a cell value. **Insert image over cells** leaves it floating above the grid.
Pasting with Ctrl+V creates the floating kind.
The script can read only an in-cell image. The `OverGridImage` class used for
floating images doesn't expose the image bytes.
Select the cell you want the image in, then choose **Insert → Image → Insert
image in cell**.
Paste the image URL, or upload a file, and click **Insert**.
The image should sit within the cell borders and resize with the row and
column.
If it floats above the grid, open its three-dot menu and choose **Put image in
cell**.
## Step 2: Write the Script
We'll add the script in sections. If you copied the completed version above, use
this section as a reference.
### 2.1 Open the Script Editor
From the spreadsheet, choose **Extensions → Apps Script**. This creates a script
bound to this spreadsheet and opens it in a new tab.
The editor opens `Code.gs` with an empty `myFunction`. Delete it, then add each
code block below.
### 2.2 Configuration and the Menu
In `Code.gs`, add the API configuration and spreadsheet menu:
```javascript Code.gs theme={null}
const BASE_URL = "https://prod.visionapi.unsiloed.ai";
function unsiloedApiKey() {
const apiKey = PropertiesService.getScriptProperties()
.getProperty("UNSILOED_API_KEY");
if (!apiKey) throw new Error("Add UNSILOED_API_KEY to the script properties.");
return apiKey;
}
function onOpen() {
SpreadsheetApp.getUi().createMenu("Unsiloed")
.addItem("Extract Table From Image", "extractTable")
.addToUi();
}
```
Apps Script runs the reserved `onOpen` function when the spreadsheet opens, which
adds the **Unsiloed** menu. The `unsiloedApiKey` helper reads the current key from
a script property and reports a clear error if it is missing. [Step
3](#step-3-store-your-api-key) configures that property.
### 2.3 Read the Image Out of the Cell
Add `imageBlob` below `onOpen`. The `getValue()` method returns a `CellImage`, so
the function fetches its bytes from a Google-hosted URL:
```javascript Code.gs theme={null}
// The cell holds a link to the image, not the bytes themselves.
function imageBlob(image) {
const blob = UrlFetchApp.fetch(image.getContentUrl()).getBlob();
const extensions = {
"image/png": "png",
"image/jpeg": "jpg",
"image/tiff": "tiff"
};
const extension = extensions[blob.getContentType()];
if (!extension) throw new Error("Use a PNG, JPEG, or TIFF image.");
return blob.setName("table." + extension);
}
```
Keep the `setName` call. The `/parse` endpoint chooses a decoder from the file
extension, so an unnamed blob is rejected. The MIME-type lookup accepts the three
supported image formats and gives the blob a matching filename.
A blob without an extension still returns `200` and a job ID, but the job later
fails with `Unsupported file type`. Check the job status, not only the
submission response.
### 2.4 Send the Image to Parse
Add `parseImage` below `imageBlob`. Because `/parse` is asynchronous, the function
submits the image and polls the job endpoint until parsing succeeds or fails:
```javascript Code.gs theme={null}
function parseImage(blob) {
const headers = { "api-key": unsiloedApiKey() };
const started = UrlFetchApp.fetch(BASE_URL + "/parse", {
method: "post",
headers,
payload: { file: blob }
});
const jobId = JSON.parse(started.getContentText()).job_id;
for (let attempt = 0; attempt < 100; attempt++) {
Utilities.sleep(3000);
const response = UrlFetchApp.fetch(BASE_URL + "/parse/" + jobId, {
headers
});
const job = JSON.parse(response.getContentText());
if (job.status === "Succeeded") return job;
if (job.status === "Failed" || job.status === "Cancelled") {
throw new Error(job.message || "Parsing " + job.status.toLowerCase() + ".");
}
}
throw new Error("Parsing took too long.");
}
```
The `UrlFetchApp` service builds a multipart request from the blob in `payload`
and uses the blob name as its filename.
Menu-driven scripts can run for six minutes, while spreadsheet custom functions
stop after 30 seconds. In our tests, a single-page image usually finishes in
about 20 seconds, but processing time varies. A menu command leaves more room
for slower jobs.
### 2.5 Turn the Table Into Rows
Add `tableRows` below `parseImage`. It finds the first `Table` segment and converts
its `html` field into rows:
```javascript Code.gs theme={null}
// No schema needed: /parse returns whatever table the image happens to hold.
function tableRows(job) {
const table = job.chunks
.flatMap(chunk => chunk.segments)
.find(segment => segment.segment_type === "Table");
if (!table) throw new Error("No table found in that image.");
if (!table.html) throw new Error("The detected table contained no HTML.");
const htmlRows = table.html.match(/
/gi);
if (!htmlRows) throw new Error("The detected table contained no rows.");
const rows = htmlRows.map(tr =>
(tr.match(//gi) || []).map(safeCellText));
if (!rows.some(row => row.length)) {
throw new Error("The detected table contained no cells.");
}
// setValues needs every row the same length, but spanning cells can make rows short.
const width = Math.max(...rows.map(row => row.length));
return rows.map(row => row.concat(Array(width - row.length).fill("")));
}
function safeCellText(cell) {
const text = cell
.replace(/ /gi, "\n")
.replace(/<[^>]+>/g, "")
.replace(/ /gi, " ")
.replace(/"/gi, '"')
.replace(/'|'/gi, "'")
.replace(/</gi, "<")
.replace(/>/gi, ">")
.replace(/(\d+);/g, (_, code) => String.fromCodePoint(Number(code)))
.replace(/([0-9a-f]+);/gi, (_, code) => String.fromCodePoint(parseInt(code, 16)))
.replace(/&/gi, "&")
.trim();
// Prevent document text from becoming a spreadsheet formula.
return /^[=+\-@]/.test(text) ? "'" + text : text;
}
```
Two details matter here:
* The parser validates that the result contains rows and cells, preserves line
breaks, and decodes common named and numeric HTML entities.
* Formula-leading text gets an apostrophe prefix before `setValues`, preventing
extracted document content from executing as a spreadsheet formula.
* Spanning cells can make some rows shorter. Because `setValues` rejects uneven
rows, the function pads missing cells at the end. It doesn't reproduce complex
`rowspan` or middle-column `colspan` layouts.
`find` takes the first table on the page. See
[Extend the Google Sheets Integration](#extend-the-google-sheets-integration) for
handling several.
### 2.6 Write the Rows to the Sheet
Add `extractTable` below `tableRows`. This menu handler validates the selection,
runs the helpers, and writes the result:
```javascript Code.gs theme={null}
function extractTable() {
const image = SpreadsheetApp.getActiveSheet().getActiveCell().getValue();
if (!image || image.valueType !== SpreadsheetApp.ValueType.IMAGE) {
throw new Error("Select a cell holding an image placed in the cell, not over the grid.");
}
const rows = tableRows(parseImage(imageBlob(image)));
const spreadsheet = SpreadsheetApp.getActiveSpreadsheet();
const sheetName = spreadsheet.getSheetByName("Extracted")
? "Extracted " + Date.now()
: "Extracted";
const sheet = spreadsheet.insertSheet(sheetName);
sheet.getRange(1, 1, rows.length, rows[0].length).setValues(rows);
}
```
The `valueType` check accepts only in-cell images. Each run creates a new sheet,
using a timestamp in the name when `Extracted` already exists, so it never clears
an earlier result. A single `setValues` call writes the full table without making
a separate Sheets request for every cell.
Save `Code.gs`.
## Step 3: Store Your API Key
Store the API key outside the source code so it isn't included when you share or
copy the script.
In the Apps Script editor, open **Project Settings** from the left sidebar and
scroll to **Script Properties**. Click **Add script property**. If the project
already has properties, click **Edit script properties** first.
Name it `UNSILOED_API_KEY` and paste your key as the value.
Click **Save script properties**.
Spreadsheet editors can view and modify its bound Apps Script project. Use a
dedicated API key, restrict edit access to trusted collaborators, and rotate
the key if the spreadsheet is shared unexpectedly.
## Step 4: Authorize the Script
The first run asks for permission to edit the spreadsheet and call the Unsiloed
API.
Back in the spreadsheet, reload the page. A new **Unsiloed** menu appears in
the menu bar. Select the cell holding your image and choose **Unsiloed →
Extract Table From Image**.
Google Sheets shows **Authorization required**.
Click **OK**, then pick your Google account in the window that opens.
Google warns that it hasn't verified the app because this is your unpublished
script. Click **Advanced**, then **Go to \ (unsafe)**.
Review the two permissions it asks for and click **Allow**:
* Spreadsheet access to read the image and write the rows
* External service access to call the Unsiloed API
## Step 5: Extract the Table
With the script authorized, the menu command is the whole workflow. Select the
cell holding the image and choose **Unsiloed → Extract Table From Image**.
In our tests, the script writes the rows in about 20 seconds for a single-page
image. Processing time varies with the image and current service load.
### Sample Output
For the sample fund performance page, the `Extracted` sheet preserves the group
rows and the em dash for the missing five-year return:
## What to Expect From the Output
The rows preserve the table text rather than normalizing it:
* **Values arrive as text.** `7.76%` keeps its percent sign, and an em dash remains
an em dash. Formula-leading values are escaped before writing. Convert the
values in Sheets if you need numbers.
* **Layout rows come through.** A grouping row like `Class A Shares` arrives as a
label with empty cells beside it, exactly as it sits on the page.
* **Row grouping can vary.** A label and its values might share a row or split
across two rows. Don't write formulas that assume a fixed row offset.
If you need typed values and a fixed set of columns, use
[`/v2/extract`](/docs/document-processing/extraction/extraction) with a schema instead.
This requires a schema but returns typed values in predictable columns.
## Troubleshoot the Script
The selected cell has no in-cell image. Either the wrong cell is selected, or
the image is floating over the grid rather than in a cell. See
[Step 1](#step-1-put-the-image-in-a-cell).
`/parse` found no table. Try a larger or sharper image, especially if the
layout has no ruling lines or consistent columns.
Use a PNG, JPEG, or TIFF image. The script rejects other MIME types before
submission and gives supported blobs a matching file extension.
Reload the spreadsheet to run `onOpen` and rebuild the menu. If it still shows
old items, run `onOpen` once from the Apps Script editor.
The API key is missing or wrong. Check the `UNSILOED_API_KEY` script property,
then run the command again.
## Extend the Google Sheets Integration
The script takes the first table it finds. To pull every table out of a
multi-table page, collect all the `Table` segments instead of calling `find`, and
write each one below the last.
To process many images, place one image per row, loop over the range, and write
each result to its own sheet. Keep enough margin below the six-minute Apps Script
limit for slower jobs and retries. For larger batches, submit every job first,
then poll them together so they run concurrently.
What `/parse` returns for a document, segment by segment.
The full list of segment types, including `Table`.
Pull typed fields against a schema when you need numbers rather than text.
The full request and response specs for `/parse`.
# Classification Overview
Source: https://docs.unsiloed.ai/docs/document-processing/classification/classification
Route documents to predefined categories with a confidence score on every page.
Classification answers a narrow question: what kind of document is this? The `/classify` endpoint takes a document (PDF or image) and a list of candidate categories, then returns per-page predictions, an overall classification for the document, and a confidence score on every value.
Use it at the front of an ingestion pipeline when you're processing mixed batches and need to route each document (or each page) to the right downstream handler. For splitting a single mixed bundle into separate files, see [Splitting](/docs/document-processing/splitting/splitting).
## How It Works
After you submit a document and categories, the API:
1. Analyzes every page of the document.
2. Scores each page against your candidate categories using a vision-language model.
3. Returns the per-page classification, the overall document classification, and confidence scores throughout.
## Common Use Cases
* **Healthcare:** sort medical records, lab reports, prescriptions, and insurance forms
* **Business:** categorize invoices, receipts, purchase orders, and contracts
* **Financial:** sort bank statements, tax forms, and financial reports
* **Legal:** identify contracts, agreements, legal notices, and compliance forms
## Dig Deeper
Submit a document, define categories, and read back classifications with confidence scores.
Browse the canonical classification response with examples for each job state.
For the full request and response specification, see the [Classify API reference](/docs/api-reference/classification/classify-document).
# Getting Started With Classification
Source: https://docs.unsiloed.ai/docs/document-processing/classification/quickstart
Submit a document and a list of candidate categories to /classify, then read back the predicted category.
Classification picks the best-fit label for a document from a list of candidate categories we supply. The endpoint returns the matched category and a confidence score, ready to feed into routing logic. For raw Markdown or structured field extraction instead, see the [Parse quickstart](/docs/quickstart) or the [Extraction quickstart](/docs/document-processing/extraction/quickstart).
The walkthrough below builds a script that submits a PDF and our candidate categories to `/classify`, waits for the verdict, and saves the matched category and confidence score to disk. The accordion below has the full script if you'd rather copy and run it directly.
Set `UNSILOED_API_KEY` in your environment and save the document you want to classify as `document.pdf` in the same directory before running.
```python classify_document.py theme={null}
import json
import os
import time
import requests
API_KEY = os.environ["UNSILOED_API_KEY"]
BASE_URL = "https://prod.visionapi.unsiloed.ai"
categories = [
{"name": "Sales Report"},
{"name": "Invoice"},
{"name": "Medical Record"},
]
with open("document.pdf", "rb") as f:
response = requests.post(
f"{BASE_URL}/classify",
headers={"api-key": API_KEY},
files={"pdf_file": ("document.pdf", f, "application/pdf")},
data={"categories": json.dumps(categories)},
)
response.raise_for_status()
job_id = response.json()["job_id"]
print(f"Job submitted: {job_id}")
max_attempts = 60 # roughly 5 minutes at 5 seconds per poll
attempts = 0
while True:
result = requests.get(
f"{BASE_URL}/classify/{job_id}",
headers={"api-key": API_KEY},
).json()
print(f"Status: {result['status']}")
if result["status"] == "completed":
break
if result["status"] == "failed":
raise RuntimeError(result.get("error", "classify job failed"))
attempts += 1
if attempts >= max_attempts:
raise TimeoutError("Classify job did not finish within 5 minutes")
time.sleep(5)
with open("classification.json", "w") as f:
json.dump(result, f, indent=2)
classification = result["result"]
print(f"Classification: {classification['classification']} ({classification['confidence']:.2%} confidence)")
```
Save this as `script.mjs` or set `"type": "module"` in your `package.json`. Requires Node.js 18 or newer for the global `fetch`, `FormData`, and `Blob`.
```javascript script.mjs theme={null}
import fs from "node:fs";
const API_KEY = process.env.UNSILOED_API_KEY;
const BASE_URL = "https://prod.visionapi.unsiloed.ai";
const categories = [
{ name: "Sales Report", description: "Sales performance summaries with regional or quarterly data" },
{ name: "Invoice", description: "Bill of sale with line items" },
{ name: "Medical Record", description: "Patient health records" },
];
const form = new FormData();
form.append("pdf_file", new Blob([fs.readFileSync("document.pdf")]), "document.pdf");
form.append("categories", JSON.stringify(categories));
const response = await fetch(`${BASE_URL}/classify`, {
method: "POST",
headers: { "api-key": API_KEY },
body: form,
});
if (!response.ok) throw new Error(`${response.status}: ${await response.text()}`);
const { job_id } = await response.json();
console.log(`Job submitted: ${job_id}`);
const maxAttempts = 60; // roughly 5 minutes at 5 seconds per poll
let attempts = 0;
let result;
while (true) {
const res = await fetch(`${BASE_URL}/classify/${job_id}`, {
headers: { "api-key": API_KEY },
});
result = await res.json();
console.log(`Status: ${result.status}`);
if (result.status === "completed") break;
if (result.status === "failed") throw new Error(result.error || "classify job failed");
if (++attempts >= maxAttempts) throw new Error("Classify job did not finish within 5 minutes");
await new Promise((r) => setTimeout(r, 5000));
}
fs.writeFileSync("classification.json", JSON.stringify(result, null, 2));
const { classification, confidence } = result.result;
console.log(`Classification: ${classification} (${(confidence * 100).toFixed(2)}% confidence)`);
```
```bash theme={null}
# Submit the document and capture the job_id from the response:
resp=$(curl -sX POST "https://prod.visionapi.unsiloed.ai/classify" \
-H "api-key: $UNSILOED_API_KEY" \
-F "pdf_file=@document.pdf" \
-F 'categories=[{"name":"Sales Report"},{"name":"Invoice"},{"name":"Medical Record"}]')
JOB_ID=$(echo "$resp" | grep -o '"job_id":"[^"]*"' | cut -d'"' -f4)
echo "Job submitted: $JOB_ID"
# Poll until the job finishes, with a 5-minute timeout:
attempts=0
max_attempts=60
while true; do
resp=$(curl -sX GET "https://prod.visionapi.unsiloed.ai/classify/$JOB_ID" \
-H "api-key: $UNSILOED_API_KEY")
status=$(echo "$resp" | grep -o '"status":"[^"]*"' | head -1 | cut -d'"' -f4)
echo "Status: $status"
[ "$status" = "completed" ] && break
[ "$status" = "failed" ] && { echo "Job failed"; exit 1; }
attempts=$((attempts + 1))
[ "$attempts" -ge "$max_attempts" ] && { echo "Classify job did not finish within 5 minutes"; exit 1; }
sleep 5
done
# Save the full response to disk:
echo "$resp" > classification.json
```
## Step 1: Set Up Your Environment
Before writing any code, gather three things: an API key, a document, and the runtime for the chosen language.
### 1.1 Get an Unsiloed AI API Key
To get API access, [sign up on Unsiloed AI](https://cal.com/aman-mishra-p0ry57/15min). Export your key as an environment variable named `UNSILOED_API_KEY` so it stays out of source control:
```bash theme={null}
export UNSILOED_API_KEY="your-api-key"
```
### 1.2 Pick a Document to Classify
The `/classify` endpoint supports PDF, DOCX, PPTX, JPG, PNG, and other formats. The walkthrough below assumes a PDF saved as `document.pdf` in your working directory. To use a different format, update the filename and content type in the snippets to match your file.
If you don't have a document handy, download our [sample PDF](https://raw.githubusercontent.com/Unsiloed-AI/cookbook/c585446e46e4be2790c6c29fe2a7a3a1b346191d/sample-documents/sample-classify.pdf) (a one-page lab report from Riverside Diagnostic Laboratory) and save it as `document.pdf`. The walkthrough scores it against three candidate categories so we can see a clear winner.
### 1.3 Install Dependencies
You need Python 3.8 or newer. Install the `requests` package:
```bash theme={null}
pip install requests
```
You need Node.js 18 or newer for the global `fetch`, `FormData`, and `Blob`. No external packages needed.
You need cURL, which is preinstalled on macOS and most Linux distributions. No external packages needed.
## Step 2: Submit a Document With Categories
The request bundles two fields: `pdf_file` for the document and `categories` for a JSON-stringified array of category objects, each with a `name` and an optional `description`. The categories list is the model's entire vocabulary for this call, so clear and distinct names matter more than they might appear. The endpoint returns a `job_id` to poll. All requests go to `https://prod.visionapi.unsiloed.ai` with the API key in the `api-key` header.
### 2.1 Set Up the Script
Create a file called `classify_document.py` and start with the imports, configuration, and category list:
```python classify_document.py theme={null}
import json
import os
import time
import requests
API_KEY = os.environ["UNSILOED_API_KEY"]
BASE_URL = "https://prod.visionapi.unsiloed.ai"
categories = [
{"name": "Sales Report"},
{"name": "Invoice"},
{"name": "Medical Record"},
]
```
`API_KEY` reads your key from the environment so it doesn't get hard-coded into the file, and `BASE_URL` points at the Unsiloed AI production endpoint. The `categories` list defines the candidate labels the model picks from. Only the names guide the result; a `description` key is accepted but not used by classification.
Create a file called `script.mjs` and start with the imports, configuration, and category list:
```javascript script.mjs theme={null}
import fs from "node:fs";
const API_KEY = process.env.UNSILOED_API_KEY;
const BASE_URL = "https://prod.visionapi.unsiloed.ai";
const categories = [
{ name: "Sales Report", description: "Sales performance summaries with regional or quarterly data" },
{ name: "Invoice", description: "Bill of sale with line items" },
{ name: "Medical Record", description: "Patient health records" },
];
```
`API_KEY` reads your key from the environment so it doesn't get hard-coded into the file, and `BASE_URL` points at the Unsiloed AI production endpoint. The `categories` list defines the candidate labels the model picks from. Only the names guide the result; a `description` key is accepted but not used by classification.
cURL doesn't need a setup step. Each command below inlines the API key, base URL, and category list directly.
### 2.2 Upload the Document
Send the file and the JSON-encoded category list as a multipart upload to `/classify`. The document goes under `pdf_file` and the categories under `categories`.
Continue the file by uploading the document:
```python classify_document.py theme={null}
with open("document.pdf", "rb") as f:
response = requests.post(
f"{BASE_URL}/classify",
headers={"api-key": API_KEY},
files={"pdf_file": ("document.pdf", f, "application/pdf")},
data={"categories": json.dumps(categories)},
)
response.raise_for_status()
```
`raise_for_status()` throws an `HTTPError` on any non-2xx response, so there's no need to check `.status_code` separately.
Continue the file by uploading the document:
```javascript script.mjs theme={null}
const form = new FormData();
form.append("pdf_file", new Blob([fs.readFileSync("document.pdf")]), "document.pdf");
form.append("categories", JSON.stringify(categories));
const response = await fetch(`${BASE_URL}/classify`, {
method: "POST",
headers: { "api-key": API_KEY },
body: form,
});
if (!response.ok) throw new Error(`${response.status}: ${await response.text()}`);
```
`fetch` doesn't throw on non-2xx responses by default, so we check `response.ok` and throw the error explicitly.
Run:
```bash theme={null}
curl -X POST "https://prod.visionapi.unsiloed.ai/classify" \
-H "api-key: $UNSILOED_API_KEY" \
-F "pdf_file=@document.pdf" \
-F 'categories=[{"name":"Sales Report"},{"name":"Invoice"},{"name":"Medical Record"}]'
```
The response prints to stdout. We need the `job_id` field for the next step.
### 2.3 Capture the Job ID
Next, read and print the `job_id`:
```python classify_document.py theme={null}
job_id = response.json()["job_id"]
print(f"Job submitted: {job_id}")
```
Run the script:
```bash theme={null}
python classify_document.py
```
The output should be a single line like `Job submitted: 2c231adf-ad5e-4e2e-8c0c-10cd7025c09b`.
Next, read and log the `job_id`:
```javascript script.mjs theme={null}
const { job_id } = await response.json();
console.log(`Job submitted: ${job_id}`);
```
Run the script:
```bash theme={null}
node script.mjs
```
The output should be a single line like `Job submitted: 2c231adf-ad5e-4e2e-8c0c-10cd7025c09b`.
The response body from the POST above looks like:
```json theme={null}
{
"job_id": "2c231adf-ad5e-4e2e-8c0c-10cd7025c09b",
"status": "processing",
"message": "Classification started",
"quota_remaining": 7704
}
```
Copy the `job_id` value to paste into the polling command in the next step.
## Step 3: Poll for Results
The job runs asynchronously. GET `/classify/{job_id}` repeatedly until the status is `completed`, then save the classification to disk.
A status of `completed` means the result is ready. A status of `failed` means the job errored. Any other value (such as `processing`) means the job is still running.
### 3.1 Write the Polling Loop
Drop in a polling loop. The `max_attempts` cap stops the loop if the job hangs:
```python classify_document.py theme={null}
max_attempts = 60 # roughly 5 minutes at 5 seconds per poll
attempts = 0
while True:
result = requests.get(
f"{BASE_URL}/classify/{job_id}",
headers={"api-key": API_KEY},
).json()
print(f"Status: {result['status']}")
if result["status"] == "completed":
break
if result["status"] == "failed":
raise RuntimeError(result.get("error", "classify job failed"))
attempts += 1
if attempts >= max_attempts:
raise TimeoutError("Classify job did not finish within 5 minutes")
time.sleep(5)
```
Drop in a polling loop. The `maxAttempts` cap stops the loop if the job hangs:
```javascript script.mjs theme={null}
const maxAttempts = 60; // roughly 5 minutes at 5 seconds per poll
let attempts = 0;
let result;
while (true) {
const res = await fetch(`${BASE_URL}/classify/${job_id}`, {
headers: { "api-key": API_KEY },
});
result = await res.json();
console.log(`Status: ${result.status}`);
if (result.status === "completed") break;
if (result.status === "failed") throw new Error(result.error || "classify job failed");
if (++attempts >= maxAttempts) throw new Error("Classify job did not finish within 5 minutes");
await new Promise((r) => setTimeout(r, 5000));
}
```
Replace `JOB_ID` below with the value you captured from Step 2.3, then run this loop. It polls every 5 seconds and gives up after 5 minutes if the job hasn't completed:
```bash theme={null}
JOB_ID="paste-job-id-here"
attempts=0
max_attempts=60 # roughly 5 minutes at 5 seconds per poll
while true; do
resp=$(curl -sX GET "https://prod.visionapi.unsiloed.ai/classify/$JOB_ID" \
-H "api-key: $UNSILOED_API_KEY")
status=$(echo "$resp" | grep -o '"status":"[^"]*"' | head -1 | cut -d'"' -f4)
echo "Status: $status"
[ "$status" = "completed" ] && break
[ "$status" = "failed" ] && { echo "Job failed"; exit 1; }
attempts=$((attempts + 1))
[ "$attempts" -ge "$max_attempts" ] && { echo "Classify job did not finish within 5 minutes"; exit 1; }
sleep 5
done
```
The loop keeps the latest response body in `$resp` for the next step.
### 3.2 Save the Classification
Persist the result to disk so downstream code can read it. The full response, including the per-page breakdown, goes to `classification.json`.
Finally, write the result to disk and print a summary:
```python classify_document.py theme={null}
with open("classification.json", "w") as f:
json.dump(result, f, indent=2)
classification = result["result"]
print(f"Classification: {classification['classification']} ({classification['confidence']:.2%} confidence)")
```
Run the script:
```bash theme={null}
python classify_document.py
```
You should see one or two `Status: processing` lines, then `Status: completed`, then a summary line like `Classification: Medical Record (100.00% confidence)`. The `classification.json` file appears in the working directory.
Finally, write the result to disk and log a summary:
```javascript script.mjs theme={null}
fs.writeFileSync("classification.json", JSON.stringify(result, null, 2));
const { classification, confidence } = result.result;
console.log(`Classification: ${classification} (${(confidence * 100).toFixed(2)}% confidence)`);
```
Run the script:
```bash theme={null}
node script.mjs
```
You should see one or two `Status: processing` lines, then `Status: completed`, then a summary line like `Classification: Medical Record (100.00% confidence)`. The `classification.json` file appears in the working directory.
The polling loop in Step 3.1 left the full response in `$resp`. Write it to disk:
```bash theme={null}
echo "$resp" > classification.json
```
The `classification.json` file now holds the full response. The overall label lives under `result.classification` and the per-page breakdown under `result.page_results`.
## Error Responses
Failures fall into two buckets: HTTP errors raised before the job is queued, and a `failed` status on a job that started but could not complete.
### HTTP Errors
The `/classify` endpoint returns JSON error bodies under a `detail` field. The common cases are:
* **`401 Unauthorized`:** `{"detail":"Invalid API key"}`. The `api-key` header is missing or wrong.
* **`400 Bad Request`:** `{"detail":"Either pdf_file or file_url must be provided"}` or `{"detail":"At least one category is required"}`. The submit form is missing a required field.
* **`422 Unprocessable Entity`:** `{"detail":[{"type":"missing","loc":["body","categories"],"msg":"Field required","input":null}]}`. A required form field, usually `categories`, is missing entirely.
* **`404 Not Found`:** `{"detail":"Job not found"}`. The `job_id` you polled doesn't exist.
### Failed Jobs
A job that was accepted but could not be processed returns `status: "failed"` on the polling endpoint. The response shape matches a successful one, but `result` is absent and the `error` field describes what went wrong:
```json theme={null}
{
"job_id": "660e8400-e29b-41d4-a716-446655440001",
"status": "failed",
"progress": "Classification failed",
"error": "Invalid PDF format"
}
```
## Response Shape
A completed job returns job metadata plus a nested `result` object that contains the overall classification, a confidence score, and per-page results.
```json theme={null}
{
"job_id": "2c231adf-ad5e-4e2e-8c0c-10cd7025c09b",
"status": "completed",
"progress": "Classification completed",
"error": null,
"result": {
"success": true,
"classification": "Medical Record",
"confidence": 1.0,
"total_pages": 1,
"processed_pages": 1,
"page_results": [
{
"page": 1,
"success": true,
"classification": "Medical Record",
"raw_result": "Medical Record",
"confidence": 1.0
}
]
}
}
```
The fields you use depend on what you're building. They fall into three broad categories:
**For routing decisions:**
* **`result.classification`:** the overall predicted category for the document, drawn from the `name` values you submitted. This is the field the walkthrough prints.
* **`result.confidence`:** confidence score for the overall classification, on a 0-1 scale. Treat it as a soft signal: high values rarely need review, low values flag documents worth a human look.
**For per-page handling and mixed-content documents:**
* **`result.page_results[]`:** the per-page classifications the overall result is built from
* **`page_results[].page`:** 1-indexed page number
* **`page_results[].classification`:** the predicted category for that page
* **`page_results[].raw_result`:** the model's raw output before normalization to a category name; usually identical to `classification`
* **`page_results[].confidence`:** the page-level confidence score on a 0-1 scale
**For job tracking:**
* **`status`:** `completed`, `failed`, or an in-progress value such as `processing`
* **`progress`:** human-readable progress message
* **`error`:** error message if the job failed, otherwise `null`
* **`result.total_pages` and `result.processed_pages`:** how much of the document the classifier got through
### Sample Output
Running the script against the [sample lab report](https://raw.githubusercontent.com/Unsiloed-AI/cookbook/c585446e46e4be2790c6c29fe2a7a3a1b346191d/sample-documents/sample-classify.pdf) and the three categories above writes the verdict to `classification.json`:
```json theme={null}
{
"job_id": "2c231adf-ad5e-4e2e-8c0c-10cd7025c09b",
"status": "completed",
"progress": "Classification completed",
"error": null,
"result": {
"success": true,
"classification": "Medical Record",
"confidence": 1.0,
"total_pages": 1,
"processed_pages": 1,
"page_results": [
{
"page": 1,
"success": true,
"classification": "Medical Record",
"raw_result": "Medical Record",
"confidence": 1.0
}
]
}
}
```
The Riverside Diagnostic lab report lands cleanly in the `Medical Record` bucket with full confidence. Swap in your own document and category list to see how the classifier handles ambiguous cases.
## Next Steps
For more on classification, including category design tips and the canonical response shape, see the [Classification overview](/docs/document-processing/classification/classification).
Understand how the classifier scores pages and when to reach for it.
Browse the full classification response with examples for each job state.
Browse the full request and response specs for the classify endpoint.
Split a mixed bundle into separate documents by section.
# Response Format
Source: https://docs.unsiloed.ai/docs/document-processing/classification/response-format
Response shape from the /classify endpoint, with examples for each job state.
A completed classification job returns the overall document classification, a confidence score, and per-page results. The example below is a real response from classifying an invoice against four candidate categories.
```json theme={null}
{
"job_id": "97ba5215-2367-47f8-9298-da43c0722d20",
"status": "completed",
"progress": "Classification completed",
"error": null,
"result": {
"success": true,
"classification": "Invoice",
"confidence": 0.9999994487765023,
"total_pages": 1,
"processed_pages": 1,
"page_results": [
{
"page": 1,
"success": true,
"classification": "Invoice",
"raw_result": "Invoice",
"confidence": 0.9999994487765023
}
]
}
}
```
## Top-Level Fields
* **`job_id`:** unique identifier for the classification job
* **`status`:** job state (`completed`, `failed`, or an in-progress value such as `processing`)
* **`progress`:** human-readable progress message
* **`error`:** error message if the job failed, otherwise `null`
* **`result`:** the classification result object (described below)
## Result Object
* **`success`:** whether the classification operation succeeded
* **`classification`:** the overall predicted category for the document
* **`confidence`:** confidence score for the overall classification (0–1)
* **`total_pages`:** total number of pages in the document
* **`processed_pages`:** number of pages successfully processed
* **`page_results`:** per-page classifications (described below)
## Page Result Object
Each item in `page_results` describes the classification of one page:
* **`page`:** page number, 1-indexed
* **`success`:** whether classification succeeded for that page
* **`classification`:** the predicted category for the page
* **`raw_result`:** the model's raw output before normalization to a category name; usually identical to `classification`
* **`confidence`:** confidence score for the page's classification (0–1)
## Other Job States
The submit response (before polling kicks in) returns a smaller body:
```json theme={null}
{
"job_id": "660e8400-e29b-41d4-a716-446655440001",
"status": "processing",
"message": "Classification started",
"quota_remaining": 400
}
```
A job that's still running returns the current progress without a `result`:
```json theme={null}
{
"job_id": "660e8400-e29b-41d4-a716-446655440001",
"status": "processing",
"progress": "Analyzing document...",
"error": null
}
```
A failed job returns an `error` message:
```json theme={null}
{
"job_id": "660e8400-e29b-41d4-a716-446655440001",
"status": "failed",
"progress": "Classification failed",
"error": "Invalid PDF format"
}
```
For the full request and response specification, see the [Classify API reference](/docs/api-reference/classification/classify-document).
# Extract Overview
Source: https://docs.unsiloed.ai/docs/document-processing/extraction/extraction
Pull typed fields out of a document using a JSON schema, with a confidence score on every value.
Where [parsing](/docs/document-processing/parsing/parsing) returns the whole document broken into Markdown chunks, extraction pulls only the specific fields you ask for. You hand the `/v2/extract` endpoint a JSON schema describing the fields you want, and the API returns them filled in with values.
Extraction is the right operation when you already know what matters: line items from an invoice, headline terms from a contract, fields from a lab report. Each value comes with a confidence score, so you can flag uncertain extractions for human review before they hit a database.
## From a Fund Report to Typed Fields
The mutual fund performance page below has a header, a growth-of-\$1,000,000 chart, and an average annual total return table broken out by share class. Extraction reads all of that and returns just the fields named in our schema, with a confidence score on every value. The rows of the returns table come back as structured objects of their own.
For this page we asked for the fund name, the returns as-of date, and every share-class row of the returns table. Here's what came back:
| Field | Extracted value | Confidence |
| - | - | - |
| `fund_name` | Calamos Market Neutral Income Fund | 99.5% |
| `returns_as_of_date` | 10/31/23 | 99.5% |
The four share-class rows came back inside the `share_classes` array, each as its own structured object with a confidence score on every value:
| Share class | Inception | 1-Year | 5-Year | 10-Year / Since Inception |
| - | - | - | - | - |
| Class A Shares | 9/4/90 | 7.76% | 3.25% | 3.28% |
| Class C Shares | 2/16/00 | 6.93 | 2.47 | 2.50 |
| Class I Shares | 5/10/00 | 8.07 | 3.52 | 3.53 |
| Class R6 Shares | 6/23/20 | 8.08 | — | 3.46 |
See the [Quickstart](/docs/document-processing/extraction/quickstart) to run an extraction end to end in Python, JavaScript, and cURL, or the [Schemas](/docs/document-processing/extraction/schemas) reference for the JSON Schema rules behind it.
## Defining a Schema
Extraction schemas follow the JSON Schema spec. The most important thing in the schema is the `description` you give each field. Keep descriptions as detailed and specific as possible. Clear, pointed descriptions help the model correctly locate and extract the intended information, especially in complex or ambiguous documents.
See the [Schemas](/docs/document-processing/extraction/schemas) reference for the rules, supported types, and worked examples.
## How Extraction Works
Once you submit a schema:
1. Unsiloed locates each named field in the document.
2. Extracts a value of the right type for each one.
3. Returns the values as structured JSON, each with a confidence score.
## Schema Tips
* Keep field names simple and descriptive.
* Use nested objects to reflect document structure.
* Avoid free-form schemas; strict schemas produce better results.
* Prefer arrays for repeated sections (line items, directors, transactions).
## Dig Deeper
JSON Schema rules, supported types, and worked examples for invoices and SEC filings.
The canonical extraction response with a field-by-field reference.
For the full request and response specification, see the [Extract API reference](/docs/api-reference/extraction/extract-data).
# Getting Started With Extract
Source: https://docs.unsiloed.ai/docs/document-processing/extraction/quickstart
Submit a document and a JSON schema to /v2/extract and read back typed fields with confidence scores.
Extraction pulls typed fields out of a document against a JSON schema we define, returning each leaf tagged with its own confidence score. For raw Markdown or just a category label instead, see the [Parse quickstart](/docs/quickstart) or the [Classification quickstart](/docs/document-processing/classification/quickstart).
By the end, we'll have a script that gives an invoice PDF and a schema to `/v2/extract`, waits for the job to finish, and writes the matched fields back as a clean JSON object with per-field confidence scores. Grab the full script from the dropdown below if you'd rather skip the walkthrough.
Set `UNSILOED_API_KEY` in your environment and save the document you want to extract from as `document.pdf` in the same directory before running.
```python extract_document.py theme={null}
import json
import os
import time
import requests
API_KEY = os.environ["UNSILOED_API_KEY"]
BASE_URL = "https://prod.visionapi.unsiloed.ai"
schema = {
"type": "object",
"properties": {
"vendor_name": {
"type": "string",
"description": "Name of the company issuing the invoice (the seller)",
},
"invoice_number": {
"type": "string",
"description": "Unique invoice identifier shown on the document",
},
"issue_date": {
"type": "string",
"description": "Date the invoice was issued",
},
"total_due": {
"type": "number",
"description": "Final total amount due in US dollars, including tax",
},
"line_items": {
"type": "array",
"description": "One row per line item in the invoice table",
"items": {
"type": "object",
"properties": {
"description": {"type": "string", "description": "Description of the product or service"},
"quantity": {"type": "number", "description": "Quantity of the item ordered"},
"unit_price": {"type": "number", "description": "Price per unit in US dollars"},
"subtotal": {"type": "number", "description": "Line subtotal in US dollars (quantity x unit_price)"},
},
"required": ["description", "quantity", "unit_price", "subtotal"],
"additionalProperties": False,
},
},
},
"required": ["vendor_name", "invoice_number", "issue_date", "total_due", "line_items"],
"additionalProperties": False,
}
with open("document.pdf", "rb") as f:
response = requests.post(
f"{BASE_URL}/v2/extract",
headers={"api-key": API_KEY},
files={"pdf_file": ("document.pdf", f, "application/pdf")},
data={"schema_data": json.dumps(schema)},
)
response.raise_for_status()
job_id = response.json()["job_id"]
print(f"Job submitted: {job_id}")
max_attempts = 60 # roughly 5 minutes at 5 seconds per poll
attempts = 0
while True:
result = requests.get(
f"{BASE_URL}/extract/{job_id}",
headers={"api-key": API_KEY},
).json()
print(f"Status: {result['status']}")
if result["status"] == "completed":
break
if result["status"] == "failed":
raise RuntimeError(result.get("error", "extract job failed"))
attempts += 1
if attempts >= max_attempts:
raise TimeoutError("Extract job did not finish within 5 minutes")
time.sleep(5)
with open("result.json", "w") as f:
json.dump(result, f, indent=2)
print(f"Saved extracted fields to result.json")
```
Save this as `script.mjs` or set `"type": "module"` in your `package.json`. Requires Node.js 18 or newer for the global `fetch`, `FormData`, and `Blob`.
```javascript script.mjs theme={null}
import fs from "node:fs";
const API_KEY = process.env.UNSILOED_API_KEY;
const BASE_URL = "https://prod.visionapi.unsiloed.ai";
const schema = {
type: "object",
properties: {
vendor_name: { type: "string", description: "Name of the company issuing the invoice (the seller)" },
invoice_number: { type: "string", description: "Unique invoice identifier shown on the document" },
issue_date: { type: "string", description: "Date the invoice was issued" },
total_due: { type: "number", description: "Final total amount due in US dollars, including tax" },
line_items: {
type: "array",
description: "One row per line item in the invoice table",
items: {
type: "object",
properties: {
description: { type: "string", description: "Description of the product or service" },
quantity: { type: "number", description: "Quantity of the item ordered" },
unit_price: { type: "number", description: "Price per unit in US dollars" },
subtotal: { type: "number", description: "Line subtotal in US dollars (quantity x unit_price)" },
},
required: ["description", "quantity", "unit_price", "subtotal"],
additionalProperties: false,
},
},
},
required: ["vendor_name", "invoice_number", "issue_date", "total_due", "line_items"],
additionalProperties: false,
};
const form = new FormData();
form.append("pdf_file", new Blob([fs.readFileSync("document.pdf")]), "document.pdf");
form.append("schema_data", JSON.stringify(schema));
const response = await fetch(`${BASE_URL}/v2/extract`, {
method: "POST",
headers: { "api-key": API_KEY },
body: form,
});
if (!response.ok) throw new Error(`${response.status}: ${await response.text()}`);
const { job_id } = await response.json();
console.log(`Job submitted: ${job_id}`);
const maxAttempts = 60; // roughly 5 minutes at 5 seconds per poll
let attempts = 0;
let result;
while (true) {
const res = await fetch(`${BASE_URL}/extract/${job_id}`, {
headers: { "api-key": API_KEY },
});
result = await res.json();
console.log(`Status: ${result.status}`);
if (result.status === "completed") break;
if (result.status === "failed") throw new Error(result.error || "extract job failed");
if (++attempts >= maxAttempts) throw new Error("Extract job did not finish within 5 minutes");
await new Promise((r) => setTimeout(r, 5000));
}
fs.writeFileSync("result.json", JSON.stringify(result, null, 2));
console.log("Saved extracted fields to result.json");
```
```bash theme={null}
# Write the schema to a file so we can pass it cleanly:
cat > schema.json <<'EOF'
{
"type": "object",
"properties": {
"vendor_name": { "type": "string", "description": "Name of the company issuing the invoice (the seller)" },
"invoice_number": { "type": "string", "description": "Unique invoice identifier shown on the document" },
"issue_date": { "type": "string", "description": "Date the invoice was issued" },
"total_due": { "type": "number", "description": "Final total amount due in US dollars, including tax" },
"line_items": {
"type": "array",
"description": "One row per line item in the invoice table",
"items": {
"type": "object",
"properties": {
"description": { "type": "string", "description": "Description of the product or service" },
"quantity": { "type": "number", "description": "Quantity of the item ordered" },
"unit_price": { "type": "number", "description": "Price per unit in US dollars" },
"subtotal": { "type": "number", "description": "Line subtotal in US dollars (quantity x unit_price)" }
},
"required": ["description", "quantity", "unit_price", "subtotal"],
"additionalProperties": false
}
}
},
"required": ["vendor_name", "invoice_number", "issue_date", "total_due", "line_items"],
"additionalProperties": false
}
EOF
# Submit the document with the schema and capture the job_id:
resp=$(curl -sX POST "https://prod.visionapi.unsiloed.ai/v2/extract" \
-H "api-key: $UNSILOED_API_KEY" \
-F "pdf_file=@document.pdf" \
-F "schema_data=$(cat schema.json)")
JOB_ID=$(echo "$resp" | grep -o '"job_id":"[^"]*"' | cut -d'"' -f4)
echo "Job submitted: $JOB_ID"
# Poll until the job finishes, with a 5-minute timeout:
attempts=0
max_attempts=60
while true; do
resp=$(curl -sX GET "https://prod.visionapi.unsiloed.ai/extract/$JOB_ID" \
-H "api-key: $UNSILOED_API_KEY")
status=$(echo "$resp" | grep -o '"status":"[^"]*"' | head -1 | cut -d'"' -f4)
echo "Status: $status"
[ "$status" = "completed" ] && break
[ "$status" = "failed" ] && { echo "Job failed"; exit 1; }
attempts=$((attempts + 1))
[ "$attempts" -ge "$max_attempts" ] && { echo "Extract job did not finish within 5 minutes"; exit 1; }
sleep 5
done
# Save the full response to disk:
echo "$resp" > result.json
```
## Step 1: Set Up Your Environment
Before writing any code, we need three things: an API key, a document, and the runtime for our chosen language.
### 1.1 Get an Unsiloed AI API Key
To get API access, [sign up on Unsiloed AI](https://cal.com/aman-mishra-p0ry57/15min). Export your key as an environment variable named `UNSILOED_API_KEY` so it stays out of source control:
```bash theme={null}
export UNSILOED_API_KEY="your-api-key"
```
### 1.2 Pick a Document to Extract Fields From
The `/v2/extract` endpoint supports PDF, DOCX, PPTX, JPG, PNG, and other formats. The walkthrough below assumes a PDF saved as `document.pdf` in your working directory. To use a different format, update the filename in the snippets to match your file.
If you don't have a document handy, download our [sample invoice PDF](https://raw.githubusercontent.com/Unsiloed-AI/cookbook/c585446e46e4be2790c6c29fe2a7a3a1b346191d/sample-documents/sample-extract.pdf) (a one-page invoice from Northwind Office Supplies with five line items) and save it as `document.pdf`. The schema in this guide targets the vendor, invoice number, issue date, total, and the line item table on that invoice.
### 1.3 Install Dependencies
You need Python 3.8 or newer. Install the `requests` package:
```bash theme={null}
pip install requests
```
You need Node.js 18 or newer for the global `fetch`, `FormData`, and `Blob`. No external packages needed.
You need cURL, which is preinstalled on macOS and most Linux distributions. No external packages needed.
## Step 2: Submit a Document With a Schema
The request bundles two fields: `pdf_file` for the document and `schema_data` for the JSON schema as a string. The schema is the interesting half. Anything we describe there, from a single total to a nested array of line items, comes back typed and scored in the exact shape we asked for. The endpoint returns a `job_id` we can poll. All requests go to `https://prod.visionapi.unsiloed.ai` with the API key in the `api-key` header.
### 2.1 Set Up the Script
Create a file called `extract_document.py` and start with the imports and configuration:
```python extract_document.py theme={null}
import json
import os
import time
import requests
API_KEY = os.environ["UNSILOED_API_KEY"]
BASE_URL = "https://prod.visionapi.unsiloed.ai"
```
`API_KEY` reads your key from the environment so it doesn't get hard-coded into the file, and `BASE_URL` points at the Unsiloed AI production endpoint. Both appear in every request below.
Create a file called `script.mjs` and start with the imports and configuration:
```javascript script.mjs theme={null}
import fs from "node:fs";
const API_KEY = process.env.UNSILOED_API_KEY;
const BASE_URL = "https://prod.visionapi.unsiloed.ai";
```
`API_KEY` reads your key from the environment so it doesn't get hard-coded into the file, and `BASE_URL` points at the Unsiloed AI production endpoint. Both appear in every request below.
cURL doesn't need a setup step. Each command below inlines the API key and base URL directly.
### 2.2 Define the Schema
The schema tells the API which fields to pull out and what shape they should take. The clearer the `description` on each field, the better the model locates and types each value. For our sample invoice, we want the vendor name, the invoice number, the issue date, the total due, and the five rows of the line item table.
Continue the file by defining the schema as a Python dict:
```python extract_document.py theme={null}
schema = {
"type": "object",
"properties": {
"vendor_name": {
"type": "string",
"description": "Name of the company issuing the invoice (the seller)",
},
"invoice_number": {
"type": "string",
"description": "Unique invoice identifier shown on the document",
},
"issue_date": {
"type": "string",
"description": "Date the invoice was issued",
},
"total_due": {
"type": "number",
"description": "Final total amount due in US dollars, including tax",
},
"line_items": {
"type": "array",
"description": "One row per line item in the invoice table",
"items": {
"type": "object",
"properties": {
"description": {"type": "string", "description": "Description of the product or service"},
"quantity": {"type": "number", "description": "Quantity of the item ordered"},
"unit_price": {"type": "number", "description": "Price per unit in US dollars"},
"subtotal": {"type": "number", "description": "Line subtotal in US dollars (quantity x unit_price)"},
},
"required": ["description", "quantity", "unit_price", "subtotal"],
"additionalProperties": False,
},
},
},
"required": ["vendor_name", "invoice_number", "issue_date", "total_due", "line_items"],
"additionalProperties": False,
}
```
The schema is plain JSON Schema with strict-mode rules: `additionalProperties: false` at every object level (which prevents the model from inventing fields we didn't ask for), and a `required` list naming the must-have fields. See the [Schemas](/docs/document-processing/extraction/schemas) reference for the full ruleset.
Continue the file by defining the schema as a JavaScript object:
```javascript script.mjs theme={null}
const schema = {
type: "object",
properties: {
vendor_name: { type: "string", description: "Name of the company issuing the invoice (the seller)" },
invoice_number: { type: "string", description: "Unique invoice identifier shown on the document" },
issue_date: { type: "string", description: "Date the invoice was issued" },
total_due: { type: "number", description: "Final total amount due in US dollars, including tax" },
line_items: {
type: "array",
description: "One row per line item in the invoice table",
items: {
type: "object",
properties: {
description: { type: "string", description: "Description of the product or service" },
quantity: { type: "number", description: "Quantity of the item ordered" },
unit_price: { type: "number", description: "Price per unit in US dollars" },
subtotal: { type: "number", description: "Line subtotal in US dollars (quantity x unit_price)" },
},
required: ["description", "quantity", "unit_price", "subtotal"],
additionalProperties: false,
},
},
},
required: ["vendor_name", "invoice_number", "issue_date", "total_due", "line_items"],
additionalProperties: false,
};
```
The schema is plain JSON Schema with strict-mode rules: `additionalProperties: false` at every object level (which prevents the model from inventing fields we didn't ask for), and a `required` list naming the must-have fields. See the [Schemas](/docs/document-processing/extraction/schemas) reference for the full ruleset.
Write the schema to a file to pass it to `curl` as the value of the `schema_data` form field:
```bash theme={null}
cat > schema.json <<'EOF'
{
"type": "object",
"properties": {
"vendor_name": { "type": "string", "description": "Name of the company issuing the invoice (the seller)" },
"invoice_number": { "type": "string", "description": "Unique invoice identifier shown on the document" },
"issue_date": { "type": "string", "description": "Date the invoice was issued" },
"total_due": { "type": "number", "description": "Final total amount due in US dollars, including tax" },
"line_items": {
"type": "array",
"description": "One row per line item in the invoice table",
"items": {
"type": "object",
"properties": {
"description": { "type": "string", "description": "Description of the product or service" },
"quantity": { "type": "number", "description": "Quantity of the item ordered" },
"unit_price": { "type": "number", "description": "Price per unit in US dollars" },
"subtotal": { "type": "number", "description": "Line subtotal in US dollars (quantity x unit_price)" }
},
"required": ["description", "quantity", "unit_price", "subtotal"],
"additionalProperties": false
}
}
},
"required": ["vendor_name", "invoice_number", "issue_date", "total_due", "line_items"],
"additionalProperties": false
}
EOF
```
The schema follows JSON Schema with strict-mode rules: `additionalProperties: false` at every object level, and a `required` list naming the must-have fields. See the [Schemas](/docs/document-processing/extraction/schemas) reference for the full ruleset.
### 2.3 Upload the Document
Send the file and the schema as a multipart upload to `/v2/extract`. The endpoint expects the document under the form field name `pdf_file` and the schema under `schema_data` (as a JSON string).
Next, upload the document and the schema together:
```python extract_document.py theme={null}
with open("document.pdf", "rb") as f:
response = requests.post(
f"{BASE_URL}/v2/extract",
headers={"api-key": API_KEY},
files={"pdf_file": ("document.pdf", f, "application/pdf")},
data={"schema_data": json.dumps(schema)},
)
response.raise_for_status()
```
The `raise_for_status()` call throws an `HTTPError` on any non-2xx response, so we don't need to check `.status_code` ourselves. The `json.dumps(schema)` call serializes the dict because the endpoint expects `schema_data` as a string, not a nested form field.
Next, upload the document and the schema together:
```javascript script.mjs theme={null}
const form = new FormData();
form.append("pdf_file", new Blob([fs.readFileSync("document.pdf")]), "document.pdf");
form.append("schema_data", JSON.stringify(schema));
const response = await fetch(`${BASE_URL}/v2/extract`, {
method: "POST",
headers: { "api-key": API_KEY },
body: form,
});
if (!response.ok) throw new Error(`${response.status}: ${await response.text()}`);
```
`fetch` doesn't throw on non-2xx responses by default, so we check `response.ok` and raise the error ourselves. The `JSON.stringify(schema)` call serializes the schema because the endpoint expects `schema_data` as a string, not a nested form field.
Run:
```bash theme={null}
curl -X POST "https://prod.visionapi.unsiloed.ai/v2/extract" \
-H "api-key: $UNSILOED_API_KEY" \
-F "pdf_file=@document.pdf" \
-F "schema_data=$(cat schema.json)"
```
The response prints to stdout. We need the `job_id` field for the next step.
### 2.4 Capture the Job ID
Then read and print the `job_id`:
```python extract_document.py theme={null}
job_id = response.json()["job_id"]
print(f"Job submitted: {job_id}")
```
Run the script:
```bash theme={null}
python extract_document.py
```
The output should be a single line like `Job submitted: a90e48c4-f564-435e-9bf2-ab6eb5a0376d`.
Then read and log the `job_id`:
```javascript script.mjs theme={null}
const { job_id } = await response.json();
console.log(`Job submitted: ${job_id}`);
```
Run the script:
```bash theme={null}
node script.mjs
```
The output should be a single line like `Job submitted: a90e48c4-f564-435e-9bf2-ab6eb5a0376d`.
The response body from the POST above looks like:
```json theme={null}
{
"job_id": "a90e48c4-f564-435e-9bf2-ab6eb5a0376d",
"status": "queued",
"message": "PDF citation processing started",
"quota_remaining": 7705
}
```
Copy the `job_id` value; you'll paste it into the polling command in the next step.
## Step 3: Poll for Results
The job runs asynchronously. We GET `/extract/{job_id}` repeatedly until the status is `completed`, then save the extracted fields to disk.
A status of `completed` means the result is ready; `failed` means the job errored; any other value (`queued`, `processing`, and so on) means the job is still running.
### 3.1 Write the Polling Loop
Next, drop in a polling loop. The `max_attempts` cap stops the loop if the job hangs:
```python extract_document.py theme={null}
max_attempts = 60 # roughly 5 minutes at 5 seconds per poll
attempts = 0
while True:
result = requests.get(
f"{BASE_URL}/extract/{job_id}",
headers={"api-key": API_KEY},
).json()
print(f"Status: {result['status']}")
if result["status"] == "completed":
break
if result["status"] == "failed":
raise RuntimeError(result.get("error", "extract job failed"))
attempts += 1
if attempts >= max_attempts:
raise TimeoutError("Extract job did not finish within 5 minutes")
time.sleep(5)
```
Next, drop in a polling loop. The `maxAttempts` cap stops the loop if the job hangs:
```javascript script.mjs theme={null}
const maxAttempts = 60; // roughly 5 minutes at 5 seconds per poll
let attempts = 0;
let result;
while (true) {
const res = await fetch(`${BASE_URL}/extract/${job_id}`, {
headers: { "api-key": API_KEY },
});
result = await res.json();
console.log(`Status: ${result.status}`);
if (result.status === "completed") break;
if (result.status === "failed") throw new Error(result.error || "extract job failed");
if (++attempts >= maxAttempts) throw new Error("Extract job did not finish within 5 minutes");
await new Promise((r) => setTimeout(r, 5000));
}
```
Replace `JOB_ID` below with the value you captured from Step 2.4, then run this loop. It polls every 5 seconds and gives up after 5 minutes if the job hasn't completed:
```bash theme={null}
JOB_ID="paste-job-id-here"
attempts=0
max_attempts=60 # roughly 5 minutes at 5 seconds per poll
while true; do
resp=$(curl -sX GET "https://prod.visionapi.unsiloed.ai/extract/$JOB_ID" \
-H "api-key: $UNSILOED_API_KEY")
status=$(echo "$resp" | grep -o '"status":"[^"]*"' | head -1 | cut -d'"' -f4)
echo "Status: $status"
[ "$status" = "completed" ] && break
[ "$status" = "failed" ] && { echo "Job failed"; exit 1; }
attempts=$((attempts + 1))
[ "$attempts" -ge "$max_attempts" ] && { echo "Extract job did not finish within 5 minutes"; exit 1; }
sleep 5
done
```
The loop keeps the latest response body in `$resp` for the next step.
### 3.2 Save the Extracted Fields
Finally, persist the result to disk. The response is already structured JSON, so we write it straight to `result.json`.
Finally, write the result to disk:
```python extract_document.py theme={null}
with open("result.json", "w") as f:
json.dump(result, f, indent=2)
print(f"Saved extracted fields to result.json")
```
Run the script:
```bash theme={null}
python extract_document.py
```
You should see a few `Status: processing` lines, then `Status: completed`, then the summary line. The `result.json` file appears in the working directory.
Finally, write the result to disk:
```javascript script.mjs theme={null}
fs.writeFileSync("result.json", JSON.stringify(result, null, 2));
console.log("Saved extracted fields to result.json");
```
Run the script:
```bash theme={null}
node script.mjs
```
You should see a few `Status: processing` lines, then `Status: completed`, then the summary line. The `result.json` file appears in the working directory.
The polling loop in Step 3.1 left the full response in `$resp`. Write it to disk:
```bash theme={null}
echo "$resp" > result.json
```
The `result.json` file now holds the full response. Extracted fields sit under `result.{field_name}`, each with a `value` and a `score`.
## Error Responses
Failures fall into two buckets: HTTP errors raised before the job is queued, and a `failed` status on a job that started but couldn't complete.
### HTTP Errors
The `/v2/extract` endpoint returns JSON bodies on HTTP errors, with a single `detail` field describing the problem. The common cases:
* **`401 Unauthorized`:** body is `{"detail": "Invalid API key"}`. The `api-key` header is missing or wrong.
* **`400 Bad Request` (missing file):** body is `{"detail": "Either pdf_file or file_url must be provided"}`. The `pdf_file` form field is missing.
* **`400 Bad Request` (bad JSON):** body is `{"detail": "schema_data must be valid JSON"}`. The `schema_data` field isn't parseable JSON.
* **`400 Bad Request` (unreadable PDF):** body starts with `{"detail": "Error extracting data with citations: Failed to get PDF page count..."}`. The uploaded file isn't a valid PDF.
* **`422 Unprocessable Entity`:** body lists the missing or malformed form fields. Usually thrown when `schema_data` is absent.
* **`404 Not Found`:** body is `{"detail": "Job not found"}`. The `job_id` you polled doesn't exist.
### Failed Jobs
A job that was accepted but couldn't be processed comes back with `status: "failed"` on a subsequent poll. The response shape mirrors a completed one, with an `error` field describing what went wrong:
```json theme={null}
{
"job_id": "7b31a7d7-e810-4a0b-931e-fbed0879bab2",
"status": "failed",
"file_name": "document.pdf",
"error": "Failed to extract structured data from document"
}
```
## Response Shape
A completed response contains job metadata plus a `result` object with one entry per top-level field in your schema. Each entry has the extracted `value` and a `score` between 0 and 1. For arrays of objects, the array itself has a `score`, and every property inside each row carries its own score as well.
```json theme={null}
{
"job_id": "cec5dcb5-53c6-47d5-afe7-28b2182171fb",
"status": "completed",
"file_name": "document.pdf",
"file_url": "https://example-bucket.s3.amazonaws.com/...",
"created_at": "2026-05-27T08:36:58.336535+00:00",
"updated_at": "2026-05-27T08:37:19.436604+00:00",
"metadata": {
"page_count": 1,
"order": ["vendor_name", "invoice_number", "issue_date", "total_due", "line_items"],
"schema": { "...": "..." }
},
"result": {
"vendor_name": { "value": "Northwind Office Supplies", "score": 0.96 },
"invoice_number": { "value": "INV-2026-00487", "score": 0.97 },
"issue_date": { "value": "April 14, 2026", "score": 0.98 },
"total_due": { "value": 3705.1, "score": 0.96 },
"line_items": {
"score": 0.98,
"value": [
{
"description": { "value": "Ergonomic Mesh Office Chair", "score": 0.95, "citation": null },
"quantity": { "value": 4, "score": 0.96, "citation": null },
"unit_price": { "value": 289.0, "score": 0.97, "citation": null },
"subtotal": { "value": 1156.0, "score": 0.98, "citation": null }
},
"...four more rows..."
]
}
}
}
```
The fields you'll actually use depend on what you're building. They fall into three broad categories:
**For typed values and validation:**
* **`result.{field_name}.value`:** the extracted data, typed to match your schema (`string`, `number`, `boolean`, `object`, or `array`)
* **`result.{field_name}.score`:** confidence score between 0 and 1, higher is better. Use it to flag uncertain values for human review.
* **`result.{array_field}.value[].{property}.citation`:** reserved slot for source citations on array rows; `null` for now
**For schema and ordering:**
* **`metadata.schema`:** an echo of the schema you submitted, useful for round-tripping or auditing
* **`metadata.order`:** the original order of top-level fields in your schema, since JSON objects don't preserve insertion order across all clients
* **`metadata.page_count`:** number of pages in the uploaded document
**For job and audit tracking:**
* **`job_id`:** unique identifier for the extraction job
* **`status`:** `completed`, `failed`, or an in-progress value (`queued`, `processing`)
* **`file_name`:** name of the uploaded file
* **`file_url`:** temporary signed S3 URL to the uploaded file
* **`created_at`, `updated_at`:** ISO 8601 timestamps for submission and the most recent status change
### Sample Output
Running the script against the [sample invoice](https://raw.githubusercontent.com/Unsiloed-AI/cookbook/c585446e46e4be2790c6c29fe2a7a3a1b346191d/sample-documents/sample-extract.pdf) writes the JSON above to `result.json`. Every field comes back with its own confidence score, so flagging uncertain values becomes a per-field check rather than a re-read of the document. The fields extracted from the sample:
| Field | Extracted value | Confidence |
| - | - | - |
| `vendor_name` | Northwind Office Supplies | 96% |
| `invoice_number` | INV-2026-00487 | 97% |
| `issue_date` | April 14, 2026 | 98% |
| `total_due` | 3705.10 | 96% |
| `line_items` | 5 rows | 98% |
And the five rows of `line_items`, each a structured object in its own right:
| Description | Quantity | Unit Price | Subtotal |
| - | - | - | - |
| Ergonomic Mesh Office Chair | 4 | 289.00 | 1,156.00 |
| Adjustable Standing Desk (60" x 30") | 2 | 549.00 | 1,098.00 |
| LED Desk Lamp with USB Charging | 6 | 42.50 | 255.00 |
| Acoustic Panel, 24" Hexagon (Pack of 4) | 3 | 78.00 | 234.00 |
| Wireless Mechanical Keyboard | 5 | 129.99 | 649.95 |
Every cell in the table has its own score in the underlying JSON, so downstream code can flag individual uncertain values without rejecting the whole row.
## Next Steps
For more on extraction, including schema rules, supported types, and the full response reference, see the [Extract overview](/docs/document-processing/extraction/extraction).
JSON Schema rules, supported types, and worked examples for invoices and SEC filings.
The canonical extraction response with a field-by-field reference.
Browse the full request and response specs for `/v2/extract`.
Check limits, supported formats, and answers to common questions.
# Response Format
Source: https://docs.unsiloed.ai/docs/document-processing/extraction/response-format
Response shape from the /v2/extract endpoint, with a field-by-field reference.
A completed extraction job returns job metadata plus a `result` object with one entry per top-level field in your schema. The example below is from a real invoice extraction.
```json theme={null}
{
"job_id": "283c77a1-ae1b-4b96-b89e-1c5e63fd89aa",
"status": "completed",
"file_name": "invoice.pdf",
"file_url": "https://example-bucket.s3.amazonaws.com/...",
"created_at": "2026-05-25T15:38:11.821Z",
"updated_at": "2026-05-25T15:38:42.196Z",
"metadata": {
"page_count": 1,
"order": ["invoice_number", "vendor_name", "total_amount_due"],
"schema": { "...": "..." }
},
"result": {
"invoice_number": {
"value": "NC-2025-00417",
"score": 0.97
},
"vendor_name": {
"value": "NORTHWIND CONSULTING LLC",
"score": 0.95
},
"total_amount_due": {
"value": 12586.0,
"score": 0.97
}
}
}
```
## Top-Level Fields
* **`job_id`:** unique identifier for the extraction job
* **`status`:** job state (`completed`, `failed`, or an in-progress value such as `processing`)
* **`file_name`:** name of the uploaded file
* **`file_url`:** temporary signed S3 URL to the uploaded file
* **`created_at`:** ISO 8601 timestamp when the job was submitted
* **`updated_at`:** ISO 8601 timestamp of the most recent status update
* **`metadata`:** object containing `page_count`, the original `order` of fields in your schema, and an echo of the `schema` you submitted
* **`result`:** object containing one entry per top-level field in your schema
## Extracted Field Structure
Each entry in the `result` object describes one extracted field:
* **`value`:** the extracted data, typed to match your schema (`string`, `number`, `boolean`, `object`, or `array`)
* **`score`:** confidence score between 0 and 1, higher is better
## Array and Nested Object Fields
When your schema includes an array of objects (like `line_items` in an invoice), each row comes back as a structured object where every property has its own `value` and `score`:
```json theme={null}
"line_items": {
"score": 0.99,
"value": [
{
"description": { "value": "Strategic planning workshop facilitation", "score": 0.98 },
"hours": { "value": 12, "score": 0.97 },
"rate": { "value": 275.0, "score": 0.97 },
"amount": { "value": 3300.0, "score": 0.98 }
}
]
}
```
The top-level `score` on the array (`0.99` above) reflects whether the parser identified the array structure correctly. Each cell carries its own score, so downstream code can flag individual uncertain values without rejecting the whole row.
## Error Responses
If the schema fails validation or the job errors during processing, the API returns a JSON body with an error description. See the [Extract API reference](/docs/api-reference/extraction/extract-data) for the exact error shapes and status codes.
# Schemas
Source: https://docs.unsiloed.ai/docs/document-processing/extraction/schemas
JSON Schema rules and patterns for defining what /v2/extract should pull out of a document.
Every extraction schema follows JSON Schema with strict-mode rules. These rules apply at every level (the root, every nested object, and every array's `items` definition) and keep the output deterministic and well-typed.
Prefer not to write JSON by hand? The [Unsiloed dashboard](https://app.unsiloed.ai/playground/Extractor) has a schema builder with Manual and Auto-Suggest modes. In Auto-Suggest, describe the fields you want, upload an example document, and the dashboard generates a schema you can export and pass to `/v2/extract`.
## Core Requirements
**1. Root Object**
Every schema starts with `"type": "object"`. Arrays and primitives aren't allowed at the top level.
```json theme={null}
{
"type": "object",
"properties": {
// Define your fields here
},
"required": [...],
"additionalProperties": false
}
```
**2. Properties**
Define all fields you want to extract using the `"properties"` key. Each field must specify a `"type"` and should include a clear `"description"`.
```json theme={null}
{
"type": "object",
"properties": {
"field_name_1": {
"type": "string",
"description": "Clear description of what to extract"
},
"field_name_2": {
"type": "number",
"description": "Description with units or context"
}
},
"required": [...],
"additionalProperties": false
}
```
**3. Required Fields**
Specify mandatory fields using the `"required"` array. Field names must exactly match those defined in `"properties"`.
```json theme={null}
{
"type": "object",
"properties": {
"mandatory_field": { "type": "string", "description": "This field is required" },
"another_required_field": { "type": "string", "description": "This is also required" }
},
"required": ["mandatory_field", "another_required_field"],
"additionalProperties": false
}
```
**4. Additional Properties**
Always set `"additionalProperties": false` at every object level to ensure only specified fields appear in output.
```json theme={null}
{
"type": "object",
"properties": {
"items": {
"type": "array",
"items": {
"type": "object",
"properties": {
"field_name": { "type": "string" }
},
"required": [...],
"additionalProperties": false // Required in array items
}
}
},
"required": [...],
"additionalProperties": false // Required at root level
}
```
## Supported Types
Extraction schemas support four field types:
**String:** For text, dates, IDs, names, addresses, and any textual data
```json theme={null}
{
"field_name": {
"type": "string",
"description": "Description of the text field"
}
}
```
**Number:** For integers and decimals like prices, quantities, counts, and measurements
```json theme={null}
{
"field_name": {
"type": "number",
"description": "Description with units (e.g., USD, kg)"
}
}
```
**Boolean:** For true/false values such as status flags and yes/no fields
```json theme={null}
{
"field_name": {
"type": "boolean",
"description": "Description of the boolean condition"
}
}
```
**Array:** For repeating items like line items or lists. Must include `items` to define the structure of array elements
```json theme={null}
{
"field_name": {
"type": "array",
"description": "Description of the array items",
"items": {
"type": "object",
"properties": {
"item_field1": { "type": "string", "description": "..." },
"item_field2": { "type": "string", "description": "..." }
},
"required": [...],
"additionalProperties": false
}
}
}
```
## Building Schemas
### Primitive Types
Primitive fields use `string`, `number`, or `boolean` as their `type`. These are the building blocks of your schema.
```json theme={null}
{
"type": "object",
"properties": {
"invoice_number": {
"type": "string",
"description": "Invoice number"
},
"total_amount": {
"type": "number",
"description": "Total amount in USD"
},
"is_paid": {
"type": "boolean",
"description": "Payment status"
}
},
"required": ["invoice_number", "total_amount", "is_paid"],
"additionalProperties": false
}
```
### Arrays of Objects
Use arrays when you have repeating data like line items, transactions, or people.
```json theme={null}
{
"type": "object",
"properties": {
"line_items": {
"type": "array",
"description": "Invoice line items",
"items": {
"type": "object",
"properties": {
"description": {
"type": "string",
"description": "Item description"
},
"quantity": {
"type": "number",
"description": "Quantity"
},
"price": {
"type": "number",
"description": "Unit price"
}
},
"required": ["description", "quantity", "price"],
"additionalProperties": false
}
}
},
"additionalProperties": false
}
```
### Nested Arrays
For hierarchical data, nest an array inside the objects of another array.
```json theme={null}
{
"type": "object",
"properties": {
"orders": {
"type": "array",
"description": "Customer orders",
"items": {
"type": "object",
"properties": {
"order_id": {
"type": "string",
"description": "Order ID"
},
"shipments": {
"type": "array",
"description": "Shipments for this order",
"items": {
"type": "object",
"properties": {
"tracking_number": {
"type": "string",
"description": "Tracking number"
},
"carrier": {
"type": "string",
"description": "Shipping carrier"
}
},
"required": ["tracking_number", "carrier"],
"additionalProperties": false
}
}
},
"required": ["order_id", "shipments"],
"additionalProperties": false
}
}
},
"additionalProperties": false
}
```
A common schema for extracting data from US invoices.
```json theme={null}
{
"type": "object",
"properties": {
"document_type": {
"type": "string",
"description": "Type of document (invoice, receipt, etc.)"
},
"invoice_header": {
"type": "object",
"description": "Invoice header details",
"properties": {
"invoice_number": {
"type": "string",
"description": "Invoice number"
},
"invoice_date": {
"type": "string",
"description": "Invoice issue date"
},
"due_date": {
"type": "string",
"description": "Payment due date"
}
},
"required": ["invoice_number", "invoice_date"],
"additionalProperties": false
},
"vendor": {
"type": "object",
"description": "Vendor information",
"properties": {
"vendor_name": {
"type": "string",
"description": "Legal business name"
},
"vendor_address": {
"type": "string",
"description": "Business address"
},
"vendor_email": {
"type": "string",
"description": "Accounts receivable contact email"
}
},
"required": ["vendor_name"],
"additionalProperties": false
},
"line_items": {
"type": "array",
"description": "List of billed items",
"items": {
"type": "object",
"properties": {
"description": {
"type": "string",
"description": "Item or service description"
},
"quantity": {
"type": "number",
"description": "Quantity billed"
},
"unit_price": {
"type": "number",
"description": "Price per unit in USD"
},
"line_total": {
"type": "number",
"description": "Total cost for this line item"
}
},
"required": ["description", "quantity", "unit_price"],
"additionalProperties": false
}
},
"invoice_totals": {
"type": "object",
"description": "Invoice totals",
"properties": {
"subtotal": {
"type": "number",
"description": "Subtotal before tax"
},
"sales_tax": {
"type": "number",
"description": "Sales tax amount"
},
"total_amount_due": {
"type": "number",
"description": "Final amount due in USD"
}
},
"required": ["total_amount_due"],
"additionalProperties": false
}
},
"required": ["document_type", "invoice_header", "invoice_totals"],
"additionalProperties": false
}
```
A schema for extracting governance and ownership information from US SEC filings.
```json theme={null}
{
"type": "object",
"properties": {
"board_of_directors": {
"type": "array",
"description": "Board of Directors",
"items": {
"type": "object",
"properties": {
"director_name": {
"type": "string",
"description": "Full name of board member"
}
},
"required": ["director_name"],
"additionalProperties": false
}
},
"major_shareholders": {
"type": "array",
"description": "Major shareholders and ownership",
"items": {
"type": "object",
"properties": {
"shareholder_name": {
"type": "string",
"description": "Name of shareholder"
},
"ownership_percentage": {
"type": "string",
"description": "Percentage ownership"
}
},
"required": ["shareholder_name", "ownership_percentage"],
"additionalProperties": false
}
}
},
"required": ["board_of_directors", "major_shareholders"],
"additionalProperties": false
}
```
# Element Types
Source: https://docs.unsiloed.ai/docs/document-processing/parsing/element-types
The segment types returned by the /parse endpoint, with example response shapes pulled from real parsed documents.
Every segment in a parsed document carries a `segment_type` field naming the layout region it came from. The parser recognizes the types listed below, divided into **text elements** (regions whose meaning lives in their characters and structure) and **visual elements** (regions whose meaning lives in their layout, image content, or rendered form).
Two of them (`KeyValuePair` and `Signature`) are not produced by the default layout model. You get them by passing `segmentation_version=v2` — the v2 layout model is form-aware and detects these regions alongside the standard set.
All segments share the same core fields: `bbox`, `confidence`, `content`, `markdown`, `html`, `ocr`, and location metadata. What changes by type is what those fields contain, and a couple of types omit specific fields entirely. The sections below show a real response sample for each type.
## Text Elements
These segments carry their meaning in text, so the `markdown` and `html` fields use semantic markup like headers, italics, list syntax, and footnote references to reflect each type.
### Text
Regular paragraph and inline text. The `content` field carries the plain text, `markdown` is the same with line breaks preserved, and `html` wraps any line breaks in ` `.
```json theme={null}
{
"segment_id": "13d55851-d0fc-4999-a508-ab82d9a64443",
"segment_type": "Text",
"content": "The following table summarises regional sales performance for Q1\n2024.",
"markdown": "The following table summarises regional sales performance for Q1\n2024.",
"html": "The following table summarises regional sales performance for Q1 \n2024.",
"bbox": { "left": 60.8, "top": 152.6, "width": 714.6, "height": 24.3 },
"page_number": 1,
"page_width": 1191.0,
"page_height": 1684.0,
"confidence": 0.99
}
```
### Title
Document titles and main headings. Rendered as a top-level Markdown header (`#`) and `
` in HTML, distinct from `SectionHeader` which uses `##`/`
",
"bbox": { "left": 427.6, "top": 67.8, "width": 344.7, "height": 36.5 },
"page_number": 1,
"page_width": 1191.0,
"page_height": 1684.0,
"confidence": 0.35
}
```
### ListItem
Bulleted and numbered list entries. The `markdown` field renders the item with a leading dash, and `html` wraps the entry in `
` (with a nested `` if the source list was numbered).
```json theme={null}
{
"segment_id": "32217825-2ddc-4121-be70-2b3e23e2ab97",
"segment_type": "ListItem",
"content": "1. Operating Conditions — 1966",
"markdown": "- 1. Operating Conditions — 1966",
"html": "
\n
Operating Conditions — 1966
\n
",
"bbox": { "left": 441.9, "top": 430.0, "width": 288.8, "height": 20.7 },
"page_number": 3,
"page_width": 1222.0,
"page_height": 1576.0,
"confidence": 0.95
}
```
### Caption
Text captions associated with images, figures, or tables. The `markdown` field wraps the caption in italics (`_..._`), and `html` wraps it in a `` for downstream styling.
```json theme={null}
{
"segment_id": "63d646e8-0cb1-4325-8759-86625a51b0f9",
"segment_type": "Caption",
"content": "Figure 1: The Transformer - model\narchitecture.",
"markdown": "_Figure 1: The Transformer - model\narchitecture._",
"html": "Figure 1: The Transformer - model \narchitecture.",
"bbox": { "left": 418.8, "top": 808.4, "width": 385.1, "height": 21.0 },
"page_number": 3,
"page_width": 1224.0,
"page_height": 1584.0,
"confidence": 1.0
}
```
### Footnote
Footnote text and references. The `markdown` field uses Markdown footnote syntax (`[^...]`), and `html` wraps the body in a ``.
```json theme={null}
{
"segment_id": "840069a9-fb5e-4dcb-ad37-459bd4ff29f1",
"segment_type": "Footnote",
"content": "∗Equal contribution. Listing order is random. Jakob proposed replacing RNNs with self-attention...",
"markdown": "[^∗Equal contribution. Listing order is random. Jakob proposed replacing RNNs with self-attention...]",
"html": "∗Equal contribution. Listing order is random...",
"bbox": { "left": 214.4, "top": 1196.5, "width": 793.9, "height": 176.2 },
"page_number": 1,
"page_width": 1224.0,
"page_height": 1584.0,
"confidence": 1.0
}
```
### PageHeader
Header content at the top of a page, such as library stamps, document titles repeating across pages, or running headers. The `markdown` and `html` fields carry the raw text without semantic markup. Often worth filtering out for clean RAG ingestion.
```json theme={null}
{
"segment_id": "0f1fce17-da69-4698-b7a6-3abf954ee41e",
"segment_type": "PageHeader",
"content": "CLEVELAND PUBLIC LIBRARY BUSINESS INF. BUR.\nCORPORATION FILE",
"markdown": "CLEVELAND PUBLIC LIBRARY BUSINESS INF. BUR.\nCORPORATION FILE",
"html": "CLEVELAND PUBLIC LIBRARY BUSINESS INF. BUR. \nCORPORATION FILE",
"bbox": { "left": 998.3, "top": 2.6, "width": 193.5, "height": 76.7 },
"page_number": 1,
"page_width": 1224.0,
"page_height": 1576.0,
"confidence": 0.34
}
```
### PageFooter
Footer content at the bottom of a page, typically page numbers, copyright notices, or document IDs. Like `PageHeader`, often filtered out before embedding.
```json theme={null}
{
"segment_id": "997cbb53-0d58-49dc-b10e-e997ee14aafc",
"segment_type": "PageFooter",
"content": "5",
"markdown": "5",
"html": "5",
"bbox": { "left": 611.0, "top": 1410.1, "width": 12.4, "height": 21.6 },
"page_number": 7,
"page_width": 1224.0,
"page_height": 1576.0,
"confidence": 0.94
}
```
### KeyValuePair
A form field region: labels together with the values filled in against them, such as `Passport No : J8369854` or a row of checkboxes. Returned when you pass `segmentation_version=v2`.
A label and its value stay in the same segment rather than being split across neighbouring segments. The `markdown` field carries the pairs, checkbox and radio state included, and the `image` field carries a signed URL to the cropped region when you poll with `include_url=true`.
```json theme={null}
{
"segment_id": "85fcc4f1-f637-4d57-a698-c4bc4bfb0d3e",
"segment_type": "KeyValuePair",
"content": "Passport No : J8369854\nDate of Issue : 14/03/2019\nStatus Single Married",
"markdown": "**Passport No :** J8369854\n**Date of Issue :** 14/03/2019\n**Status :** ☒ Single ☐ Married",
"html": "
Passport No : J8369854
\n
Date of Issue : 14/03/2019
\n
Status : ☒ Single ☐ Married
",
"image": "https://s3.us-east-1.amazonaws.com/...",
"bbox": { "left": 95.5, "top": 217.9, "width": 610.4, "height": 92.2 },
"page_number": 1,
"page_width": 1191.0,
"page_height": 1684.0,
"confidence": 1.0
}
```
The `html` above is what you get with `segment_analysis={"KeyValuePair": {"html": "VLM"}}`. At the default `html: "Auto"`, the HTML comes back as a single `
` wrapper.
## Visual Elements
These segments carry their meaning in visual content or layout. The `markdown` and `html` fields contain either rendered structured content (Markdown tables, LaTeX) or AI-generated descriptions for image regions.
### Table
Tabular data with structured rows and columns. The `markdown` field carries the Markdown pipe-table syntax, `html` carries a full `
` with `` and ``, and `content` is a flat plain-text approximation. The `image` field, returned when you poll with `include_url=true`, contains a signed URL to a cropped image of the table region, useful for verifying parses visually or feeding the original table to an image-input model.
```json theme={null}
{
"segment_id": "4f4b54bc-793e-49cc-b0a3-113bbb5484be",
"segment_type": "Table",
"content": "Region Sales Rep Units Sold Revenue ($) Target ($) % of Target\nNorth Alice Brown 1,240 186,000 175,000 106%\n...",
"markdown": "| Region | Sales Rep | Units Sold | Revenue ($) | Target ($) | % of Target |\n| --- | --- | --- | --- | --- | --- |\n| North | Alice Brown | 1,240 | 186,000 | 175,000 | 106% |\n| ... | ... | ... | ... | ... | ... |",
"html": "
\n \n
Region
Sales Rep
Units Sold
...
\n \n \n
North
Alice Brown
...
\n ...\n \n
",
"image": "https://s3.us-east-1.amazonaws.com/...",
"bbox": { "left": 54.4, "top": 208.5, "width": 1026.5, "height": 246.5 },
"page_number": 1,
"page_width": 1191.0,
"page_height": 1684.0,
"confidence": 0.99
}
```
### Picture
Images, charts, illustrations, and diagrams. The `image` field, returned when you poll with `include_url=true`, contains a signed URL to the cropped picture itself. The `markdown` and `html` fields contain an AI-generated description of the image (not the image bytes), making the picture's visual content searchable and embeddable as text alongside the rest of the document. The `content` field carries any OCR text detected inside the picture region, such as chart labels and axis values, and is omitted when the picture yields no legible text.
```json theme={null}
{
"segment_id": "38cee134-39a2-49b9-9702-e3697f238f5e",
"segment_type": "Picture",
"markdown": "# Image Description\n\nThe image is a financial line chart titled **GROWTH OF $1,000,000: FOR THE 10-YEAR PERIOD ENDED 10/31/23**.\n\n## Legend\n\n- **Calamos Market Neutral Income Fund (I shares at NAV)**\n- **Bloomberg U.S. Government/Credit Index**\n- **Bloomberg Short Treasury 1-3 Month Index**\n\n...\n\n## Visible End Values\n\n- **$1,414,955**\n- **$1,120,466**\n- **$1,113,559**",
"html": "
Image Description
The image is a financial line chart titled GROWTH OF $1,000,000: FOR THE 10-YEAR PERIOD ENDED 10/31/23...
",
"image": "https://s3.us-east-1.amazonaws.com/...",
"bbox": { "left": 70.4, "top": 178.2, "width": 1046.1, "height": 236.4 },
"page_number": 1,
"page_width": 1188.0,
"page_height": 1548.0,
"confidence": 0.87
}
```
### Formula
Mathematical equations and expressions. The most distinctive type: the `markdown` and `html` fields contain LaTeX wrapped in `$...$`, ready to render with KaTeX, MathJax, or any other LaTeX-aware tool. The `content` field carries a plain-text OCR approximation of the equation, which is usually less reliable than the LaTeX representation.
```json theme={null}
{
"segment_id": "b4ccd4cf-01ae-4e92-b881-3bb1c335e8b3",
"segment_type": "Formula",
"content": "V ) = softmax(QKT )V (1)\nAttention(Q, K, √\ndk",
"markdown": "$\\mathrm{Attention}(Q, K, V) = \\mathrm{softmax}\\left(\\frac{QK^T}{\\sqrt{d_k}}\\right)V$",
"html": "
",
"bbox": { "left": 438.4, "top": 928.8, "width": 569.9, "height": 52.7 },
"page_number": 4,
"page_width": 1224.0,
"page_height": 1584.0,
"confidence": 1.0
}
```
### Signature
A handwritten signature region. Returned when you pass `segmentation_version=v2`. Like `Picture`, the `markdown` and `html` fields contain an AI-generated description of what the handwriting looks like, useful as searchable text, and the `content` field holds whatever OCR text the region yields, which for cursive strokes is usually nothing, so the field is often omitted. Unlike `Picture`, a Signature segment never carries an `image` URL, only the description and bounding box.
```json theme={null}
{
"segment_id": "b328b32d-1c37-41f2-a5f7-0366a870d238",
"segment_type": "Signature",
"markdown": "## Image Description\n\nThe image shows a handwritten word in dark ink on a light background.\n\n### Visible Text\n- **Dhote.**\n\n### Details\n- The handwriting is cursive and slightly slanted...",
"html": "
Image Description
\n
The image shows a handwritten word in dark ink on a light background.
\n
Visible Text
\n
Dhote.
...",
"bbox": { "left": 96.7, "top": 1398.4, "width": 84.2, "height": 50.8 },
"page_number": 1,
"page_width": 1191.0,
"page_height": 1684.0,
"confidence": 1.0
}
```
For the full segment shape and configuration options, see the [Parse API reference](/docs/api-reference/parser/parse-document).
# Excel Parsing
Source: https://docs.unsiloed.ai/docs/document-processing/parsing/excel
Parse .xls and .xlsx workbooks into structured tables with cell-range references.
Spreadsheets don't behave like documents. There are no pages to lay out, no reading order to reconstruct, and no OCR to run — a workbook is a grid of typed cells spread across one or more sheets, often with hidden rows, formulas, merged ranges, and styling that carry meaning. Excel parsing has its own pipeline and its own endpoint: **`POST /parse/excel`**.
Excel files go to `POST /parse/excel`, not `POST /parse`. Submitting a `.xls` or `.xlsx` to `/parse` returns a `400` with `error: "excel_not_supported_here"` pointing you here. Conversely, `/parse/excel` only accepts Excel files and rejects everything else.
## How It Differs from `/parse`
The Excel endpoint shares the same asynchronous submit-and-poll flow and the same result shape as `/parse` — you get a `job_id`, poll `GET /parse/{job_id}` until `Succeeded`, and read back chunks of segments. But none of the PDF parsing options carry over: `ocr_engine`, `layout_analysis`, `segment_processing`, `agentic_ocr`, and the [processing modes](/docs/document-processing/parsing/processing-modes) have no effect on a workbook. Excel parsing has its own, separate set of configuration options (below).
Each sheet is converted to structured tables. The parser preserves the grid, lets you drop hidden content, controls how large tables are split, and tags every table segment with the spreadsheet cells it came from.
## Submitting a Workbook
Provide either a `file` upload or a `url` to a workbook in cloud storage — not both. Every configuration field is optional and defaults to the pipeline's own default.
```python Python theme={null}
import json
import os
import time
import requests
API_KEY = os.environ["UNSILOED_API_KEY"]
BASE_URL = "https://prod.visionapi.unsiloed.ai"
with open("workbook.xlsx", "rb") as f:
response = requests.post(
f"{BASE_URL}/parse/excel",
headers={"api-key": API_KEY},
files={"file": ("workbook.xlsx", f,
"application/vnd.openxmlformats-officedocument.spreadsheetml.sheet")},
data={
"exclude_hidden": "true", # drop hidden sheets/rows/cols
"split_large_tables": "true",
"max_rows_per_segment": "50",
},
)
response.raise_for_status()
job_id = response.json()["job_id"]
print(f"Job submitted: {job_id}")
while True:
result = requests.get(
f"{BASE_URL}/parse/{job_id}",
headers={"api-key": API_KEY},
).json()
if result["status"] == "Succeeded":
break
if result["status"] == "Failed":
raise RuntimeError(result.get("message", "parse job failed"))
time.sleep(5)
with open("result.json", "w") as f:
json.dump(result, f, indent=2)
```
```javascript JavaScript theme={null}
import fs from "node:fs";
const API_KEY = process.env.UNSILOED_API_KEY;
const BASE_URL = "https://prod.visionapi.unsiloed.ai";
const form = new FormData();
form.append("file", new Blob([fs.readFileSync("workbook.xlsx")]), "workbook.xlsx");
form.append("exclude_hidden", "true");
form.append("split_large_tables", "true");
form.append("max_rows_per_segment", "50");
const submit = await fetch(`${BASE_URL}/parse/excel`, {
method: "POST",
headers: { "api-key": API_KEY },
body: form,
});
const { job_id } = await submit.json();
console.log(`Job submitted: ${job_id}`);
let result;
while (true) {
result = await (
await fetch(`${BASE_URL}/parse/${job_id}`, { headers: { "api-key": API_KEY } })
).json();
if (result.status === "Succeeded") break;
if (result.status === "Failed") throw new Error(result.message ?? "parse job failed");
await new Promise((r) => setTimeout(r, 5000));
}
```
```bash cURL theme={null}
# Submit
curl -X POST "https://prod.visionapi.unsiloed.ai/parse/excel" \
-H "api-key: $UNSILOED_API_KEY" \
-F "file=@workbook.xlsx" \
-F "exclude_hidden=true" \
-F "split_large_tables=true" \
-F "max_rows_per_segment=50"
# Poll (substitute the job_id from the submit response)
curl "https://prod.visionapi.unsiloed.ai/parse/$JOB_ID" \
-H "api-key: $UNSILOED_API_KEY"
```
## Configuration
All fields are optional. Booleans are sent as the strings `"true"` / `"false"` in the multipart form.
### Hidden Content
| Parameter | Type | Default | What it does |
| - | - | - | - |
| `exclude_hidden` | boolean | `false` | Drop hidden sheets, rows, columns, and styling from the output (gates the four fields below). |
| `exclude_hidden_sheets` | boolean | `true` | When excluding hidden content, also drop hidden sheets. |
| `exclude_hidden_rows` | boolean | `true` | When excluding hidden content, also drop hidden rows. |
| `exclude_hidden_cols` | boolean | `true` | When excluding hidden content, also drop hidden columns. |
| `exclude_images` | boolean | `false` | When excluding hidden content, also drop embedded/pasted images. |
### Tables
| Parameter | Type | Default | What it does |
| - | - | - | - |
| `split_large_tables` | boolean | `true` | Break big tables into smaller segments so each stays a manageable size. |
| `max_rows_per_segment` | integer | `50` | Maximum rows per segment when `split_large_tables` is enabled. |
| `table_clustering` | string | `"accurate"` | How aggressively to group adjacent ranges into tables: `accurate` (full analysis, best but slower), `fast`, or `off`. |
The `exclude_hidden_*` and `exclude_images` toggles only take effect when `exclude_hidden` is `true`. Leave `exclude_hidden` off to keep everything.
## What an Excel Parse Returns
The response is the same job/chunk/segment shape as a PDF parse (see [Response Format](/docs/document-processing/parsing/response-format)), with workbook content surfaced as `Table` segments. Each sheet's tables come back as `markdown` and `html`, and large tables are split according to `split_large_tables` / `max_rows_per_segment`.
The Excel-specific addition is **`cell_references`** on each segment:
* **`cell_references`:** spreadsheet cell-range references for the segment, each an object of `{ sheet, address, ref }` — the sheet name, the cell or range address (e.g. `Sheet1!B2:D10`), and the referenced value. This is how you trace a parsed table back to exact cells in the workbook.
`cell_references` is distinct from `references`, which carries research-paper citations for document parses. On Excel segments, spreadsheet cell ranges live in `cell_references`; `references` stays `null`.
## Dig Deeper
The canonical job/chunk/segment shape shared with PDF parses.
Parse workbooks straight from cloud storage with a `url` instead of an upload.
# Parse Overview
Source: https://docs.unsiloed.ai/docs/document-processing/parsing/parsing
Turn documents into ordered Markdown chunks with bounding boxes, OCR, and labeled layout segments.
Just want to run a parse? The [Quickstart](/docs/quickstart) walks through the submit-and-poll flow end to end in Python, JavaScript, and cURL.
If you don't yet know what's important in a document, parsing is where you start. You hand the `/parse` endpoint a document (PDF, image, spreadsheet, or Office document) and it returns the whole thing broken into labeled layout regions: paragraphs, tables, figures, titles, captions, list items, headers, and footers. Tables stay in one piece, and reading order is preserved across multi-column layouts. That makes parsing the right operation for loading documents into a search index, an embedding store for RAG, or a long-term archive.
For pulling specific fields out by name (invoice totals, contract dates, lab values), use [Extraction](/docs/document-processing/extraction/extraction) instead.
Parsing a spreadsheet? Excel workbooks have their own pipeline and endpoint; see [Excel Parsing](/docs/document-processing/parsing/excel). Submitting a `.xls`/`.xlsx` to `/parse` returns a `400` pointing you to `POST /parse/excel`.
## From a Dense Fund Filing to Structured Markdown
The page below is a real document we ran through `/parse`: a Schedule of Investments from the Calamos mutual funds' 2023 annual report. It's the kind of page standard parsers choke on, with two columns of holdings, each split into nested sub-tables (convertible bonds, sector-grouped preferred stocks, a government security, purchased options, and two forward foreign-currency-contract tables), threaded with footnote symbols, subtotals, and running totals. The filing also ships with a scrambled embedded text layer, so we parse it with [`ocr_strategy=force_ocr`](/docs/api-reference/parser/parse-document) to take the text from OCR of the rendered page instead of the corrupt layer. The layout model handles the column structure itself: it detects the two columns and orders segments top to bottom within each, and for layouts that still come back out of order, the [`enhance_reading_order`](/docs/api-reference/parser/parse-document) parameter re-sorts the segments explicitly. Parsing returns each column as its own table, with the sector groups, footnote markers, and totals intact.
Here's a slice of the Markdown that comes back, showing the sector-grouped convertible preferred stocks with footnote symbols and subtotals preserved:
```markdown theme={null}
| NUMBER OF SHARES | | VALUE |
| --- | --- | --- |
| **CONVERTIBLE PREFERRED STOCKS (3.4%)** | | |
| | Financials (2.0%) | |
| 127,810 | Apollo Global Management, Inc. 6.750%, 07/31/26 | 6,148,939 |
| 10,300 | Bank of America Corp.~‡‡ 7.250% | 10,847,960 |
| | | 16,996,899 |
| | Industrials (0.5%) | |
| 86,940 | Chart Industries, Inc. 6.750%, 12/15/25 | 4,273,971 |
| | Utilities (0.9%) | |
| 187,200 | NextEra Energy, Inc.~ 6.926%, 09/01/25 | 7,027,488 |
| | **TOTAL CONVERTIBLE PREFERRED STOCKS** (Cost $32,846,094) | 28,298,358 |
```
## What a Parse Returns
A successful parse organizes the document into **chunks**, each containing an array of **segments**:
* A **segment** is a single labeled layout region: one paragraph, one table, one image. Each segment carries a bounding box, a `segment_type` (`Text`, `Table`, `Picture`, and others), a confidence score, and the region rendered as Markdown, HTML, and plain text. Tables and pictures also include word-level OCR boxes and a signed image URL for the cropped region.
* A **chunk** groups one or more adjacent segments and includes an `embed` field with the chunk's Markdown rolled into one string, ready to drop straight into a vector store.
See [Response Format](/docs/document-processing/parsing/response-format) for the full field reference, and [Element Types](/docs/document-processing/parsing/element-types) for the complete list of `segment_type` values.
Parsing forms or applications? Pass [`segmentation_version=v2`](/docs/api-reference/parser/parse-document#segmentation-version) to run a form-aware layout model. It adds `KeyValuePair` segments that keep each printed label with the value filled in against it — including checkbox state — instead of scattering labels and values across separate `Text` segments.
## How a Parse Job Works
`/parse` is asynchronous. You submit the document and get back a `job_id`, then poll `GET /parse/{job_id}` until the status is `Succeeded`. Behind the scenes the parser runs OCR on every region, classifies each region's type with a vision model, and reconstructs the reading order across the page.
The [Quickstart](/docs/quickstart) walks through the submit-and-poll flow in Python, JavaScript, and cURL. If your documents live in cloud storage (S3, GCS, Azure Blob, Supabase), you can skip the upload step and point `/parse` at a [presigned URL](/docs/document-processing/parsing/presigned-urls) instead.
## Dig Deeper
The full list of `segment_type` values the parser returns.
The canonical response shape with a field-by-field reference.
Preset parameter bundles (Fast, Accurate, Agentic Lite, Agentic) and the parameters they set.
Parse documents straight from cloud storage.
The dedicated `/parse/excel` endpoint for `.xls` and `.xlsx` workbooks.
For the full request and response specification, see the [Parse API reference](/docs/api-reference/parser/parse-document).
# Presigned URLs
Source: https://docs.unsiloed.ai/docs/document-processing/parsing/presigned-urls
Parse documents from a URL instead of uploading them in the request body.
Instead of uploading a file with multipart form data, you can point `/parse` at a presigned URL using the `url` field. This is the right pattern when your documents already live in cloud storage (S3, GCS, Azure Blob, Supabase, etc.) — no need to download them locally first.
The submit-and-poll flow is identical to the upload path; only the request body differs.
```python Python theme={null}
import os
import requests
import time
API_KEY = os.environ["UNSILOED_API_KEY"]
BASE_URL = "https://prod.visionapi.unsiloed.ai"
document_url = "https://example.com/path/to/your/document.pdf"
response = requests.post(
f"{BASE_URL}/parse",
headers={"api-key": API_KEY},
data={"url": document_url},
)
response.raise_for_status()
job_id = response.json()["job_id"]
print(f"Job submitted: {job_id}")
while True:
result = requests.get(
f"{BASE_URL}/parse/{job_id}",
headers={"api-key": API_KEY},
).json()
print(f"Status: {result['status']}")
if result["status"] == "Succeeded":
break
if result["status"] == "Failed":
raise RuntimeError(result.get("message", "parse job failed"))
time.sleep(5)
print(f"Total chunks: {result['total_chunks']}")
for chunk in result["chunks"]:
print(chunk["embed"][:100])
```
```javascript JavaScript theme={null}
const API_KEY = process.env.UNSILOED_API_KEY;
const BASE_URL = "https://prod.visionapi.unsiloed.ai";
const documentUrl = "https://example.com/path/to/your/document.pdf";
const form = new FormData();
form.append("url", documentUrl);
const response = await fetch(`${BASE_URL}/parse`, {
method: "POST",
headers: { "api-key": API_KEY },
body: form,
});
if (!response.ok) throw new Error(`${response.status}: ${await response.text()}`);
const { job_id } = await response.json();
console.log(`Job submitted: ${job_id}`);
let result;
while (true) {
const res = await fetch(`${BASE_URL}/parse/${job_id}`, {
headers: { "api-key": API_KEY },
});
result = await res.json();
console.log(`Status: ${result.status}`);
if (result.status === "Succeeded") break;
if (result.status === "Failed") throw new Error(result.message || "parse job failed");
await new Promise((r) => setTimeout(r, 5000));
}
console.log(`Total chunks: ${result.total_chunks}`);
for (const chunk of result.chunks) {
console.log(chunk.embed.slice(0, 100));
}
```
```bash cURL theme={null}
curl -X POST "https://prod.visionapi.unsiloed.ai/parse" \
-H "api-key: $UNSILOED_API_KEY" \
-F "url=https://example.com/path/to/your/document.pdf"
# Then poll GET /parse/{job_id} until "status" is "Succeeded".
```
For the full set of parameters accepted by `/parse`, see the [Parse API reference](/docs/api-reference/parser/parse-document).
# Processing Modes
Source: https://docs.unsiloed.ai/docs/document-processing/parsing/processing-modes
Preset parameter bundles for the /parse endpoint: Fast, Accurate, Agentic Lite, and Agentic.
Instead of configuring `ocr_engine`, `segment_processing`, `agentic_ocr`, and `merge_tables` individually, you can apply a preset mode that bundles a recommended combination of all four. The parser ships with four:
Fast
Optimized for speed. Best for standard documents with clean text and simple tables.
Accurate
Balanced accuracy and performance. Ideal for complex layouts and detailed tables.
Agentic Lite
An agent reviews each segment, flags low-confidence text, and re-reads it against the source to resolve errors at faster speeds. Best for text-heavy documents.
Agentic
An agent reviews every segment, reasons over ambiguous or low-confidence text, and re-reads until resolved for maximum accuracy. Best for critical documents requiring the highest fidelity.
## Choosing a Mode
The four modes trade speed for fidelity. The response shape is the same across them in most cases, though jobs submitted with `merge_tables` (Agentic Lite and Agentic) can return a cached cross-page merged result instead (see the [parse job status reference](/docs/api-reference/parser/get-parse-job-status)). Start with Fast and step up only when the output misses detail your pipeline depends on.
| Mode | Speed | Table handling | Reach for it when |
| - | - | - | - |
| **Fast** | Fastest | Single-pass tables | Documents are clean and digital: typed invoices, receipts, and simple tables. |
| **Accurate** | Moderate | Higher-fidelity single-pass tables | Layouts are complex or multi-column, but tables stay within a page. |
| **Agentic Lite** | Slower | Cross-page tables merged; each segment agentically re-read | Text-heavy documents that need verification, including tables spanning pages. |
| **Agentic** | Slowest | Cross-page tables merged; every segment re-read and reasoned over | Every value has to be right: contracts, regulatory filings, and low-quality scans. |
Cross-page table merging (`merge_tables`) is on only for **Agentic Lite** and **Agentic** — Fast and Accurate leave it off. If a table spans pages and you need it stitched into one segment, use an Agentic mode or pass `merge_tables=true` alongside your chosen mode's parameters.
## Mode Parameters
Each mode is shorthand for a specific set of parameters. To use a mode, pass the values from its column in your `/parse` request.
| Parameter | Fast | Accurate | Agentic Lite | Agentic |
| - | - | - | - | - |
| `ocr_engine` | `UnsiloedBeta` | `UnsiloedBeta` | `UnsiloedBeta` | `UnsiloedBeta` |
| `segment_processing.Table.html` | `VLM` | `VLM` | `VLM` | `VLM` |
| `segment_processing.Table.model_id` | `astra_lite` | `astra_v2` | `astra` | `astra_v3` |
| `agentic_ocr` | `null` | `null` | `standard` | `advanced` |
| `merge_tables` | `false` | `false` | `true` | `true` |
Modes are preset configurations. You can override any individual parameter after applying one. See the [Parse API reference](/docs/api-reference/parser/parse-document) for the full parameter list.
Modes don't set `segmentation_version` — it's orthogonal and composes with all four. Add `segmentation_version=v2` to any mode above when the documents are forms and you want `KeyValuePair` segments. See [Segmentation Version](/docs/api-reference/parser/parse-document#segmentation-version).
## Applying a Mode
The submit-and-poll flow is the same as in the [Quickstart](/docs/quickstart); only the request body differs. The snippets below show each mode in Python — translate to JavaScript or cURL by changing the request body shape, not the parameter values.
```python Fast theme={null}
import os
import requests
API_KEY = os.environ["UNSILOED_API_KEY"]
BASE_URL = "https://prod.visionapi.unsiloed.ai"
with open("document.pdf", "rb") as f:
response = requests.post(
f"{BASE_URL}/parse",
headers={"api-key": API_KEY},
files={"file": ("document.pdf", f, "application/pdf")},
data={
"ocr_engine": "UnsiloedBeta",
"merge_tables": "false",
"agentic_ocr": "",
"segment_processing": '{"Table": {"html": "VLM", "model_id": "astra_lite"}}',
},
)
response.raise_for_status()
job_id = response.json()["job_id"]
print(f"Job submitted: {job_id}")
```
```python Accurate theme={null}
import os
import requests
API_KEY = os.environ["UNSILOED_API_KEY"]
BASE_URL = "https://prod.visionapi.unsiloed.ai"
with open("document.pdf", "rb") as f:
response = requests.post(
f"{BASE_URL}/parse",
headers={"api-key": API_KEY},
files={"file": ("document.pdf", f, "application/pdf")},
data={
"ocr_engine": "UnsiloedBeta",
"merge_tables": "false",
"agentic_ocr": "",
"segment_processing": '{"Table": {"html": "VLM", "model_id": "astra_v2"}}',
},
)
response.raise_for_status()
job_id = response.json()["job_id"]
print(f"Job submitted: {job_id}")
```
```python Agentic Lite theme={null}
import os
import requests
API_KEY = os.environ["UNSILOED_API_KEY"]
BASE_URL = "https://prod.visionapi.unsiloed.ai"
with open("document.pdf", "rb") as f:
response = requests.post(
f"{BASE_URL}/parse",
headers={"api-key": API_KEY},
files={"file": ("document.pdf", f, "application/pdf")},
data={
"ocr_engine": "UnsiloedBeta",
"merge_tables": "true",
"agentic_ocr": "standard",
"segment_processing": '{"Table": {"html": "VLM", "model_id": "astra"}}',
},
)
response.raise_for_status()
job_id = response.json()["job_id"]
print(f"Job submitted: {job_id}")
```
```python Agentic theme={null}
import os
import requests
API_KEY = os.environ["UNSILOED_API_KEY"]
BASE_URL = "https://prod.visionapi.unsiloed.ai"
with open("document.pdf", "rb") as f:
response = requests.post(
f"{BASE_URL}/parse",
headers={"api-key": API_KEY},
files={"file": ("document.pdf", f, "application/pdf")},
data={
"ocr_engine": "UnsiloedBeta",
"merge_tables": "true",
"agentic_ocr": "advanced",
"segment_processing": '{"Table": {"html": "VLM", "model_id": "astra_v3"}}',
},
)
response.raise_for_status()
job_id = response.json()["job_id"]
print(f"Job submitted: {job_id}")
```
# Response Format
Source: https://docs.unsiloed.ai/docs/document-processing/parsing/response-format
Canonical response shape from the /parse endpoint, with a field-by-field reference.
A successful `/parse` job returns the document organized into chunks. Each chunk has an `embed` Markdown string (concatenated content from its segments, ready for embedding) and an array of segments with bounding boxes and metadata. The example below is a real response from a single-page test document.
```json theme={null}
{
"job_id": "a0f51f79-6eb8-412a-9afa-924ddfbf9578",
"status": "Succeeded",
"message": "Task succeeded",
"file_name": "document.pdf",
"file_type": "application/pdf",
"page_count": 1,
"total_chunks": 1,
"credit_used": 1,
"merge_tables": false,
"created_at": "2026-05-22T11:29:17.964433Z",
"started_at": "2026-05-22T11:29:18.040527Z",
"finished_at": "2026-05-22T11:29:32.109418Z",
"pdf_url": "https://s3.us-east-1.amazonaws.com/...",
"file_url": "https://s3.us-east-1.amazonaws.com/...",
"configuration": {
"layout_analysis": "smart_layout_detection",
"ocr_engine": "UnsiloedHawk",
"ocr_strategy": "auto_detection",
"merge_tables": false,
"...": "..."
},
"metadata": {},
"chunks": [
{
"chunk_id": "6b2eca3a-d14f-4164-ba9a-0a3a58fcaf45",
"chunk_length": 117,
"embed": "## Q1 2024 Sales Report\nThe following table summarises regional sales performance...",
"segments": [
{
"segment_id": "034a37e7-6e4b-45dd-802c-e648d6c16498",
"segment_type": "SectionHeader",
"content": "Q1 2024 Sales Report",
"markdown": "## Q1 2024 Sales Report",
"html": "
",
"image": "https://s3.us-east-1.amazonaws.com/...",
"bbox": { "left": 54.4, "top": 208.5, "width": 1026.5, "height": 246.5 },
"page_number": 1,
"page_width": 1191.0,
"page_height": 1684.0,
"confidence": 0.99,
"ocr": [],
"references": null
}
]
}
]
}
```
## Top-Level Fields
These fall into three groups: identification, status, and timing; parsed content; and job configuration and metering.
### Identification, Status, and Timing
* **`job_id`:** unique identifier for the parsing job
* **`status`:** job state (`Succeeded`, `Failed`, or an in-progress value such as `Starting` or `Processing`)
* **`message`:** human-readable status message (`"Task succeeded"` when the job completes)
* **`file_name`:** name of the uploaded file
* **`file_type`:** MIME type of the uploaded file (e.g., `application/pdf`)
* **`created_at`:** ISO 8601 timestamp when the job was created
* **`started_at`:** ISO 8601 timestamp when processing began
* **`finished_at`:** ISO 8601 timestamp when processing completed
### Parsed Content
* **`chunks`:** array of content chunks
* **`total_chunks`:** total number of chunks
* **`page_count`:** total number of pages in the document
* **`pdf_url`:** temporary signed S3 URL to the processed PDF, or `null` unless `include_url=true` (see [URL fields](#url-fields) below)
* **`file_url`:** temporary signed S3 URL to the original uploaded file, or `null` unless `include_url=true`
### Job Configuration and Metering
* **`configuration`:** the full configuration object used for this parse (OCR engine, layout strategy, segment processing settings, etc.); see the [Parse API reference](/docs/api-reference/parser/parse-document) for every option
* **`metadata`:** additional job metadata; usually an empty object
* **`merge_tables`:** whether tables were merged across pages
* **`credit_used`:** credits consumed by this job
## Chunk Fields
* **`chunk_id`:** unique identifier for the chunk
* **`chunk_length`:** character length of the chunk's `embed` content
* **`embed`:** combined Markdown content from all segments in the chunk, ready for embedding into a vector store
* **`segments`:** array of layout segments within the chunk
## Segment Fields
* **`segment_id`:** unique identifier for the segment
* **`segment_type`:** element classification; see the [Element Types](/docs/document-processing/parsing/element-types) reference for the full list
* **`content`:** plain-text content of the segment (omitted for `Signature` segments)
* **`markdown`:** Markdown-formatted content
* **`html`:** HTML-formatted content
* **`image`:** signed S3 URL to a cropped image of the segment (present for most types; omitted for `Signature`), or `null` unless `include_url=true` (see [URL fields](#url-fields) below)
* **`bbox`:** bounding box relative to the page, with `left`, `top`, `width`, `height` in render pixels
* **`page_number`:** page where the segment appears
* **`page_width` / `page_height`:** dimensions in pixels of the rendered page the bounding boxes are measured against; use the ratio of `page_width` to the page's width in PDF points to convert coordinates back to points
* **`confidence`:** model confidence score (0–1) for element detection
* **`ocr`:** array of word-level OCR results
* **`references`:** references to related segments; typically `null`
## OCR Item Fields
Each item in a segment's `ocr` array describes one word the OCR engine recognized within that segment. The bounding box is relative to the segment's cropped image, not the full page.
* **`text`:** the recognized word or token
* **`bbox`:** bounding box relative to the segment's image, with `left`, `top`, `width`, `height`
* **`confidence`:** per-word model confidence (0–1), or `null` when not reported
* **`color`:** optional `r`, `g`, `b`, and `hex` sub-fields, present only when `extract_colors: true` is set in the parse configuration
## URL Fields
By default, every file URL in the response is returned as `null` so the response (and any log that captures it) never exposes your storage bucket, region, or path. The gated fields are:
* `pdf_url`
* `file_url`
* `output_file_url`
* `exports` (presigned export download URLs)
* segment `image` (cropped segment images)
* `configuration.input_file_url`
To receive the real URLs, opt in when polling for results with either the `include_url=true` query parameter or the `include-url: true` header on `GET /parse/{job_id}`:
```bash cURL theme={null}
curl "https://prod.visionapi.unsiloed.ai/parse/$JOB_ID?include_url=true" \
-H "api-key: $UNSILOED_API_KEY"
```
```python Python theme={null}
result = requests.get(
f"{BASE_URL}/parse/{job_id}",
params={"include_url": "true"},
headers={"api-key": API_KEY},
).json()
```
`include_url` does not rewrite or re-sign URLs; when set to `true` they are returned exactly as generated. Presigned URLs are time-limited, so fetch any files you need promptly.
# Getting Started With Splitting
Source: https://docs.unsiloed.ai/docs/document-processing/splitting/quickstart
Submit a bundled PDF to /splitter and download one PDF per matched category.
Mail rooms, AP scans, and patient intake forms often arrive as one big PDF with several different documents stacked together. The `/splitter` endpoint takes that bundle and a list of categories, then returns one labeled PDF per matched category. For other endpoints, see the [Parse quickstart](/docs/quickstart), [Extraction quickstart](/docs/document-processing/extraction/quickstart), or [Classification quickstart](/docs/document-processing/classification/quickstart).
This walkthrough builds a script that ships a bundled PDF and a category list to `/splitter`, waits for the split to complete, and writes one labeled PDF per matched category into a local `split_files/` directory, ready to drop into per-category downstream pipelines. If you'd rather just copy the whole script, it's in the dropdown below.
Set `UNSILOED_API_KEY` in your environment and save the bundled PDF as `bundle.pdf` in the same directory before running.
```python split_bundle.py theme={null}
import json
import os
import time
import requests
API_KEY = os.environ["UNSILOED_API_KEY"]
BASE_URL = "https://prod.visionapi.unsiloed.ai"
categories = [
{"name": "Invoice", "description": "Vendor invoices with itemized charges and a total due"},
{"name": "Receipt", "description": "Point-of-sale receipts with line items, tax, and payment method"},
{"name": "Purchase Order", "description": "Buyer-issued purchase orders authorizing goods or services from a vendor"},
]
with open("bundle.pdf", "rb") as f:
response = requests.post(
f"{BASE_URL}/splitter",
headers={"api-key": API_KEY},
files={"file": ("bundle.pdf", f, "application/pdf")},
data={"categories": json.dumps(categories)},
)
response.raise_for_status()
job_id = response.json()["job_id"]
print(f"Job submitted: {job_id}")
max_attempts = 60 # roughly 5 minutes at 5 seconds per poll
attempts = 0
while True:
result = requests.get(
f"{BASE_URL}/splitter/{job_id}",
headers={"api-key": API_KEY},
).json()
print(f"Status: {result['status']}")
if result["status"] == "completed":
break
if result["status"] == "failed":
raise RuntimeError(result.get("error", "split job failed"))
attempts += 1
if attempts >= max_attempts:
raise TimeoutError("Split job did not finish within 5 minutes")
time.sleep(5)
with open("result.json", "w") as f:
json.dump(result, f, indent=2)
os.makedirs("split_files", exist_ok=True)
for file_info in result["result"]["files"]:
pdf_bytes = requests.get(file_info["full_path"]).content
out_path = os.path.join("split_files", file_info["name"])
with open(out_path, "wb") as out:
out.write(pdf_bytes)
print(f"Saved {out_path} (confidence={file_info['confidence_score']:.2%})")
```
Save this as `script.mjs` or set `"type": "module"` in your `package.json`. Requires Node.js 18 or newer for the global `fetch`, `FormData`, and `Blob`.
```javascript script.mjs theme={null}
import fs from "node:fs";
import path from "node:path";
const API_KEY = process.env.UNSILOED_API_KEY;
const BASE_URL = "https://prod.visionapi.unsiloed.ai";
const categories = [
{ name: "Invoice", description: "Vendor invoices with itemized charges and a total due" },
{ name: "Receipt", description: "Point-of-sale receipts with line items, tax, and payment method" },
{ name: "Purchase Order", description: "Buyer-issued purchase orders authorizing goods or services from a vendor" },
];
const form = new FormData();
form.append("file", new Blob([fs.readFileSync("bundle.pdf")]), "bundle.pdf");
form.append("categories", JSON.stringify(categories));
const response = await fetch(`${BASE_URL}/splitter`, {
method: "POST",
headers: { "api-key": API_KEY },
body: form,
});
if (!response.ok) throw new Error(`${response.status}: ${await response.text()}`);
const { job_id } = await response.json();
console.log(`Job submitted: ${job_id}`);
const maxAttempts = 60; // roughly 5 minutes at 5 seconds per poll
let attempts = 0;
let result;
while (true) {
const res = await fetch(`${BASE_URL}/splitter/${job_id}`, {
headers: { "api-key": API_KEY },
});
result = await res.json();
console.log(`Status: ${result.status}`);
if (result.status === "completed") break;
if (result.status === "failed") throw new Error(result.error || "split job failed");
if (++attempts >= maxAttempts) throw new Error("Split job did not finish within 5 minutes");
await new Promise((r) => setTimeout(r, 5000));
}
fs.writeFileSync("result.json", JSON.stringify(result, null, 2));
fs.mkdirSync("split_files", { recursive: true });
for (const file of result.result.files) {
const res = await fetch(file.full_path);
const buf = Buffer.from(await res.arrayBuffer());
const outPath = path.join("split_files", file.name);
fs.writeFileSync(outPath, buf);
console.log(`Saved ${outPath} (confidence=${(file.confidence_score * 100).toFixed(2)}%)`);
}
```
```bash theme={null}
# Submit the bundle and capture the job_id from the response:
resp=$(curl -sX POST "https://prod.visionapi.unsiloed.ai/splitter" \
-H "api-key: $UNSILOED_API_KEY" \
-F "file=@bundle.pdf" \
-F 'categories=[{"name":"Invoice","description":"Vendor invoices with itemized charges and a total due"},{"name":"Receipt","description":"Point-of-sale receipts with line items, tax, and payment method"},{"name":"Purchase Order","description":"Buyer-issued purchase orders authorizing goods or services from a vendor"}]')
JOB_ID=$(echo "$resp" | grep -o '"job_id":"[^"]*"' | cut -d'"' -f4)
echo "Job submitted: $JOB_ID"
# Poll until the job finishes, with a 5-minute timeout:
attempts=0
max_attempts=60
while true; do
resp=$(curl -sX GET "https://prod.visionapi.unsiloed.ai/splitter/$JOB_ID" \
-H "api-key: $UNSILOED_API_KEY")
status=$(echo "$resp" | grep -o '"status":"[^"]*"' | head -1 | cut -d'"' -f4)
echo "Status: $status"
[ "$status" = "completed" ] && break
[ "$status" = "failed" ] && { echo "Job failed"; exit 1; }
attempts=$((attempts + 1))
[ "$attempts" -ge "$max_attempts" ] && { echo "Split job did not finish within 5 minutes"; exit 1; }
sleep 5
done
# Save the full response and download each split file:
echo "$resp" > result.json
mkdir -p split_files
echo "$resp" \
| python3 -c "import json,sys; [print(f['name'], f['full_path']) for f in json.load(sys.stdin)['result']['files']]" \
| while read name url; do
curl -s "$url" -o "split_files/$name"
echo "Saved split_files/$name"
done
```
## Step 1: Set Up Your Environment
Before writing any code, we need three things: an API key, a bundled PDF, and the runtime for our chosen language.
### 1.1 Get an Unsiloed AI API Key
To get API access, [sign up on Unsiloed AI](https://cal.com/aman-mishra-p0ry57/15min). Export your key as an environment variable named `UNSILOED_API_KEY` so it stays out of source control:
```bash theme={null}
export UNSILOED_API_KEY="your-api-key"
```
### 1.2 Pick a Bundled PDF
The `/splitter` endpoint is designed for PDFs that contain more than one logical document. This walkthrough assumes a multi-document PDF saved as `bundle.pdf` in your working directory.
If you don't have one handy, download our [sample bundle](https://raw.githubusercontent.com/Unsiloed-AI/cookbook/c585446e46e4be2790c6c29fe2a7a3a1b346191d/sample-documents/sample-split.pdf) (a three-page accounts-payable batch scan: an invoice, a receipt, and a purchase order) and save it as `bundle.pdf`.
### 1.3 Install Dependencies
You need Python 3.8 or newer. Install the `requests` package:
```bash theme={null}
pip install requests
```
You need Node.js 18 or newer for the global `fetch`, `FormData`, and `Blob`. No external packages needed.
You need cURL, which is preinstalled on macOS and most Linux distributions. We also use Python 3 for one short JSON-parsing line in the download loop.
## Step 2: Submit the Bundle
Two form fields go up: `file` for the bundled PDF and `categories` for a JSON-stringified array of the labels the splitter can choose from. The categories list is the only vocabulary the splitter uses. Pages that don't fit any category are still grouped under the closest match, so the list needs to cover everything that could plausibly appear in the bundle. The endpoint returns a `job_id` to poll. All requests go to `https://prod.visionapi.unsiloed.ai` with the API key in the `api-key` header.
### 2.1 Set Up the Script
Create a file called `split_bundle.py` and start with the imports and configuration:
```python split_bundle.py theme={null}
import json
import os
import time
import requests
API_KEY = os.environ["UNSILOED_API_KEY"]
BASE_URL = "https://prod.visionapi.unsiloed.ai"
```
`API_KEY` reads your key from the environment so it doesn't get hard-coded into the file, and `BASE_URL` points at the Unsiloed AI production endpoint. Both appear in every request below.
Create a file called `script.mjs` and start with the imports and configuration:
```javascript script.mjs theme={null}
import fs from "node:fs";
import path from "node:path";
const API_KEY = process.env.UNSILOED_API_KEY;
const BASE_URL = "https://prod.visionapi.unsiloed.ai";
```
`API_KEY` reads your key from the environment so it doesn't get hard-coded into the file, and `BASE_URL` points at the Unsiloed AI production endpoint. Both appear in every request below.
cURL doesn't need a setup step. Each command below inlines the API key and base URL directly.
### 2.2 Define the Categories
Decide which document types the bundle might contain. Each category is an object with a `name` and an optional `description`; richer descriptions help the splitter pick the right label when categories are similar.
Add the category list to the script:
```python split_bundle.py theme={null}
categories = [
{"name": "Invoice", "description": "Vendor invoices with itemized charges and a total due"},
{"name": "Receipt", "description": "Point-of-sale receipts with line items, tax, and payment method"},
{"name": "Purchase Order", "description": "Buyer-issued purchase orders authorizing goods or services from a vendor"},
]
```
Add the category list to the script:
```javascript script.mjs theme={null}
const categories = [
{ name: "Invoice", description: "Vendor invoices with itemized charges and a total due" },
{ name: "Receipt", description: "Point-of-sale receipts with line items, tax, and payment method" },
{ name: "Purchase Order", description: "Buyer-issued purchase orders authorizing goods or services from a vendor" },
];
```
For cURL we pass the categories inline as a JSON string on the request itself, so there's nothing to set up here. The submission command in the next step shows the full form.
Include every document type the bundle might contain. Pages that don't match any category are still grouped under the closest match.
### 2.3 Upload the Bundle
Send the file and categories as a multipart upload to `/splitter`. The endpoint expects the document under the form field name `file` and the categories as a JSON-encoded string under `categories`.
Continue the script by uploading the bundle:
```python split_bundle.py theme={null}
with open("bundle.pdf", "rb") as f:
response = requests.post(
f"{BASE_URL}/splitter",
headers={"api-key": API_KEY},
files={"file": ("bundle.pdf", f, "application/pdf")},
data={"categories": json.dumps(categories)},
)
response.raise_for_status()
```
The `raise_for_status()` call throws an `HTTPError` on any non-2xx response, so we don't need to check `.status_code` ourselves.
Continue the script by uploading the bundle:
```javascript script.mjs theme={null}
const form = new FormData();
form.append("file", new Blob([fs.readFileSync("bundle.pdf")]), "bundle.pdf");
form.append("categories", JSON.stringify(categories));
const response = await fetch(`${BASE_URL}/splitter`, {
method: "POST",
headers: { "api-key": API_KEY },
body: form,
});
if (!response.ok) throw new Error(`${response.status}: ${await response.text()}`);
```
`fetch` doesn't throw on non-2xx responses by default, so we check `response.ok` and raise the error ourselves.
Run:
```bash theme={null}
curl -X POST "https://prod.visionapi.unsiloed.ai/splitter" \
-H "api-key: $UNSILOED_API_KEY" \
-F "file=@bundle.pdf" \
-F 'categories=[{"name":"Invoice","description":"Vendor invoices with itemized charges and a total due"},{"name":"Receipt","description":"Point-of-sale receipts with line items, tax, and payment method"},{"name":"Purchase Order","description":"Buyer-issued purchase orders authorizing goods or services from a vendor"}]'
```
The response prints to stdout. We need the `job_id` field for the next step.
### 2.4 Capture the Job ID
Read and print the `job_id`:
```python split_bundle.py theme={null}
job_id = response.json()["job_id"]
print(f"Job submitted: {job_id}")
```
Run the script:
```bash theme={null}
python split_bundle.py
```
The output should be a single line like `Job submitted: 887f26e6-d089-47f6-8def-afe84de40ecd`.
Read and log the `job_id`:
```javascript script.mjs theme={null}
const { job_id } = await response.json();
console.log(`Job submitted: ${job_id}`);
```
Run the script:
```bash theme={null}
node script.mjs
```
The output should be a single line like `Job submitted: 887f26e6-d089-47f6-8def-afe84de40ecd`.
The response body from the POST above looks like:
```json theme={null}
{
"job_id": "887f26e6-d089-47f6-8def-afe84de40ecd",
"status": "processing",
"quota_remaining": 7698
}
```
Copy the `job_id` value; you'll paste it into the polling command in the next step.
## Step 3: Poll and Download the Split Files
The job runs asynchronously. We GET `/splitter/{job_id}` repeatedly until the status is `completed`, then download each split PDF using the signed URL in the response.
The status values the polling loop handles:
* **`completed`:** the split files are ready to download
* **`failed`:** the job errored; check the `error` field for details
* **`queued`:** the job is waiting to be picked up
* **`processing`:** the job is still running
### 3.1 Write the Polling Loop
Add a polling loop. The `max_attempts` cap stops the loop if the job hangs:
```python split_bundle.py theme={null}
max_attempts = 60 # roughly 5 minutes at 5 seconds per poll
attempts = 0
while True:
result = requests.get(
f"{BASE_URL}/splitter/{job_id}",
headers={"api-key": API_KEY},
).json()
print(f"Status: {result['status']}")
if result["status"] == "completed":
break
if result["status"] == "failed":
raise RuntimeError(result.get("error", "split job failed"))
attempts += 1
if attempts >= max_attempts:
raise TimeoutError("Split job did not finish within 5 minutes")
time.sleep(5)
```
Add a polling loop. The `maxAttempts` cap stops the loop if the job hangs:
```javascript script.mjs theme={null}
const maxAttempts = 60; // roughly 5 minutes at 5 seconds per poll
let attempts = 0;
let result;
while (true) {
const res = await fetch(`${BASE_URL}/splitter/${job_id}`, {
headers: { "api-key": API_KEY },
});
result = await res.json();
console.log(`Status: ${result.status}`);
if (result.status === "completed") break;
if (result.status === "failed") throw new Error(result.error || "split job failed");
if (++attempts >= maxAttempts) throw new Error("Split job did not finish within 5 minutes");
await new Promise((r) => setTimeout(r, 5000));
}
```
Replace `JOB_ID` below with the value you captured from Step 2.4, then run this loop. It polls every 5 seconds and gives up after 5 minutes if the job hasn't completed:
```bash theme={null}
JOB_ID="paste-job-id-here"
attempts=0
max_attempts=60 # roughly 5 minutes at 5 seconds per poll
while true; do
resp=$(curl -sX GET "https://prod.visionapi.unsiloed.ai/splitter/$JOB_ID" \
-H "api-key: $UNSILOED_API_KEY")
status=$(echo "$resp" | grep -o '"status":"[^"]*"' | head -1 | cut -d'"' -f4)
echo "Status: $status"
[ "$status" = "completed" ] && break
[ "$status" = "failed" ] && { echo "Job failed"; exit 1; }
attempts=$((attempts + 1))
[ "$attempts" -ge "$max_attempts" ] && { echo "Split job did not finish within 5 minutes"; exit 1; }
sleep 5
done
```
The loop keeps the latest response body in `$resp` for the next step.
### 3.2 Download the Split PDFs
Each entry in `result.result.files` has a presigned `full_path` URL that downloads the split PDF. The code below saves the metadata to `result.json` and writes each split file into a `split_files/` directory.
Add the download step:
```python split_bundle.py theme={null}
with open("result.json", "w") as f:
json.dump(result, f, indent=2)
os.makedirs("split_files", exist_ok=True)
for file_info in result["result"]["files"]:
pdf_bytes = requests.get(file_info["full_path"]).content
out_path = os.path.join("split_files", file_info["name"])
with open(out_path, "wb") as out:
out.write(pdf_bytes)
print(f"Saved {out_path} (confidence={file_info['confidence_score']:.2%})")
```
Run the script:
```bash theme={null}
python split_bundle.py
```
You should see a few `Status: processing` lines, then `Status: completed`, then one `Saved` line per matched category.
Add the download step:
```javascript script.mjs theme={null}
fs.writeFileSync("result.json", JSON.stringify(result, null, 2));
fs.mkdirSync("split_files", { recursive: true });
for (const file of result.result.files) {
const res = await fetch(file.full_path);
const buf = Buffer.from(await res.arrayBuffer());
const outPath = path.join("split_files", file.name);
fs.writeFileSync(outPath, buf);
console.log(`Saved ${outPath} (confidence=${(file.confidence_score * 100).toFixed(2)}%)`);
}
```
Run the script:
```bash theme={null}
node script.mjs
```
You should see a few `Status: processing` lines, then `Status: completed`, then one `Saved` line per matched category.
The polling loop in Step 3.1 left the full response in `$resp`. Save the metadata, then download each split file using the presigned `full_path` URLs:
```bash theme={null}
echo "$resp" > result.json
mkdir -p split_files
echo "$resp" \
| python3 -c "import json,sys; [print(f['name'], f['full_path']) for f in json.load(sys.stdin)['result']['files']]" \
| while read name url; do
curl -s "$url" -o "split_files/$name"
echo "Saved split_files/$name"
done
```
The `split_files/` directory now holds one PDF per matched category, named after it.
## Error Responses
Failures fall into two buckets: HTTP errors raised before the job is queued, and a `failed` status on a job that started but couldn't complete.
### HTTP Errors
The `/splitter` endpoint returns JSON error bodies with a `detail` field. The common cases:
* **`401 Unauthorized`:** body is `{"detail":"Invalid API key"}`. The `api-key` header is missing or wrong.
* **`400 Bad Request`:** body is `{"detail":"Invalid JSON format for categories: ..."}`. The `categories` form field isn't valid JSON.
* **`422 Unprocessable Entity`:** body is `{"detail":[{"type":"missing","loc":["body","categories"],"msg":"Field required","input":null}]}`. A required form field (usually `file` or `categories`) is missing.
* **`400 Bad Request`:** body is `{"detail":"At least one category is required"}`. The `categories` array is empty.
* **`400 Bad Request`:** body is `{"detail":"Failed to process file: Failed to get PDF page count: ..."}`. The upload isn't a readable PDF.
* **`404 Not Found`:** body is `{"detail":"Job not found"}`. The `job_id` you polled doesn't exist.
### Failed Jobs
A job that was accepted but couldn't be processed comes back with `status: "failed"` and a populated `error` field. The walkthrough's polling loop raises on this case so you see the message instead of waiting out the timeout:
```json theme={null}
{
"job_id": "7b31a7d7-e810-4a0b-931e-fbed0879bab2",
"status": "failed",
"progress": "Splitting failed",
"error": "Failed to classify pages",
"file_url": "https://example-bucket.s3.amazonaws.com/user_uploads/...",
"file_name": "bundle.pdf",
"parameters": { "classes": ["Invoice", "Receipt", "Purchase Order"], "page_count": 3 },
"result": null
}
```
## Response Shape
A completed response contains job metadata, an echo of the input `parameters`, and a `result.files[]` array with one entry per matched category. Each entry carries a presigned `full_path` URL we can download directly.
```json theme={null}
{
"job_id": "887f26e6-d089-47f6-8def-afe84de40ecd",
"status": "completed",
"progress": "Starting document splitting...",
"error": null,
"file_url": "https://example-bucket.s3.amazonaws.com/user_uploads/...",
"file_name": "bundle.pdf",
"parameters": {
"classes": ["Invoice", "Receipt", "Purchase Order"],
"page_count": 3,
"enable_reordering": false,
"category_descriptions": {
"Invoice": "Vendor invoices with itemized charges and a total due",
"Receipt": "Point-of-sale receipts with line items, tax, and payment method",
"Purchase Order": "Buyer-issued purchase orders authorizing goods or services from a vendor"
}
},
"result": {
"success": true,
"message": "Successfully split PDF into 3 files",
"files": [
{
"name": "Invoice.pdf",
"path": "Invoice.pdf",
"type": "file",
"fileId": "359afb3c-2554-4acd-9cb3-be4044d7ec97",
"full_path": "https://example-bucket.s3.amazonaws.com/files/...",
"confidence_score": 0.9999976308610644
},
{
"name": "Receipt.pdf",
"path": "Receipt.pdf",
"type": "file",
"fileId": "eca3f2db-e113-4b5f-92db-aab89d417114",
"full_path": "https://example-bucket.s3.amazonaws.com/files/...",
"confidence_score": 0.9999976308610644
},
{
"name": "Purchase Order.pdf",
"path": "Purchase Order.pdf",
"type": "file",
"fileId": "1af2c343-9a86-46c5-b8f0-c48cc04d1b6f",
"full_path": "https://example-bucket.s3.amazonaws.com/files/...",
"confidence_score": 0.9999976308610644
}
]
}
}
```
The fields fall into three broad categories:
**For downloading the split PDFs:**
* **`result.files[].full_path`:** presigned S3 URL to download the split PDF; this is what the walkthrough fetches into `split_files/`
* **`result.files[].name`:** filename derived from the matched category, suitable for saving to disk
* **`result.files[].confidence_score`:** the splitter's confidence in the classification, on a 0-1 scale; use it to flag low-confidence splits for human review
* **`result.files[].fileId`:** unique identifier for the split file, useful for tracking or deduplicating downstream
**For echoing the request back:**
* **`parameters.classes`:** the category names you submitted
* **`parameters.category_descriptions`:** the descriptions you submitted, keyed by category name
* **`parameters.page_count`:** number of pages in the uploaded PDF
* **`parameters.enable_reordering`:** whether the splitter reordered pages within each category after classification; defaults to `false`
* **`file_url`:** signed URL to the original uploaded bundle
* **`file_name`:** name of the uploaded bundle
**For job and progress tracking:**
* **`status`:** `completed`, `failed`, or an in-progress value such as `processing`
* **`progress`:** human-readable progress message
* **`error`:** error message if the job failed, otherwise `null`
* **`result.success`:** whether the split operation succeeded
* **`result.message`:** human-readable success or failure message
### Sample Output
Running the script against the [sample AP batch](https://raw.githubusercontent.com/Unsiloed-AI/cookbook/c585446e46e4be2790c6c29fe2a7a3a1b346191d/sample-documents/sample-split.pdf) produces a `split_files/` directory with one PDF per matched category:
```
split_files/
├── Invoice.pdf # page 1 — the Greenfield Print & Bindery invoice
├── Receipt.pdf # page 2 — the Cooper's Office Supply receipt
└── Purchase Order.pdf # page 3 — the Lighthouse Studios LLC purchase order
```
All three files report a confidence score above 99%. Open them to confirm each page of the bundle landed in the matching category file.
## Next Steps
For more on splitting, including the underlying classification step and the full response shape, see the [Splitting overview](/docs/document-processing/splitting/splitting) and the [Response Format](/docs/document-processing/splitting/response-format) reference.
Learn how the splitter groups pages and where to use it in a pipeline.
Classify a single document against candidate categories instead of splitting a bundle.
Browse the full request and response specs for the splitting endpoint.
Check limits, supported formats, and answers to common questions.
# Response Format
Source: https://docs.unsiloed.ai/docs/document-processing/splitting/response-format
Response shape from the /splitter endpoint, with a field-by-field reference.
A completed splitting job returns one entry per split file in `result.files`, each with a download URL and a confidence score for the classification. The example below is a real response from splitting a bundle of pages from a mutual fund annual report that contained a fund performance summary, a schedule of investments, and a financial statement.
```json theme={null}
{
"job_id": "c68959b4-87cc-4cd5-984e-47f14dfad6cb",
"status": "completed",
"progress": "Starting document splitting...",
"error": null,
"file_url": "https://example-bucket.s3.amazonaws.com/user_uploads/...?AWSAccessKeyId=...&Signature=...&Expires=...",
"file_name": "calamos-fund-report-bundle.pdf",
"parameters": {
"classes": ["Fund Performance", "Schedule of Investments", "Financial Statements"],
"page_count": 3,
"enable_reordering": false,
"category_descriptions": {
"Fund Performance": "Fund performance summary with growth chart and average annual total return table",
"Schedule of Investments": "Itemized list of portfolio holdings with values",
"Financial Statements": "Statements of assets and liabilities, operations, and changes in net assets"
}
},
"result": {
"success": true,
"message": "Successfully split PDF into 3 files",
"files": [
{
"name": "Fund Performance.pdf",
"path": "Fund Performance.pdf",
"type": "file",
"fileId": "cbe0c91c-8577-4bfb-a43a-b40f3932592c",
"full_path": "https://example-bucket.s3.amazonaws.com/files/...?AWSAccessKeyId=...&Signature=...&Expires=...",
"confidence_score": 0.9999979255377075
},
{
"name": "Schedule of Investments.pdf",
"path": "Schedule of Investments.pdf",
"type": "file",
"fileId": "dc0955b6-3bd7-4373-b356-cce3c8cf31f0",
"full_path": "https://example-bucket.s3.amazonaws.com/files/...?AWSAccessKeyId=...&Signature=...&Expires=...",
"confidence_score": 0.9999979255377075
},
{
"name": "Financial Statements.pdf",
"path": "Financial Statements.pdf",
"type": "file",
"fileId": "e41bccff-6200-484d-9e1e-92cdd02b48bc",
"full_path": "https://example-bucket.s3.amazonaws.com/files/...?AWSAccessKeyId=...&Signature=...&Expires=...",
"confidence_score": 0.9999979255377075
}
]
}
}
```
## Top-Level Fields
* **`job_id`:** unique identifier for the split job
* **`status`:** job state. `queued` while the job waits to be picked up, `processing` while it runs, then `completed` or `failed`.
* **`progress`:** human-readable progress message
* **`error`:** error message if the job failed, otherwise `null`
* **`file_url`:** presigned download URL for the original uploaded file. Expires roughly an hour after the response is generated. Re-issue `GET /splitter/{job_id}` to get a fresh URL.
* **`file_name`:** name of the original file
* **`parameters`:** echo of the parameters used for this job (described below)
* **`result`:** the split files (described below)
## Parameters Object
The `parameters` block echoes the inputs used for the split, including defaults applied by the API.
* **`classes`:** the category names you submitted
* **`page_count`:** number of pages in the uploaded document
* **`enable_reordering`:** whether the splitter reordered pages within each category after classification. Only applied to categories with more than one matched page. Defaults to `false`.
* **`category_descriptions`:** the descriptions you submitted for each category
## Result Object
* **`success`:** whether the split operation succeeded
* **`message`:** human-readable success or failure message
* **`files`:** array of generated split files, one per matched category. Pages sharing a category are grouped into a single file, so a bundle with several schedule-of-investments pages yields one Schedule of Investments.pdf containing all of them.
## File Object
Each entry in `result.files` represents one split document.
* **`name`:** name of the split file, derived from the matched category
* **`path`:** relative path to the file
* **`type`:** file type, always `"file"`
* **`fileId`:** unique identifier for the file
* **`full_path`:** presigned download URL for the split file. Expires roughly an hour after the response is generated, so download or copy the file straight away rather than storing the URL. To get a fresh URL, re-issue `GET /splitter/{job_id}`.
* **`confidence_score`:** classification confidence for this file (0–1), averaged across all pages assigned to the category
For the full request and response specification, see the [Split API reference](/docs/api-reference/splitting/split-document).
# Splitting Overview
Source: https://docs.unsiloed.ai/docs/document-processing/splitting/splitting
Split a bundled PDF into separate documents by category.
Splitting takes a single PDF that contains several documents bundled together and breaks them apart. Real-world bundles arrive every day (a scanned batch of receipts, a stack of student exams, a patient's referral packet) and downstream systems usually need each logical document on its own.
Given a bundled PDF and a list of candidate categories (the same shape `/classify` uses), `/splitter` returns one downloadable PDF per matched category, each with a confidence score. All pages assigned to the same category are collected into that category's file: a bundle with two invoices produces one Invoice.pdf containing both, not two separate files.
## How It Works
After you submit a PDF and categories, the API:
1. Analyzes every page of the bundle.
2. Classifies each page against your candidate categories.
3. Groups adjacent pages of the same type so multi-page documents stay together.
4. Generates one PDF per matched category, named after the category.
5. Returns a download URL and confidence score for each split file.
## Common Categories
Categories are whatever you define. Common groupings include:
* **Business:** invoices, receipts, purchase orders, contracts
* **Financial:** bank statements, financial reports, tax forms
* **Legal:** contracts, agreements, legal notices, compliance forms
* **Healthcare:** medical records, insurance forms, lab reports
* **HR:** resumes, employment forms, payroll documents
* **Academic:** research papers, reports, transcripts
## Dig Deeper
Submit a bundled PDF, define categories, and read back the split files.
Browse the canonical splitting response shape with a field-by-field reference.
For the full request and response specification, see the [Split API reference](/docs/api-reference/splitting/split-document).
# General FAQ
Source: https://docs.unsiloed.ai/docs/faq/general
Frequently asked questions about Unsiloed AI's document processing platform
Find answers to the most commonly asked questions about Unsiloed AI's document processing platform.
## What is Unsiloed AI?
Unsiloed AI is building the most accurate APIs for ingesting multimodal unstructured data like PDFs, PPT, DOCX, tables, charts, and images, and converting it into structured Markdown and JSON for downstream LLMs and AI Agents.
We specialize in processing complex documents with unprecedented accuracy and speed, powering vertical AI solutions across industries from startups to NASDAQ-listed enterprises.
## What types of content can Unsiloed AI process?
Our APIs can process multimodal unstructured data including:
* **Document Formats**: PDFs, PowerPoint (PPT/PPTX), Word (DOCX), images (PNG, JPEG, TIFF)
* **Visual Elements**: Tables, charts, diagrams, graphs, images
* **Text Content**: Paragraphs, headers, lists, captions, footnotes
* **Complex Layouts**: Multi-column documents, mixed content, scanned documents
**Common Use Cases:**
* **Financial Documents**: Annual reports, earnings statements, SEC filings, investment documents
* **Business Documents**: Invoices, receipts, contracts, presentations
* **Research Documents**: Academic papers, technical reports, whitepapers
* **Enterprise Content**: Internal documents, compliance forms, operational reports
## What output formats do you provide?
Unsiloed AI converts unstructured documents into structured formats optimized for LLMs and AI Agents:
* **Markdown**: Clean, structured markdown with preserved document hierarchy
* **JSON**: Structured JSON with extracted data, confidence scores, and bounding boxes
All outputs include metadata like element types, page numbers, and confidence scores for downstream processing.
## How accurate is Unsiloed AI?
Unsiloed AI delivers industry-leading accuracy for multimodal document processing:
* **Vision Models**: State-of-the-art models trained on diverse document types
* **Benchmarks**: Consistently outperforms LlamaIndex, Gemini, Mistral, and Unstructured.io on public benchmarks
* **Multimodal Understanding**: Superior handling of tables, charts, images, and mixed content
* **Production-Ready**: Trusted by startups and NASDAQ-listed enterprises processing hundreds of thousands of documents
Accuracy is validated through confidence scores and bounding box precision for every extracted element.
## Do you offer on-premise deployment?
Yes! We offer on-premise and self-hosted deployment options for enterprises with specific security, compliance, or data sovereignty requirements:
* **Self-Hosted Deployment**: Deploy Unsiloed AI infrastructure within your own cloud environment (AWS, Azure, GCP)
* **Air-Gapped Environments**: Support for air-gapped and isolated network deployments
* **Hybrid Solutions**: Combine cloud and on-premise deployments based on your needs
* **Full Control**: Maintain complete control over data processing and storage
* **Custom Integration**: Direct integration with your existing security and monitoring tools
* **Dedicated Support**: Enterprise support with custom SLAs
Contact us at [support@unsiloed.ai](mailto:support@unsiloed.ai) to discuss on-premise deployment options.
## How do I get started?
Getting started is easy:
1. **Sign up** - [Sign up on Unsiloed AI](https://cal.com/aman-mishra-p0ry57/15min) to get started
2. **Get API access** - We'll provide you with API keys to get started
3. **Try the demo** to test document processing with your own files
4. **Integrate** using our comprehensive API documentation
5. **Scale** as your needs grow with production-grade infrastructure
## What support is available?
We offer multiple support channels:
* **Documentation**: Comprehensive guides and API reference
* **Email Support**: [support@unsiloed.ai](mailto:support@unsiloed.ai)
* **Demo**: Try our interactive demo at [unsiloed.ai/demo](https://www.unsiloed.ai/demo)
* **Enterprise Support**: Dedicated support for enterprise customers
## Pricing and Plans
We offer flexible, usage-based pricing:
* **Pay-as-you-go**: Start with credits and scale as needed
* **Enterprise Plans**: Custom pricing for high-volume usage and dedicated infrastructure
* **Quota Management**: Track usage through your platform dashboard
Contact us at [support@unsiloed.ai](mailto:support@unsiloed.ai) for pricing details and enterprise plans.
***
Contact our support team for personalized assistance
# Documentation for Coding Agents
Source: https://docs.unsiloed.ai/docs/for-coding-agents
Machine-readable documentation entry points for agents integrating the Unsiloed document-processing APIs.
Coding agents can read the Unsiloed documentation without rendering the website. Start with the focused documentation index, then open only the raw Markdown pages needed for the task.
## Start With the Focused Documentation Index
Use [`llms.txt`](https://www.unsiloed.ai/docs/llms.txt) to discover documentation pages and their raw Markdown URLs. The index is approximately 10 KB, so it provides a lower-context starting point than loading the complete documentation corpus.
Use [`llms-full.txt`](https://www.unsiloed.ai/docs/llms-full.txt) only when the task requires searching across the entire documentation set.
Every documentation page is also available as raw Markdown. For example:
| Documentation resource | Machine-readable URL |
| - | - |
| Introduction | [`/docs/index.md`](https://www.unsiloed.ai/docs/index.md) |
| Parsing quickstart | [`/docs/quickstart.md`](https://www.unsiloed.ai/docs/quickstart.md) |
| Parse overview | [`/docs/document-processing/parsing/parsing.md`](https://www.unsiloed.ai/docs/document-processing/parsing/parsing.md) |
| Extract overview | [`/docs/document-processing/extraction/extraction.md`](https://www.unsiloed.ai/docs/document-processing/extraction/extraction.md) |
| Classification overview | [`/docs/document-processing/classification/classification.md`](https://www.unsiloed.ai/docs/document-processing/classification/classification.md) |
| Splitting overview | [`/docs/document-processing/splitting/splitting.md`](https://www.unsiloed.ai/docs/document-processing/splitting/splitting.md) |
## Choose the Document Operation
Unsiloed separates document processing into four operations. Choose the operation from the result your application needs:
| Goal | Start here |
| - | - |
| Convert the complete document into ordered Markdown chunks with layout metadata | [Parse overview](https://www.unsiloed.ai/docs/document-processing/parsing/parsing.md) |
| Pull named fields into a JSON schema with confidence scores and citations | [Extract overview](https://www.unsiloed.ai/docs/document-processing/extraction/extraction.md) |
| Assign a document to one of your predefined categories | [Classification overview](https://www.unsiloed.ai/docs/document-processing/classification/classification.md) |
| Separate a bundled PDF into documents by category | [Splitting overview](https://www.unsiloed.ai/docs/document-processing/splitting/splitting.md) |
For a first parsing integration, follow the [parsing quickstart](https://www.unsiloed.ai/docs/quickstart.md). It provides Python, JavaScript, and cURL examples for submitting a document, polling the job, and reading the result. The Python and JavaScript examples save both JSON and Markdown output.
## Use the API Endpoints
Send authenticated requests to `https://prod.visionapi.unsiloed.ai` with your API key in the `api-key` header. Use these routes for new integrations:
| Operation | Submit | Retrieve the result | Reference |
| - | - | - | - |
| Parse | `POST /parse` | [`GET /parse/{job_id}`](https://www.unsiloed.ai/docs/api-reference/parser/get-parse-job-status.md) | [Parse API](https://www.unsiloed.ai/docs/api-reference/parser/parse-document.md) |
| Parse a large file | `POST /v2/parse/upload`, then `PUT` to the returned URL | [`GET /parse/{job_id}`](https://www.unsiloed.ai/docs/api-reference/parser/get-parse-job-status.md) | [Large-file upload API](https://www.unsiloed.ai/docs/api-reference/parser/parse-document-v2.md) |
| Parse Excel | `POST /parse/excel` | [`GET /parse/{job_id}`](https://www.unsiloed.ai/docs/api-reference/parser/get-parse-job-status.md) | [Excel API](https://www.unsiloed.ai/docs/api-reference/parser/parse-excel.md) |
| Extract | `POST /v2/extract` | [`GET /extract/{job_id}`](https://www.unsiloed.ai/docs/api-reference/jobs/results.md) | [Extract API](https://www.unsiloed.ai/docs/api-reference/extraction/extract-data.md) |
| Classify | `POST /classify` | [`GET /classify/{job_id}`](https://www.unsiloed.ai/docs/api-reference/classification/get-classification-status.md) | [Classification API](https://www.unsiloed.ai/docs/api-reference/classification/classify-document.md) |
| Split | `POST /splitter` | [`GET /splitter/{job_id}`](https://www.unsiloed.ai/docs/api-reference/splitting/get-split-status.md) | [Splitting API](https://www.unsiloed.ai/docs/api-reference/splitting/split-document.md) |
Accepted jobs return a `job_id` for polling. Parse jobs finish with `Succeeded`, `Failed`, or `Cancelled`. Extraction results are available when the status is `completed` or `review`, and extraction can end with `failed`. Classification and splitting jobs finish with `completed` or `failed`.
Use `POST /parse` for standard parsing unless the file needs the large-file upload flow or is an Excel workbook. For the large-file upload, send the file to `upload_url` with the returned `upload_headers`; the presigned `PUT` does not use the `api-key` header. Poll the resulting job through `GET /parse/{job_id}`.
When extraction has personally identifiable information (PII) detection enabled, it can instead return `status: "pii_blocked"` with no job. Check the submit response before polling.
## Treat OpenAPI as Supplementary Reference
Use the human-readable guides and API references linked above as the source of truth for current integrations. The aggregate OpenAPI specification still uses older Parse field names and requires a file upload, so it does not fully describe URL-only Parse requests. Verify generated clients against the relevant endpoint reference before using them:
* [Aggregate API specification](https://www.unsiloed.ai/docs/api-reference/openapi.json)
## Distinguish Documentation Access From Product Tools
The [Unsiloed Document Processing MCP server](https://www.unsiloed.ai/docs/integrations/mcp-server.md) gives an agent tools for parsing, classification, and structured extraction. It processes documents and does not search the documentation.
For documentation research, use `llms.txt` and raw Markdown pages. For document-processing actions, use the MCP server or REST API.
# Introduction
Source: https://docs.unsiloed.ai/docs/index
# Welcome to Unsiloed AI
Agentic OCR for AI pipelines that need to trust the page. Turn PDFs, scans, and forms into Markdown and structured JSON, with confidence scores and bounding boxes on every value.
Unsiloed AI parses unstructured documents (PDFs, scans, slides, spreadsheets, and 20+ file formats) into Markdown and structured JSON that LLMs and agents can use directly. The API sits between your raw files and your retrieval, extraction, or automation pipeline.
Generic OCR and text-only LLM parsers lose tables when columns wrap, mangle reading order in multi-column layouts, and produce brittle outputs on the real-world PDFs that show up in invoices, contracts, and forms. Unsiloed AI uses vision and layout models alongside OCR so the structure of the source survives the parse.
## API Capabilities
The API covers four document operations:
Convert PDFs, DOCX, PPTX, images, and more into hierarchical Markdown chunks. Tables, figures, formulas, and headers are preserved as first-class segments with bounding boxes.
Define a JSON schema and get back typed fields with word-level citations and per-field confidence scores. Useful for invoices, claims, KYC forms, and any pipeline that needs auditability.
Detect document boundaries inside merged or scanned batches and return each one separately. Works on layout, content, or custom rules.
Route incoming files to the right downstream pipeline by classifying them against a list of categories you define.
## Use Unsiloed from Claude
Add Unsiloed as a remote MCP connector in Claude.ai or Claude Desktop. Parse, classify, and extract from chat — OAuth-based, no API key pasting.
Drop-in Anthropic tool-use schemas for calling Unsiloed directly from the Claude API.
## Built for Production Pipelines
The API is designed for the things teams hit when they move document workflows out of a prototype.
Production Workloads
Asynchronous processing for large and multi-page documents
Deterministic outputs with confidence scores and word-level bounding boxes
Broad multi-format support across PDFs, DOCX, PPTX, images, and more
Scalable infrastructure for high-throughput enterprise workloads
Developer Experience
Clean REST APIs with stable versioned contracts
Schema-driven extraction with validation, confidence, and traceability
Interactive playground for testing API requests, schemas, and outputs
Predictable error handling for reliable production integrations
## Common Use Cases
* **Finance:** Parse financial statements, reports, and regulatory filings into structured, machine-readable data.
* **Legal:** Extract clauses, entities, dates, and obligations from contracts and legal documents.
* **Healthcare:** Structure clinical documents, forms, and records for downstream systems and workflows.
* **RAG & Automation:** Parse, chunk, classify, and route documents to power reliable RAG pipelines and document-driven automations.
## Next Steps
[Sign up on Unsiloed AI](https://cal.com/aman-mishra-p0ry57/15min) to receive a key.
Follow the [Quickstart](/docs/quickstart) guide to submit a document and read back chunks.
Define a JSON schema and pull typed values out of a document. See the [extraction guide](/docs/document-processing/extraction/extraction).
Browse the [API reference](/docs/api-reference/parser/parse-document) for parsing strategies, classification, splitting, and batch endpoints.
## Need Help?
Guides for parsing, extraction, classification, and splitting, plus the full API reference.
Email [support@unsiloed.ai](mailto:support@unsiloed.ai) to reach the team.
# Upgrade Your OpenClaw Agent with the Unsiloed Skill
Source: https://docs.unsiloed.ai/docs/integrations/agent-skills/openclaw
Install the Unsiloed skill so your OpenClaw agent routes every image and PDF through Unsiloed's document AI instead of its built-in vision, returning structured data with a confidence score on every field.
An OpenClaw agent reads documents with the underlying model's general-purpose vision. That works on a clean printed PDF, but on handwriting, a dense table, or a faded scan the agent makes up content and gives no signal that anything went wrong. This guide installs the Unsiloed skill into your agent so any document attachment routes through the Unsiloed API instead. The agent gets back structured data with a confidence score on every value, so its reply flags what to verify.
By the end, your agent auto-invokes Unsiloed whenever someone sends it a document or asks about one, across any connected channel.
The skill is a single `SKILL.md` file that drives the Unsiloed API with `curl` and `jq`. There's no Python or extra runtime to manage. It exposes all four Unsiloed operations (Parse, Extract, Classify, and Split), and the agent picks the right one per request.
## What Changes After You Install It
We gave an OpenClaw agent a handwritten prescription and asked which medicines were listed.
* **Default vision read:** the agent confidently returns medicine names that aren't on the page, with no confidence score and no way to tell which values to trust.
* **Routed through Unsiloed:** the agent returns the three medicines actually written on the page, each with a confidence score above 0.95.
The prescription is one case among many. The same skill handles any document that's hard to read with vision alone:
* Faded thermal receipts
* Dense financial tables that span pages
* Multi-column forms
* Handwritten invoices
* Low-contrast scans
## Prerequisites for the Unsiloed Skill
Before installing, make sure you have:
* **OpenClaw with the gateway configured:** see the [OpenClaw skills documentation](https://docs.openclaw.ai/cli/skills) to set it up.
* **An Unsiloed API key:** [sign up on Unsiloed AI](https://www.unsiloed.ai) to get one.
* **`curl` and `jq` on the machine running the gateway:** the skill declares both as requirements and won't load without them.
## Install the Skill
The skill installs from its GitHub repository through the OpenClaw CLI. Run these steps in order.
The skill lives in the Unsiloed cookbook at [`skills/unsiloed`](https://github.com/Unsiloed-AI/cookbook/tree/main/skills/unsiloed). Point the OpenClaw CLI at the repository and that subpath:
```bash theme={null}
openclaw skills install git:https://github.com/Unsiloed-AI/cookbook --subpath skills/unsiloed
```
OpenClaw clones the repo, registers the skill in your workspace, and tracks the origin, so `openclaw skills update --all` pulls future updates from the same source.
The skill reads the key from the gateway's environment. Append it to OpenClaw's global env file:
```bash theme={null}
echo 'UNSILOED_API_KEY=us_...' >> ~/.openclaw/.env
```
Replace `us_...` with your actual key.
The gateway loads env vars at startup, so restart it to pick up the new key:
```bash theme={null}
openclaw gateway restart
```
Confirm the skill loaded and its requirements pass:
```bash theme={null}
openclaw skills info unsiloed
```
The first line of the output should read:
```
unsiloed ✓ Ready
```
To see it alongside every other skill, run `openclaw skills check` — `unsiloed` should appear under **Ready and visible to model**. If it shows up under **Missing requirements** instead, the `UNSILOED_API_KEY`, `curl`, or `jq` requirement isn't satisfied yet.
## Send Your Agent a Document
Send a document to the agent through any connected channel: Telegram, a file share, or a paired terminal. The skill auto-invokes whenever you attach a document or ask a question about one, such as "what does this say", "what medicines are listed", or "extract the totals". The agent picks the right Unsiloed operation, polls the async job until it finishes, and replies in plain English. The user never sees raw JSON.
For example:
```
You: [attached prescription.jpg] What medicines are listed here?
Agent: Three medicines: Tab Azee 500mg, Tab Montair FX, and Tab Dolo 650.
All extracted with confidence above 0.95.
```
## Parse vs Extract: When the Agent Uses Each
The skill teaches the agent to choose between the Unsiloed operations. The two you'll see most often are Parse and Extract.
**Parse** reads an entire document and returns it as Markdown, with every layout region preserved (paragraphs, tables, lists, headings, captions). It's the default. For most chat questions about a document, the agent parses it, then reads and quotes what's there.
**Extract** pulls specific fields you name in advance (invoice total, patient name, date of issue) and returns each one with a confidence score from 0 to 1. Use Extract when the output is going somewhere structured, such as a spreadsheet, a database row, or an automated workflow, and you want a per-field confidence score so anything below a threshold like 0.85 can be flagged for human review. Because Extract returns the same shape every time, the agent can pipe it straight into a destination you've already wired up.
The skill defaults to Parse, so most prompts route there automatically. To trigger Extract, ask for structured fields back:
```
Give me the total and the date as JSON.
```
The skill also exposes **Classify** (label a document as one of several candidate categories) and **Split** (break a single PDF containing several documents into separate files by type). The agent reaches for these when the request clearly calls for them: "is this a receipt or a contract", or "split this bundle into separate files".
## Keeping the Skill Up to Date
Because the skill was installed from git, pull the latest version any time with:
```bash theme={null}
openclaw skills update --all
```
This updates every tracked skill from its origin. Restart the gateway afterward if the update changes anything the running agent has cached.
## Next Steps
Configure chunking strategies, segment filters, and the OCR backend behind Parse.
Define a JSON schema and pull typed fields out of a document with confidence scores.
Give Claude direct, structured access to the same operations through Anthropic tool use.
Browse the full request and response specs for every endpoint.
# Claude Integration
Source: https://docs.unsiloed.ai/docs/integrations/claude-integration
Drop-in Anthropic tool-use schemas to give Claude structured access to Unsiloed's document processing API
## Overview
This page provides ready-to-use [Anthropic tool-use](https://docs.anthropic.com/en/docs/build-with-claude/tool-use/overview) JSON schemas that let Claude interact with Unsiloed's document processing API. Copy the tool definitions into your `messages.create()` call and Claude can parse, extract, classify, and split documents on your behalf.
All Unsiloed API operations are **asynchronous**. Each tool submits a job and returns a `job_id`. Use the `unsiloed_get_job_result` tool to poll for results. Since Claude cannot upload binary files, all tools accept a **publicly accessible URL** to the document.
## Prerequisites
Before using these tool schemas, you'll need:
* An **Unsiloed API key** — [sign up at unsiloed.ai](https://www.unsiloed.ai) to get one
* An **Anthropic API key** — [get one from the Anthropic console](https://console.anthropic.com)
* **Python 3.8+** or **Node.js 18+**
```bash Python theme={null}
pip install anthropic requests
```
```bash TypeScript theme={null}
npm install @anthropic-ai/sdk
```
## Key Features
Copy-paste JSON tool definitions directly into your Anthropic API calls
Built-in polling tool to retrieve results from asynchronous operations
Parse documents, extract structured data, classify, and split PDFs
Full working example of an autonomous document processing agent
## Tool Definitions
Each tool follows [Anthropic's tool format](https://docs.anthropic.com/en/docs/build-with-claude/tool-use/overview) with `name`, `description`, and `input_schema`. Expand each tool below to see its full schema.
Parse documents (PDF, images, Office files) into structured chunks with element detection, OCR, and reading order analysis.
```json theme={null}
{
"name": "unsiloed_parse_document",
"description": "Parse a document into structured chunks with element detection, text extraction, and reading order analysis. Use this tool when the user wants to break a document into its structural components (text blocks, tables, images, headers, footers) or convert a document to markdown. You must provide a publicly accessible URL to the document. Supported formats include PDF, PNG, JPEG, TIFF, BMP, DOCX, XLSX, and PPTX. The operation is asynchronous — it returns a job_id that you must poll using unsiloed_get_job_result to retrieve the parsed content.",
"input_schema": {
"type": "object",
"properties": {
"url": {
"type": "string",
"description": "Publicly accessible URL to the document to parse."
},
"use_high_resolution": {
"type": "boolean",
"description": "Use high-resolution image processing for better OCR accuracy. Defaults to false."
},
"segmentation_method": {
"type": "string",
"enum": ["smart_layout_detection", "page_by_page"],
"description": "Document segmentation strategy. 'smart_layout_detection' (default) groups related elements into semantic chunks. 'page_by_page' creates one chunk per page."
},
"ocr_mode": {
"type": "string",
"enum": ["auto_ocr", "full_ocr"],
"description": "OCR strategy. 'auto_ocr' (default) only runs OCR when needed. 'full_ocr' forces OCR on all pages."
},
"merge_tables": {
"type": "boolean",
"description": "Merge adjacent table segments into a single table. Defaults to false."
},
"segment_filter": {
"type": "string",
"description": "Filter output to include only specific segment types. Accepts comma-separated values (e.g., 'table', 'picture', 'table,picture') or 'all' for everything. Defaults to 'all'."
},
"xml_citation": {
"type": "boolean",
"description": "Enable citation extraction from PDF documents. Extracts structured bibliography and in-text citations in markdown output. Defaults to false."
},
"output_fields": {
"type": "string",
"description": "JSON object controlling which fields are included in the response. Set fields to false to exclude them and reduce response size. Available fields: html, markdown, ocr, image, content, bbox, confidence, embed. All default to true. Example: {\"html\":true,\"markdown\":true,\"ocr\":false,\"image\":false}"
},
"segment_analysis": {
"type": "string",
"description": "JSON object controlling HTML/Markdown generation strategy and AI model per segment type. Example: {\"Table\":{\"html\":\"VLM\",\"markdown\":\"VLM\"},\"Picture\":{\"html\":\"VLM\",\"markdown\":\"VLM\"}}"
},
"page_range": {
"type": "string",
"description": "Specify which pages to process. Formats: '1-5', '2,4,6', or '[1,3,5]'. Defaults to all pages."
},
"segment_type_naming": {
"type": "string",
"enum": ["Unsiloed", "Other"],
"description": "Segment type naming convention. 'Unsiloed' (default) uses names like PageHeader, ListItem, Picture. 'Other' uses alternative names like Header, List Item, Figure."
},
"extract_colors": {
"type": "boolean",
"description": "Transfer text color from the PDF text layer to OCR results. Defaults to false."
},
"extract_links": {
"type": "boolean",
"description": "Attach hyperlink URLs from PDF annotations to OCR results. Defaults to false."
},
"export_format": {
"type": "string",
"description": "JSON array of export format(s) to generate after parsing. The exported files are available as presigned URLs in the exports field of the response. Currently supported: [\"docx\"]. Example: [\"docx\"]"
},
"error_handling": {
"type": "string",
"enum": ["Continue", "Fail"],
"description": "How to handle per-page errors. 'Continue' (default) skips failed pages and continues. 'Fail' aborts the entire job on the first error."
},
"expires_in": {
"type": "integer",
"description": "Seconds until the task and its output are automatically deleted."
},
"chunk_processing": {
"type": "string",
"description": "JSON object for chunk processing configuration."
},
"llm_processing": {
"type": "string",
"description": "JSON object for LLM processing configuration."
}
},
"required": ["url"]
}
}
```
Extract structured data from PDF documents using a custom JSON schema you define.
```json theme={null}
{
"name": "unsiloed_extract_data",
"description": "Extract structured data from a PDF document using a custom JSON schema. Use this tool when the user wants to pull specific fields (such as invoice numbers, names, dates, or line items) out of a PDF. You must provide a publicly accessible URL to the PDF and a JSON schema defining the fields to extract. The operation is asynchronous — it returns a job_id that you must poll using unsiloed_get_job_result to retrieve the extracted data. Do not use this tool for document classification or splitting.",
"input_schema": {
"type": "object",
"properties": {
"file_url": {
"type": "string",
"description": "Publicly accessible URL to the PDF file to process."
},
"schema_data": {
"type": "string",
"description": "A JSON-stringified schema defining the fields to extract. Example: {\"type\":\"object\",\"properties\":{\"invoice_number\":{\"type\":\"string\",\"description\":\"The invoice number\"},\"total\":{\"type\":\"number\",\"description\":\"Total amount\"}},\"required\":[\"invoice_number\"],\"additionalProperties\":false}"
},
"model": {
"type": "string",
"enum": ["alpha", "beta", "gamma", "delta"],
"description": "Model tier for extraction. Default is 'gamma', recommended for most use cases."
},
"enable_citations": {
"type": "boolean",
"description": "When true, returns bounding box coordinates for each extracted value, enabling you to trace data back to its location in the PDF. Defaults to false."
}
},
"required": ["file_url", "schema_data"]
}
}
```
Classify a PDF document into one of several predefined categories with confidence scoring.
```json theme={null}
{
"name": "unsiloed_classify_document",
"description": "Classify a PDF document into one of several predefined categories. Use this tool when the user wants to determine what type of document a PDF is (for example, invoice, receipt, contract, or medical record). You must provide a publicly accessible URL to the PDF and a list of candidate categories. The operation is asynchronous — it returns a job_id that you must poll using unsiloed_get_job_result to retrieve the classification result with confidence scores. Do not use this for data extraction or document splitting.",
"input_schema": {
"type": "object",
"properties": {
"file_url": {
"type": "string",
"description": "Publicly accessible URL to the PDF file to classify."
},
"categories": {
"type": "string",
"description": "A JSON-stringified array of category objects. Each object must have a 'name' field and may have an optional 'description' field for better accuracy. Example: [{\"name\":\"Invoice\",\"description\":\"Financial invoices with itemized charges\"},{\"name\":\"Receipt\"},{\"name\":\"Contract\",\"description\":\"Legal agreements\"}]"
}
},
"required": ["file_url", "categories"]
}
}
```
Split a multi-document PDF into separate files by classifying each page into categories.
```json theme={null}
{
"name": "unsiloed_split_document",
"description": "Split a multi-document PDF into separate files by classifying each page into predefined categories. Use this tool when the user has a single PDF containing multiple document types (for example, a scanned batch of invoices, receipts, and contracts) and wants them separated into individual files. You must provide a publicly accessible URL to the PDF and a list of candidate categories. The operation is asynchronous — it returns a job_id that you must poll using unsiloed_get_job_result to retrieve download links for the split files. Do not use this for single-document classification or data extraction.",
"input_schema": {
"type": "object",
"properties": {
"file_url": {
"type": "string",
"description": "Publicly accessible URL to the PDF file to split."
},
"categories": {
"type": "string",
"description": "A JSON-stringified array of category objects. Each object must have a 'name' field and may have an optional 'description' field for better accuracy. Example: [{\"name\":\"Invoice\",\"description\":\"Financial invoices\"},{\"name\":\"Contract\"},{\"name\":\"Receipt\"}]"
}
},
"required": ["file_url", "categories"]
}
}
```
Poll for the result of any asynchronous Unsiloed job.
```json theme={null}
{
"name": "unsiloed_get_job_result",
"description": "Poll for the result of an asynchronous Unsiloed job. Use this tool after calling unsiloed_parse_document, unsiloed_extract_data, unsiloed_classify_document, or unsiloed_split_document to check whether the job has completed and retrieve its results. If the status indicates the job is still processing, wait a few seconds and call this tool again. Once the job is complete, the response contains the output data. If the job failed, the response includes an error message explaining what went wrong.",
"input_schema": {
"type": "object",
"properties": {
"job_id": {
"type": "string",
"description": "The job_id returned by a previous Unsiloed tool call."
},
"job_type": {
"type": "string",
"enum": ["parse", "extract", "classify", "splitter"],
"description": "The type of job to check. Use 'parse' for parsing jobs, 'extract' for extraction jobs, 'classify' for classification jobs, and 'splitter' for splitting jobs."
}
},
"required": ["job_id", "job_type"]
}
}
```
## All Tools (Copy & Paste)
Copy this complete `tools` array and pass it directly to your Anthropic API call. The SDK examples below refer back to it as the `tools` array.
```json theme={null}
[
{
"name": "unsiloed_parse_document",
"description": "Parse a document into structured chunks with element detection, text extraction, and reading order analysis. Use this tool when the user wants to break a document into its structural components (text blocks, tables, images, headers, footers) or convert a document to markdown. You must provide a publicly accessible URL to the document. Supported formats include PDF, PNG, JPEG, TIFF, BMP, DOCX, XLSX, and PPTX. The operation is asynchronous — it returns a job_id that you must poll using unsiloed_get_job_result to retrieve the parsed content.",
"input_schema": {
"type": "object",
"properties": {
"url": {
"type": "string",
"description": "Publicly accessible URL to the document to parse."
},
"use_high_resolution": {
"type": "boolean",
"description": "Use high-resolution image processing for better OCR accuracy. Defaults to false."
},
"segmentation_method": {
"type": "string",
"enum": ["smart_layout_detection", "page_by_page"],
"description": "Document segmentation strategy. 'smart_layout_detection' (default) groups related elements into semantic chunks. 'page_by_page' creates one chunk per page."
},
"ocr_mode": {
"type": "string",
"enum": ["auto_ocr", "full_ocr"],
"description": "OCR strategy. 'auto_ocr' (default) only runs OCR when needed. 'full_ocr' forces OCR on all pages."
},
"merge_tables": {
"type": "boolean",
"description": "Merge adjacent table segments into a single table. Defaults to false."
},
"segment_filter": {
"type": "string",
"description": "Filter output to include only specific segment types. Accepts comma-separated values (e.g., 'table', 'picture', 'table,picture') or 'all' for everything. Defaults to 'all'."
},
"xml_citation": {
"type": "boolean",
"description": "Enable citation extraction from PDF documents. Extracts structured bibliography and in-text citations in markdown output. Defaults to false."
},
"output_fields": {
"type": "string",
"description": "JSON object controlling which fields are included in the response. Set fields to false to exclude them and reduce response size. Available fields: html, markdown, ocr, image, content, bbox, confidence, embed. All default to true. Example: {\"html\":true,\"markdown\":true,\"ocr\":false,\"image\":false}"
},
"segment_analysis": {
"type": "string",
"description": "JSON object controlling HTML/Markdown generation strategy and AI model per segment type. Example: {\"Table\":{\"html\":\"VLM\",\"markdown\":\"VLM\"},\"Picture\":{\"html\":\"VLM\",\"markdown\":\"VLM\"}}"
},
"page_range": {
"type": "string",
"description": "Specify which pages to process. Formats: '1-5', '2,4,6', or '[1,3,5]'. Defaults to all pages."
},
"segment_type_naming": {
"type": "string",
"enum": ["Unsiloed", "Other"],
"description": "Segment type naming convention. 'Unsiloed' (default) uses names like PageHeader, ListItem, Picture. 'Other' uses alternative names like Header, List Item, Figure."
},
"extract_colors": {
"type": "boolean",
"description": "Transfer text color from the PDF text layer to OCR results. Defaults to false."
},
"extract_links": {
"type": "boolean",
"description": "Attach hyperlink URLs from PDF annotations to OCR results. Defaults to false."
},
"export_format": {
"type": "string",
"description": "JSON array of export format(s) to generate after parsing. The exported files are available as presigned URLs in the exports field of the response. Currently supported: [\"docx\"]. Example: [\"docx\"]"
},
"error_handling": {
"type": "string",
"enum": ["Continue", "Fail"],
"description": "How to handle per-page errors. 'Continue' (default) skips failed pages and continues. 'Fail' aborts the entire job on the first error."
},
"expires_in": {
"type": "integer",
"description": "Seconds until the task and its output are automatically deleted."
},
"chunk_processing": {
"type": "string",
"description": "JSON object for chunk processing configuration."
},
"llm_processing": {
"type": "string",
"description": "JSON object for LLM processing configuration."
}
},
"required": ["url"]
}
},
{
"name": "unsiloed_extract_data",
"description": "Extract structured data from a PDF document using a custom JSON schema. Use this tool when the user wants to pull specific fields (such as invoice numbers, names, dates, or line items) out of a PDF. You must provide a publicly accessible URL to the PDF and a JSON schema defining the fields to extract. The operation is asynchronous — it returns a job_id that you must poll using unsiloed_get_job_result to retrieve the extracted data. Do not use this tool for document classification or splitting.",
"input_schema": {
"type": "object",
"properties": {
"file_url": {
"type": "string",
"description": "Publicly accessible URL to the PDF file to process."
},
"schema_data": {
"type": "string",
"description": "A JSON-stringified schema defining the fields to extract. Example: {\"type\":\"object\",\"properties\":{\"invoice_number\":{\"type\":\"string\",\"description\":\"The invoice number\"},\"total\":{\"type\":\"number\",\"description\":\"Total amount\"}},\"required\":[\"invoice_number\"],\"additionalProperties\":false}"
},
"model": {
"type": "string",
"enum": ["alpha", "beta", "gamma", "delta"],
"description": "Model tier for extraction. Default is 'gamma', recommended for most use cases."
},
"enable_citations": {
"type": "boolean",
"description": "When true, returns bounding box coordinates for each extracted value, enabling you to trace data back to its location in the PDF. Defaults to false."
}
},
"required": ["file_url", "schema_data"]
}
},
{
"name": "unsiloed_classify_document",
"description": "Classify a PDF document into one of several predefined categories. Use this tool when the user wants to determine what type of document a PDF is (for example, invoice, receipt, contract, or medical record). You must provide a publicly accessible URL to the PDF and a list of candidate categories. The operation is asynchronous — it returns a job_id that you must poll using unsiloed_get_job_result to retrieve the classification result with confidence scores. Do not use this for data extraction or document splitting.",
"input_schema": {
"type": "object",
"properties": {
"file_url": {
"type": "string",
"description": "Publicly accessible URL to the PDF file to classify."
},
"categories": {
"type": "string",
"description": "A JSON-stringified array of category objects. Each object must have a 'name' field and may have an optional 'description' field for better accuracy. Example: [{\"name\":\"Invoice\",\"description\":\"Financial invoices with itemized charges\"},{\"name\":\"Receipt\"},{\"name\":\"Contract\",\"description\":\"Legal agreements\"}]"
}
},
"required": ["file_url", "categories"]
}
},
{
"name": "unsiloed_split_document",
"description": "Split a multi-document PDF into separate files by classifying each page into predefined categories. Use this tool when the user has a single PDF containing multiple document types (for example, a scanned batch of invoices, receipts, and contracts) and wants them separated into individual files. You must provide a publicly accessible URL to the PDF and a list of candidate categories. The operation is asynchronous — it returns a job_id that you must poll using unsiloed_get_job_result to retrieve download links for the split files. Do not use this for single-document classification or data extraction.",
"input_schema": {
"type": "object",
"properties": {
"file_url": {
"type": "string",
"description": "Publicly accessible URL to the PDF file to split."
},
"categories": {
"type": "string",
"description": "A JSON-stringified array of category objects. Each object must have a 'name' field and may have an optional 'description' field for better accuracy. Example: [{\"name\":\"Invoice\",\"description\":\"Financial invoices\"},{\"name\":\"Contract\"},{\"name\":\"Receipt\"}]"
}
},
"required": ["file_url", "categories"]
}
},
{
"name": "unsiloed_get_job_result",
"description": "Poll for the result of an asynchronous Unsiloed job. Use this tool after calling unsiloed_parse_document, unsiloed_extract_data, unsiloed_classify_document, or unsiloed_split_document to check whether the job has completed and retrieve its results. If the status indicates the job is still processing, wait a few seconds and call this tool again. Once the job is complete, the response contains the output data. If the job failed, the response includes an error message explaining what went wrong.",
"input_schema": {
"type": "object",
"properties": {
"job_id": {
"type": "string",
"description": "The job_id returned by a previous Unsiloed tool call."
},
"job_type": {
"type": "string",
"enum": ["parse", "extract", "classify", "splitter"],
"description": "The type of job to check. Use 'parse' for parsing jobs, 'extract' for extraction jobs, 'classify' for classification jobs, and 'splitter' for splitting jobs."
}
},
"required": ["job_id", "job_type"]
}
}
]
```
## Usage with the Anthropic SDK
Here's how to register these tools and handle Claude's tool calls:
```python Python theme={null}
import anthropic
import requests
import json
import time
UNSILOED_API_KEY = "your-unsiloed-api-key"
UNSILOED_BASE_URL = "https://prod.visionapi.unsiloed.ai"
UNSILOED_HEADERS = {"api-key": UNSILOED_API_KEY}
# Paste the tools array from the "All Tools" section above
tools = [...]
client = anthropic.Anthropic(api_key="your-anthropic-api-key")
def process_tool_call(tool_name: str, tool_input: dict) -> str:
"""Execute an Unsiloed API tool call and return the result as a string."""
if tool_name == "unsiloed_parse_document":
data = {"url": tool_input["url"]}
if "use_high_resolution" in tool_input:
data["use_high_resolution"] = str(tool_input["use_high_resolution"]).lower()
if "segmentation_method" in tool_input:
data["segmentation_method"] = tool_input["segmentation_method"]
if "ocr_mode" in tool_input:
data["ocr_mode"] = tool_input["ocr_mode"]
if "merge_tables" in tool_input:
data["merge_tables"] = str(tool_input["merge_tables"]).lower()
if "segment_filter" in tool_input:
data["segment_filter"] = tool_input["segment_filter"]
if "xml_citation" in tool_input:
data["xml_citation"] = str(tool_input["xml_citation"]).lower()
resp = requests.post(f"{UNSILOED_BASE_URL}/parse", headers=UNSILOED_HEADERS, data=data)
elif tool_name == "unsiloed_extract_data":
data = {
"file_url": tool_input["file_url"],
"schema_data": tool_input["schema_data"],
}
if "model" in tool_input:
data["model"] = tool_input["model"]
if "enable_citations" in tool_input:
data["enable_citations"] = str(tool_input["enable_citations"]).lower()
resp = requests.post(f"{UNSILOED_BASE_URL}/v2/extract", headers=UNSILOED_HEADERS, data=data)
elif tool_name == "unsiloed_classify_document":
data = {
"file_url": tool_input["file_url"],
"categories": tool_input["categories"],
}
resp = requests.post(f"{UNSILOED_BASE_URL}/classify", headers=UNSILOED_HEADERS, data=data)
elif tool_name == "unsiloed_split_document":
data = {
"file_url": tool_input["file_url"],
"categories": tool_input["categories"],
}
resp = requests.post(f"{UNSILOED_BASE_URL}/splitter", headers=UNSILOED_HEADERS, data=data)
elif tool_name == "unsiloed_get_job_result":
job_type = tool_input["job_type"]
job_id = tool_input["job_id"]
time.sleep(5) # Brief pause before polling
resp = requests.get(f"{UNSILOED_BASE_URL}/{job_type}/{job_id}", headers=UNSILOED_HEADERS)
else:
return json.dumps({"error": f"Unknown tool: {tool_name}"})
return json.dumps(resp.json())
```
```typescript TypeScript theme={null}
import Anthropic from "@anthropic-ai/sdk";
const UNSILOED_API_KEY = "your-unsiloed-api-key";
const UNSILOED_BASE_URL = "https://prod.visionapi.unsiloed.ai";
// Paste the tools array from the "All Tools" section above
const tools: Anthropic.Tool[] = [];
const client = new Anthropic({ apiKey: "your-anthropic-api-key" });
async function processToolCall(
toolName: string,
toolInput: Record
): Promise {
const headers: Record = { "api-key": UNSILOED_API_KEY };
let resp: Response;
if (toolName === "unsiloed_parse_document") {
const formData = new FormData();
formData.append("url", toolInput.url as string);
if (toolInput.use_high_resolution !== undefined)
formData.append("use_high_resolution", String(toolInput.use_high_resolution));
if (toolInput.segmentation_method)
formData.append("segmentation_method", toolInput.segmentation_method as string);
if (toolInput.ocr_mode)
formData.append("ocr_mode", toolInput.ocr_mode as string);
if (toolInput.merge_tables !== undefined)
formData.append("merge_tables", String(toolInput.merge_tables));
if (toolInput.segment_filter)
formData.append("segment_filter", toolInput.segment_filter as string);
if (toolInput.xml_citation !== undefined)
formData.append("xml_citation", String(toolInput.xml_citation));
resp = await fetch(`${UNSILOED_BASE_URL}/parse`, {
method: "POST", headers, body: formData,
});
} else if (toolName === "unsiloed_extract_data") {
const formData = new FormData();
formData.append("file_url", toolInput.file_url as string);
formData.append("schema_data", toolInput.schema_data as string);
if (toolInput.model) formData.append("model", toolInput.model as string);
if (toolInput.enable_citations !== undefined)
formData.append("enable_citations", String(toolInput.enable_citations));
resp = await fetch(`${UNSILOED_BASE_URL}/v2/extract`, {
method: "POST", headers, body: formData,
});
} else if (toolName === "unsiloed_classify_document") {
const formData = new FormData();
formData.append("file_url", toolInput.file_url as string);
formData.append("categories", toolInput.categories as string);
resp = await fetch(`${UNSILOED_BASE_URL}/classify`, {
method: "POST", headers, body: formData,
});
} else if (toolName === "unsiloed_split_document") {
const formData = new FormData();
formData.append("file_url", toolInput.file_url as string);
formData.append("categories", toolInput.categories as string);
resp = await fetch(`${UNSILOED_BASE_URL}/splitter`, {
method: "POST", headers, body: formData,
});
} else if (toolName === "unsiloed_get_job_result") {
const jobType = toolInput.job_type as string;
const jobId = toolInput.job_id as string;
await new Promise((r) => setTimeout(r, 5000)); // Brief pause before polling
resp = await fetch(`${UNSILOED_BASE_URL}/${jobType}/${jobId}`, { headers });
} else {
return JSON.stringify({ error: `Unknown tool: ${toolName}` });
}
return JSON.stringify(await resp.json());
}
```
## Complete Agentic Loop Example
This standalone example shows a full autonomous loop where Claude processes a document end-to-end:
```python Python theme={null}
import anthropic
import requests
import json
import time
# --- Configuration ---
ANTHROPIC_API_KEY = "your-anthropic-api-key"
UNSILOED_API_KEY = "your-unsiloed-api-key"
UNSILOED_BASE_URL = "https://prod.visionapi.unsiloed.ai"
UNSILOED_HEADERS = {"api-key": UNSILOED_API_KEY}
# --- Load tools (paste the array from the "All Tools" section above) ---
tools = [...]
# --- Tool executor ---
def process_tool_call(tool_name: str, tool_input: dict) -> str:
if tool_name == "unsiloed_extract_data":
resp = requests.post(
f"{UNSILOED_BASE_URL}/v2/extract",
headers=UNSILOED_HEADERS,
data={
"file_url": tool_input["file_url"],
"schema_data": tool_input["schema_data"],
**({"model": tool_input["model"]} if "model" in tool_input else {}),
},
)
elif tool_name == "unsiloed_get_job_result":
time.sleep(5)
resp = requests.get(
f"{UNSILOED_BASE_URL}/{tool_input['job_type']}/{tool_input['job_id']}",
headers=UNSILOED_HEADERS,
)
# Add other tool handlers (parse, classify, split) as needed
else:
return json.dumps({"error": f"Unhandled tool: {tool_name}"})
return json.dumps(resp.json())
# --- Agentic loop ---
client = anthropic.Anthropic(api_key=ANTHROPIC_API_KEY)
messages = [
{
"role": "user",
"content": "Extract the invoice number, date, and total amount from this PDF: https://example.com/invoice.pdf",
}
]
print("Starting agentic loop...\n")
while True:
response = client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=4096,
tools=tools,
messages=messages,
)
# Collect assistant response
messages.append({"role": "assistant", "content": response.content})
# If Claude is done, print the final text and exit
if response.stop_reason == "end_turn":
for block in response.content:
if hasattr(block, "text"):
print(f"Claude: {block.text}")
break
# If Claude wants to use tools, execute them
if response.stop_reason == "tool_use":
tool_results = []
for block in response.content:
if block.type == "tool_use":
print(f" -> Calling {block.name}({json.dumps(block.input, indent=2)[:100]}...)")
result = process_tool_call(block.name, block.input)
tool_results.append(
{
"type": "tool_result",
"tool_use_id": block.id,
"content": result,
}
)
messages.append({"role": "user", "content": tool_results})
print("\nDone.")
```
## Error Handling
When integrating with the Unsiloed API, handle these common scenarios:
If a job's status is `"Failed"` or `"failed"`, the response includes an error message. Parse jobs use capitalized statuses (`Succeeded`, `Failed`), while extraction, classification, and splitting jobs use lowercase (`completed`, `failed`).
```json theme={null}
{
"status": "failed",
"error": "Error processing document: unsupported file format"
}
```
A `401` response means the API key is missing or invalid. Ensure you pass the key in the `api-key` header.
```json theme={null}
{
"detail": "API key is required. Please provide 'api-key' in the request header."
}
```
A `402` response means your organization has run out of credits. Check the `quota_remaining` field in successful responses to monitor usage proactively.
```json theme={null}
{
"detail": {
"message": "Insufficient quota",
"status": "QUOTA_EXCEEDED"
}
}
```
If a job hasn't completed after 5 minutes of polling, treat it as a timeout. Jobs rarely take longer than 2 minutes for standard documents.
## Best Practices
1. **Use descriptive categories** — When classifying or splitting, add `description` fields to your category objects. This significantly improves accuracy.
2. **Poll with backoff** — Wait 5-10 seconds between `unsiloed_get_job_result` calls. Tight polling wastes quota and adds no benefit.
3. **Use publicly accessible URLs** — Claude cannot upload binary files. Use presigned URLs from your cloud storage or any publicly accessible link.
4. **Keep extraction schemas focused** — Smaller, targeted extraction schemas produce better results than large catch-all schemas. Extract what you need.
5. **Handle errors gracefully** — Always check job status before processing results. Return clear error messages so Claude can inform the user.
6. **Monitor your quota** — Check the `quota_remaining` field in API responses and alert when running low.
## Next Steps
Learn about document parsing and structure analysis
Explore document classification with confidence scoring
Split multi-document PDFs into separate files
Full API reference with all endpoints and parameters
# Extract Documents from Databricks with Unsiloed
Source: https://docs.unsiloed.ai/docs/integrations/databricks
Build a Databricks ETL pipeline that reads documents from a Unity Catalog volume, extracts structured fields with Unsiloed, and writes confidence scores and citations to Delta tables.
Databricks reads documents natively with `ai_parse_document` and `ai_extract`.
This guide routes that workload to Unsiloed for a confidence score and a source
citation for every grounded value, so you can triage the values worth checking
instead of trusting them all equally.
Document processing runs inside Databricks. Unity Catalog governs the pipeline,
and you can trigger or schedule it like any other Databricks workload. The only
local tool this guide uses is the Databricks CLI for one-time secret setup.
## Why Call Unsiloed from Databricks
PDFs in a Unity Catalog volume are just bytes. You can list them and count them,
but you can't join them to anything, because none of what matters is in a column
yet.
This guide turns them into a table. Every document becomes a row, and every
extracted field carries a value and two confidence scores. When Unsiloed can ground
a value in the source document, the field also includes a citation pointing to the
region of the page it came from.
That confidence signal is the reason to route the work through Unsiloed. High-
confidence values can flow into downstream analysis, while lower-confidence values
can enter a review queue. Both land in the same table, so you can define a review
policy with a SQL filter.
## How the Pipeline Works
Documents land in a Unity Catalog volume. An ETL pipeline reads them, sends each
one to Unsiloed for extraction, and writes the results back as Delta tables in the
same schema, so the documents and the data pulled out of them stay together.
## What We'll Build
An ETL pipeline with three tables:
* **`extractions`:** one row per document, holding the complete Unsiloed result object in a `VARIANT` column
* **`extracted_fields`:** one row per successfully extracted field, including its confidence scores and citation
* **`extraction_errors`:** one row per document that could not be extracted
The pipeline reads the volume with Auto Loader, so a normal run discovers document
paths that haven't already committed. We'll build the code up a piece at a time
below. If you'd rather skip the walkthrough, take the whole file from the dropdown.
This is the complete file. Paste it into the pipeline's starter file in [Step 4](#step-4-create-and-run-the-pipeline), changing `VOLUME` and `FIELDS` to match your own volume and the fields you want.
```python my_transformation.py theme={null}
"""Turn documents in a Unity Catalog volume into a queryable Delta table with Unsiloed.
Create an ETL pipeline in Databricks, point it at this file, and click Run.
New files dropped into the volume are picked up on the next run.
"""
import json
import mimetypes
import time
import requests
from pyspark import pipelines as dp
from pyspark.sql.functions import col, element_at, split, udf
# --- Configure -------------------------------------------------------------
VOLUME = "/Volumes/workspace/unsiloed/docs"
MAX_FILE_BYTES = 50 * 1024 * 1024
MAX_CONCURRENT_EXTRACTIONS = 4
# The fields to pull from each document. Each description is an instruction to
# the model, so be specific.
FIELDS = {
"vendor_name": "Company that issued the invoice",
"invoice_number": "Invoice number or ID",
"invoice_date": "Issue date as YYYY-MM-DD",
"total_amount": "Grand total payable",
}
# ---------------------------------------------------------------------------
API_KEY = dbutils.secrets.get("unsiloed", "api_key")
BASE = "https://prod.visionapi.unsiloed.ai"
SCHEMA = json.dumps({
"type": "object",
"properties": {k: {"type": "string", "description": v} for k, v in FIELDS.items()},
})
@udf("string")
def unsiloed_extract(file_name, content):
"""Send one document to Unsiloed and return its result as JSON text."""
if content is None:
return json.dumps({"_error": "The document has no content"})
if len(content) > MAX_FILE_BYTES:
return json.dumps({"_error": f"File exceeds the {MAX_FILE_BYTES}-byte limit"})
try:
content_type = mimetypes.guess_type(file_name)[0] or "application/octet-stream"
with requests.Session() as session:
submit = session.post(
f"{BASE}/v2/extract",
headers={"api-key": API_KEY},
files={"pdf_file": (file_name, bytes(content), content_type)},
data={"schema_data": SCHEMA, "model": "gamma", "enable_citations": "true"},
timeout=180,
)
if not submit.ok:
detail = submit.text.replace("\n", " ")[:240]
return json.dumps({"_error": f"Submit HTTP {submit.status_code}: {detail}"})
job_id = submit.json()["job_id"]
deadline = time.monotonic() + 240
while time.monotonic() < deadline:
remaining = deadline - time.monotonic()
poll_response = session.get(
f"{BASE}/extract/{job_id}",
headers={"api-key": API_KEY},
timeout=max(1, min(30, remaining)),
)
if poll_response.status_code == 429:
time.sleep(min(8, max(0, deadline - time.monotonic())))
continue
poll_response.raise_for_status()
poll = poll_response.json()
if poll.get("status") in ("completed", "review"):
return json.dumps(poll.get("result") or {})
if poll.get("status") == "failed":
return json.dumps({"_error": f"Job {job_id} failed: {json.dumps(poll)[:240]}"})
time.sleep(min(4, max(0, deadline - time.monotonic())))
return json.dumps({"_error": f"Job {job_id} timed out after 240 seconds"})
except Exception as e: # one bad document must not fail the pipeline
return json.dumps({"_error": f"{type(e).__name__}: {e}"})
unsiloed_extract = unsiloed_extract.asNondeterministic()
@dp.table(
name="extractions",
comment="One row per document, containing the Unsiloed result object",
)
def extractions():
# Auto Loader tracks which files it has already seen, so each run only
# extracts documents that are new since last time.
return (
spark.readStream.format("cloudFiles")
.option("cloudFiles.format", "binaryFile")
.load(VOLUME)
.repartition(MAX_CONCURRENT_EXTRACTIONS)
.withColumn("file_name", element_at(split(col("path"), "/"), -1))
.withColumn("result_json", unsiloed_extract(col("file_name"), col("content")))
.selectExpr(
"replace(path, 'dbfs:', '') AS path",
"file_name",
"length AS size_bytes",
"try_parse_json(result_json) AS result",
"current_timestamp() AS extracted_at")
)
@dp.table(
name="extracted_fields",
comment="One row per extracted field, so you can triage on confidence",
)
def extracted_fields():
return spark.sql("""
SELECT path, file_name, key AS field,
value:value::string AS value,
value:score.extraction_score::double AS extraction_score,
value:score.grounding_score::double AS grounding_score,
value:citation.page::int AS citation_page,
to_json(value:citation.bbox) AS citation_bbox,
value:citation.page_width::double AS citation_page_width,
value:citation.page_height::double AS citation_page_height
FROM STREAM(extractions), LATERAL variant_explode(result)
WHERE result:_error IS NULL
""")
@dp.table(
name="extraction_errors",
comment="One row per document that Unsiloed could not extract",
)
def extraction_errors():
return spark.sql("""
SELECT path, file_name, result:_error::string AS error, extracted_at
FROM STREAM(extractions)
WHERE result:_error IS NOT NULL
""")
```
## Prerequisites for the Databricks Pipeline
You need four things before starting:
* A Databricks workspace with Unity Catalog and access to a SQL warehouse
* Permission to use your target catalog and create a schema, volume, pipeline,
streaming tables, and materialized views
* An Unsiloed API key from the [Unsiloed dashboard](https://app.unsiloed.ai/playground/Extractor)
* The [Databricks CLI](https://docs.databricks.com/aws/en/dev-tools/cli/),
authenticated with `databricks auth login`
You also need at least one supported document to test. The extractor accepts PDFs,
images, and Office documents. Invoice-like documents work with the example fields
in this guide; change `FIELDS` if you use another document type.
You only need the CLI once, in Step 1, to store your API key. Everything after
that happens in the Databricks UI.
## Step 1: Store Your API Key as a Secret
The pipeline reads your key from a Databricks secret scope, so it never appears
in the code.
Databricks has no menu entry for this page, so open it directly, replacing
`` with your own workspace URL:
```
https:///#secrets/createScope
```
Enter `unsiloed` as the scope name, leave **Manage Principal** set to
**Creator**, and click **Create**. Limiting management to the creator prevents
other workspace users from changing the API key or the scope permissions.
You set the value itself through the [Databricks CLI](https://docs.databricks.com/aws/en/dev-tools/cli/),
the only part of this guide that needs a terminal:
```bash theme={null}
databricks secrets put-secret unsiloed api_key
```
Paste your Unsiloed API key when the CLI prompts for the secret value. The
interactive prompt keeps the key out of your shell history and process list.
Confirm it saved:
```bash theme={null}
databricks secrets list-secrets unsiloed
```
## Step 2: Put Your Documents in a Volume
Documents go in a Unity Catalog volume, which is what Databricks uses for files.
They don't go in a table, and the pipeline reads them straight out of the volume.
In a SQL editor, run:
```sql theme={null}
CREATE SCHEMA IF NOT EXISTS workspace.unsiloed;
CREATE VOLUME IF NOT EXISTS workspace.unsiloed.docs;
```
Substitute your own catalog if you aren't using `workspace`.
In the sidebar, click **Catalog**, then expand your catalog and schema. Volumes
sit under their own **Volumes** node, separate from **Tables**, so open that and
select `docs`. Then click **Upload to this volume**.
The catalog browser needs a running SQL warehouse to list anything. If the tree
spins, start your warehouse and try again.
Drop your files in and click **Upload**. Any mix of PDFs, images, and Office
documents works.
## Step 3: Write the Pipeline
A Databricks ETL pipeline is a Python file that declares tables. You don't write
the orchestration. You write a function per table, decorate it with `@dp.table`,
and Databricks works out the dependency order and what to refresh.
**This is all one file.** The four blocks below are consecutive parts of a single
Python file, not alternatives. Append each one to the end of the last, in order.
That file is `transformations/my_transformation.py`, which Databricks creates for
you when you make the pipeline in [Step 4](#step-4-create-and-run-the-pipeline).
Draft it in your own editor as you read, then paste the finished file in when you
get there. If you'd rather build it up in place, create the pipeline first and
come back to this step.
### 3.1 Imports and Configuration
**Start the file** with the imports and the configuration block.
The `pyspark.pipelines` module declares the tables, and it's available inside a
pipeline without installing anything. We call Unsiloed with `requests`.
Four constants near the top control the pipeline. `VOLUME` is where your documents
live, and `FIELDS` defines what to extract. `MAX_FILE_BYTES` rejects files that are
too large to copy safely through a Python UDF, while `MAX_CONCURRENT_EXTRACTIONS`
bounds how many partitions submit work in parallel. Start with four concurrent
extractions and raise the value only after checking your Unsiloed rate limit.
```python my_transformation.py theme={null}
"""Turn documents in a Unity Catalog volume into a queryable Delta table with Unsiloed.
Create an ETL pipeline in Databricks, point it at this file, and click Run.
New files dropped into the volume are picked up on the next run.
"""
import json
import mimetypes
import time
import requests
from pyspark import pipelines as dp
from pyspark.sql.functions import col, element_at, split, udf
# --- Configure -------------------------------------------------------------
VOLUME = "/Volumes/workspace/unsiloed/docs"
MAX_FILE_BYTES = 50 * 1024 * 1024
MAX_CONCURRENT_EXTRACTIONS = 4
# The fields to pull from each document. Each description is an instruction to
# the model, so be specific.
FIELDS = {
"vendor_name": "Company that issued the invoice",
"invoice_number": "Invoice number or ID",
"invoice_date": "Issue date as YYYY-MM-DD",
"total_amount": "Grand total payable",
}
# ---------------------------------------------------------------------------
API_KEY = dbutils.secrets.get("unsiloed", "api_key")
BASE = "https://prod.visionapi.unsiloed.ai"
SCHEMA = json.dumps({
"type": "object",
"properties": {k: {"type": "string", "description": v} for k, v in FIELDS.items()},
})
```
The `dbutils.secrets.get` call reads the key you stored in Step 1, so the key
itself never appears in the file. The `SCHEMA` constant turns `FIELDS` into the
JSON Schema the API expects, which means adding a field later is a one-line
change.
Use `string` for amounts and dates. Real invoices carry currency symbols,
thousands separators, and European decimal commas that fail numeric coercion.
Cast them in SQL later, where a bad value stays visible instead of being
silently dropped.
### 3.2 Call Unsiloed from a UDF
**Append this below the configuration.** The extraction is an ordinary Python
function wrapped in `@udf`, which turns it into something Spark can apply to every
row. It runs in three phases: check the document, submit it, then poll until the
job finishes. The three blocks below are one continuous function, so paste them
one after the other.
#### Reject documents before spending a call
An empty file or one over the size limit can never extract, so the function
returns an error for those immediately rather than paying for a request that is
certain to fail.
```python my_transformation.py theme={null}
@udf("string")
def unsiloed_extract(file_name, content):
"""Send one document to Unsiloed and return its result as JSON text."""
if content is None:
return json.dumps({"_error": "The document has no content"})
if len(content) > MAX_FILE_BYTES:
return json.dumps({"_error": f"File exceeds the {MAX_FILE_BYTES}-byte limit"})
```
Both guards return the same `_error` shape used everywhere else, which is what
puts the document in the `extraction_errors` table later.
#### Submit the document
The `mimetypes` lookup sets the content type from the file extension, so the same
function handles PDFs, images, and Office documents without special cases. A
`requests.Session` reuses one connection for the submit and every poll that
follows.
```python my_transformation.py theme={null}
try:
content_type = mimetypes.guess_type(file_name)[0] or "application/octet-stream"
with requests.Session() as session:
submit = session.post(
f"{BASE}/v2/extract",
headers={"api-key": API_KEY},
files={"pdf_file": (file_name, bytes(content), content_type)},
data={"schema_data": SCHEMA, "model": "gamma", "enable_citations": "true"},
timeout=180,
)
if not submit.ok:
detail = submit.text.replace("\n", " ")[:240]
return json.dumps({"_error": f"Submit HTTP {submit.status_code}: {detail}"})
job_id = submit.json()["job_id"]
```
Checking `submit.ok` rather than calling `raise_for_status()` keeps a rejected
upload as a returned error instead of an exception, so the document lands in
`extraction_errors` with the HTTP status and response body attached.
#### Poll until the job finishes
Extraction is asynchronous, so the submit only returns a `job_id`. This loop asks
for the result until it arrives or the deadline passes.
```python my_transformation.py theme={null}
deadline = time.monotonic() + 240
while time.monotonic() < deadline:
remaining = deadline - time.monotonic()
poll_response = session.get(
f"{BASE}/extract/{job_id}",
headers={"api-key": API_KEY},
timeout=max(1, min(30, remaining)),
)
if poll_response.status_code == 429:
time.sleep(min(8, max(0, deadline - time.monotonic())))
continue
poll_response.raise_for_status()
poll = poll_response.json()
if poll.get("status") in ("completed", "review"):
return json.dumps(poll.get("result") or {})
if poll.get("status") == "failed":
return json.dumps({"_error": f"Job {job_id} failed: {json.dumps(poll)[:240]}"})
time.sleep(min(4, max(0, deadline - time.monotonic())))
return json.dumps({"_error": f"Job {job_id} timed out after 240 seconds"})
except Exception as e: # one bad document must not fail the pipeline
return json.dumps({"_error": f"{type(e).__name__}: {e}"})
unsiloed_extract = unsiloed_extract.asNondeterministic()
```
The deadline is measured with `time.monotonic()`, which doesn't jump if the
cluster clock changes, and every timeout is clamped to what's left of it so a
slow response can't run past the budget. A `429` means Unsiloed is rate limiting,
so the loop waits and retries rather than treating it as a failure.
Two more things about the function as a whole:
* **`enable_citations` must be `true`.** With citations off, the `alpha`, `beta`,
and `delta` tiers return a legacy flat shape (`{"value": ..., "score": }`)
with no citation, and the `extracted_fields` table would have nothing to read.
* **The UDF is marked nondeterministic.** Unsiloed scores can vary between calls,
so this stops Spark optimizing the function as though the same input always
gives the same output.
### 3.3 The `extractions` Table
**Append this below the UDF.** The `spark.readStream` call with the `cloudFiles`
format is Auto Loader. It tracks committed files, so later runs normally pick up
only new document paths. `repartition` bounds the number of partitions that can
submit documents concurrently.
```python my_transformation.py theme={null}
@dp.table(
name="extractions",
comment="One row per document, containing the Unsiloed result object",
)
def extractions():
# Auto Loader tracks which files it has already seen, so each run only
# extracts documents that are new since last time.
return (
spark.readStream.format("cloudFiles")
.option("cloudFiles.format", "binaryFile")
.load(VOLUME)
.repartition(MAX_CONCURRENT_EXTRACTIONS)
.withColumn("file_name", element_at(split(col("path"), "/"), -1))
.withColumn("result_json", unsiloed_extract(col("file_name"), col("content")))
.selectExpr(
"replace(path, 'dbfs:', '') AS path",
"file_name",
"length AS size_bytes",
"try_parse_json(result_json) AS result",
"current_timestamp() AS extracted_at")
)
```
The `binaryFile` format gives us the raw bytes of each document plus its path. We
derive `file_name` from the path, hand both to the UDF, and store the response
with `try_parse_json` so it lands as a `VARIANT` column.
That last part matters. A `VARIANT` column has no fixed shape, so adding a field
to `FIELDS` later needs no migration.
### 3.4 The Output Tables
**Append this at the end of the file.** The `extractions` table is one row per
document, which is awkward for finding the values you should check. The next table
pivots successful results to one row per field, while the final table keeps failed
documents separate.
```python my_transformation.py theme={null}
@dp.table(
name="extracted_fields",
comment="One row per extracted field, so you can triage on confidence",
)
def extracted_fields():
return spark.sql("""
SELECT path, file_name, key AS field,
value:value::string AS value,
value:score.extraction_score::double AS extraction_score,
value:score.grounding_score::double AS grounding_score,
value:citation.page::int AS citation_page,
to_json(value:citation.bbox) AS citation_bbox,
value:citation.page_width::double AS citation_page_width,
value:citation.page_height::double AS citation_page_height
FROM STREAM(extractions), LATERAL variant_explode(result)
WHERE result:_error IS NULL
""")
@dp.table(
name="extraction_errors",
comment="One row per document that Unsiloed could not extract",
)
def extraction_errors():
return spark.sql("""
SELECT path, file_name, result:_error::string AS error, extracted_at
FROM STREAM(extractions)
WHERE result:_error IS NOT NULL
""")
```
The `variant_explode` function walks whatever is inside `result` without naming a
single field, so `extracted_fields` keeps working unchanged when you revise
`FIELDS`. Each successful row includes the value, both confidence scores, and the
full citation geometry when Unsiloed could ground the value. `extraction_errors`
keeps API and file failures out of confidence-review queries.
That's the whole file. Compare it against the [full pipeline](#what-well-build)
above if you want to check you have every piece in the right order.
## Step 4: Create and Run the Pipeline
With the code written, the rest happens in the Databricks UI. Creating an ETL
pipeline gives you a starter file to paste it into.
In the sidebar, click **Jobs & Pipelines**, then click **ETL pipeline** under
Create new. Databricks creates the pipeline immediately and opens its editor,
with a Python file already in place.
The editor opens with an empty starter file at
`transformations/my_transformation.py`.
Before anything else, click the catalog and schema shown at the top right to
open the **Default location** panel, and set **Default catalog** and **Default
schema** to match your volume (`workspace` and `unsiloed`). The schema defaults
to `default`, so if you leave it there your tables land in the wrong place.
Open `transformations/my_transformation.py`. The file is empty, so the grey
text and the **Create with Genie Code** and **Use sample code** buttons are
editor prompts rather than file content, and there is nothing to delete. Paste
in the file you assembled in Step 3, then click **Run pipeline**.
The first run starts compute, reads every document in the volume, and creates
all three tables. Processing time depends on document length, model load, and
the concurrency limit. Databricks submits up to four partitions in parallel
with the configuration used here.
## Step 5: Query the Results
All three outputs are ordinary Delta tables, so query them from the SQL editor or
any connected tool:
```sql theme={null}
SELECT file_name,
result:vendor_name.value::string AS vendor,
result:total_amount.value::string AS total
FROM workspace.unsiloed.extractions
ORDER BY file_name;
```
That returns one row per document:
```
AmazonWebServices.pdf Amazon Web Services 4.11
FlipkartInvoice.pdf WS Retail Services Pvt. Ltd. 319.00
QualityHosting.pdf QualityHosting AG 34,73
coolblue1.pdf Coolblue B.V. 717,97
oyo.pdf OYO 1939
```
A grounded field arrives as an object carrying its value, both scores, and the
region it came from:
```json theme={null}
{
"total_amount": {
"value": "593.36",
"score": { "extraction_score": 0.52, "grounding_score": 0.92 },
"citation": {
"page": 1,
"bbox": [150, 577, 182, 590],
"page_width": 594.99,
"page_height": 841.89
}
}
}
```
The `bbox` array is `[x0, y0, x1, y1]` in points, measured against the
`page_width` and `page_height` in the same object, so you can scale it to whatever
size you render the page at.
## Review the Uncertain Values
The `extracted_fields` table gives one row per field, so the values worth checking
sort to the top:
```sql theme={null}
SELECT file_name, field, value, extraction_score, grounding_score, citation_page
FROM workspace.unsiloed.extracted_fields
WHERE value IS NOT NULL
ORDER BY extraction_score NULLS LAST;
```
The two scores mean different things:
* **`extraction_score`:** confidence in the value that was read
* **`grounding_score`:** confidence that the value was located in the document, at the region the citation points to
They come apart usefully. A value the model inferred rather than read off the page
scores high on extraction and low on grounding, which is exactly what a human
should confirm. If you ask for a normalized value, such as an ISO currency code
when the invoice prints a symbol, grounding collapses toward zero on every
document, because there is no literal text to cite.
Ambiguity in the document is one cause worth knowing about, because it catches
people out. An invoice printing both `Subtotaal €717,97` (including VAT) and
`Exclusief BTW €593,36` will score its subtotal low whichever one it picks, even
when the value is correct. So treat a flagged value as worth checking rather than
wrong, and note the page it came from.
Scores are not deterministic. Re-extracting the same document with the same
schema can return a different score. Use them to rank and triage, not as a fixed
threshold to assert against.
## Add More Documents
Upload another file under a new path and run the pipeline again. Auto Loader tracks
what it has already committed, so the normal case extracts only the new document
and existing rows keep their original `extracted_at`:
```
AmazonWebServices.pdf 09:40:47
FlipkartInvoice.pdf 09:40:47
QualityHosting.pdf 09:40:47
coolblue1.pdf 09:40:47
oyo.pdf 09:40:47
newarrival.pdf 09:41:53 <- only this one was extracted
```
Open **Schedule** on the pipeline and set a trigger. A run with no newly discovered
files doesn't intentionally submit work to Unsiloed, although Databricks compute
and task retries still have their own cost implications.
Auto Loader ignores overwritten files by default. To process a corrected document,
upload it under a new path or run a full refresh.
To re-extract everything after changing `FIELDS`, open the **Run pipeline** menu
and select **Run pipeline with full table refresh**.
## Troubleshoot the Pipeline
Most problems on a first run come from the schema selector or from network access
rather than from the code itself.
The **Default schema** in the pipeline's **Default location** panel is
`default` until you change it, so tables land there instead. Set it to the
schema you want before the first run, or change it in **Settings** and run a
full refresh.
You changed a table between a materialized view and a streaming table, for
example by removing `spark.readStream`. Drop the existing table and run again.
Check the volume path in `VOLUME`, and that the files are really there. Also
avoid eager actions such as `spark.read.table("...").collect()` inside a table
function. Those run while Databricks plans the pipeline graph, before any table
holds data, so they return nothing and give no error.
The row still lands in `extractions`, the error is also written to
`extraction_errors`, and the rest of the batch continues. Query the error table:
```sql theme={null}
SELECT file_name, error, extracted_at
FROM workspace.unsiloed.extraction_errors
ORDER BY extracted_at DESC;
```
The error includes the HTTP status or job failure detail returned by Unsiloed.
If one document fails repeatedly while its siblings succeed, verify its format
and open it locally before retrying.
Pipeline compute needs outbound access to `prod.visionapi.unsiloed.ai`. Some
workspaces restrict serverless egress at the account level. If requests time
out or fail DNS resolution, ask your account admin about the serverless
network policy.
## What to Know Before Running at Scale
Five things are worth settling before you increase the file or concurrency limits:
* **Billing is per page processed.** A full refresh re-extracts every document in
the volume, so check what that represents before you trigger one.
* **Model choice changes cost and accuracy.** The `gamma` tier used here is the
thorough one and suits most work. Use `delta` for complex contracts and dense
tables, and `alpha` when speed matters more than accuracy on clean documents.
* **Documents leave your workspace.** The pipeline sends each file over HTTPS to
`prod.visionapi.unsiloed.ai`. Confirm that's acceptable for your data before
running this at scale.
* **External submissions are not exactly once.** Auto Loader makes the committed
Delta output incremental, but Spark can retry an HTTP side effect. Use a durable,
idempotent submission service when duplicate jobs are unacceptable.
* **Files are loaded into UDF memory.** The 50 MB limit in this guide leaves room
for Spark, Python, and multipart-upload copies. Raise it only after testing the
memory behavior of your pipeline environment.
## See Also
The `/v2/extract` call on its own, with schema options explained.
Every parameter, model tier, and the full response format.
# Integrate Unsiloed with Google Sheets
Source: https://docs.unsiloed.ai/docs/integrations/google-sheets
Connect Google Sheets to the Unsiloed API with Apps Script, verify the connection, and choose a spreadsheet workflow.
Google Sheets formulas don't make authenticated REST API requests directly. A
spreadsheet-bound Google Apps Script provides the bridge: it reads values or
images from the sheet, calls the Unsiloed API, and writes the result back to
cells.
In this guide, we'll connect a spreadsheet to Unsiloed and verify the API key.
The resulting Apps Script project is the starting point for menu commands,
custom functions, and automated document-processing workflows.
This setup is for a script attached to one spreadsheet. If you need the same
integration across many spreadsheets, package the script as a Google Workspace
add-on instead of copying it into each file.
## How Unsiloed Works with Google Sheets
The integration follows one request path:
`Google Sheets → Apps Script → Unsiloed API → Apps Script → Google Sheets`
Apps Script can turn this connection into several spreadsheet workflows:
* **Menu commands** process a selected cell or range and write results to another
sheet. Use this approach for images, files, and jobs that require polling.
* **Custom functions** send a cell value or public URL to Unsiloed and return a
value directly into the formula cell, such as `=UNSILOED(A2)`. Custom functions
can't edit arbitrary cells, and Google stops them after 30 seconds.
* **Batch workflows** loop over a range and process several documents from one
menu command or trigger. Batch writes reduce calls to the Google Sheets service.
Use a menu command for the first integration. It can request authorization, write
to multiple cells, and wait for asynchronous Unsiloed jobs. Add a custom function
when the input and output fit naturally into a single formula.
## Requirements for the Google Sheets Integration
Before you start, gather:
* A Google account and a spreadsheet you can edit
* An Unsiloed API key from the [Unsiloed dashboard](https://app.unsiloed.ai)
* Permission to run Apps Script in your Google Workspace organization
The script uses Google's `UrlFetchApp` service to call Unsiloed. Google asks you
to authorize external requests the first time you run it.
## Connect Google Sheets to Unsiloed
We'll create a script attached to the spreadsheet, store the API key outside the
source code, and call a read-only endpoint to verify the connection.
Open the spreadsheet, then choose **Extensions → Apps Script**. Google creates
an Apps Script project attached to that spreadsheet and opens `Code.gs`.
A [spreadsheet-bound script](https://developers.google.com/apps-script/guides/bound)
can read the active cell, add menus, display sidebars, and define custom
functions for its spreadsheet.
In the Apps Script editor, open **Project Settings** from the left sidebar.
Under **Script Properties**, click **Add script property**.
Add the following property name and use your API key as its value:
```text theme={null}
UNSILOED_API_KEY
```
Click **Save script properties**. The
[Properties service](https://developers.google.com/apps-script/guides/properties)
stores the key for this script, so you don't need to paste it into `Code.gs`.
Replace the contents of `Code.gs` with the following code:
```javascript Code.gs theme={null}
const UNSILOED_USAGE_URL =
"https://platformbackend.unsiloed.ai/api/v1/org/get_usage";
function unsiloedApiKey() {
const apiKey = PropertiesService.getScriptProperties()
.getProperty("UNSILOED_API_KEY");
if (!apiKey) {
throw new Error("Add UNSILOED_API_KEY to the script properties.");
}
return apiKey;
}
function checkUnsiloedConnection() {
const response = UrlFetchApp.fetch(
UNSILOED_USAGE_URL,
{ headers: { "api-key": unsiloedApiKey() } }
);
const usage = JSON.parse(response.getContentText());
SpreadsheetApp.getActive().toast(
"Connected to " + usage.org_name,
"Unsiloed"
);
}
```
The helper reads the current property value whenever the script makes a
request. The connection check uses the read-only usage endpoint and displays
the organization name in the spreadsheet. The workflow cookbooks use the
separate document-processing API at `prod.visionapi.unsiloed.ai`.
Click **Save**, select `checkUnsiloedConnection` from the function dropdown,
and click **Run**.
Google asks for permission the first time. Choose your account, then review
the two things the script needs: access to this spreadsheet, so it can show
the toast, and permission to connect to an external service, so it can reach
Unsiloed. Click **Allow**.
Return to the spreadsheet after the function finishes. A toast should confirm
the Unsiloed organization connected to the API key.
Anyone who can edit the spreadsheet can also view its bound Apps Script project.
Use a dedicated API key, and don't share the spreadsheet with untrusted editors.
## Choose a Google Sheets Workflow
The connection is now ready for a task-specific script. Start with the table
extraction cookbook, which adds a menu command that sends an in-cell image to
`/parse` and writes the detected table to a second sheet.
Use [`/parse`](/docs/document-processing/parsing/parsing) when the sheet needs the
document's original structure, including tables and text sections. Use
[`/v2/extract`](/docs/document-processing/extraction/extraction) when each row needs
the same typed fields, such as an invoice number, date, and total.
Turn an image in a cell into editable rows with a spreadsheet menu command.
Turn a public chart image into numeric cells with `UNSILOED_CHART`.
Review the parsing workflow, supported files, and response structure.
Return a predictable set of fields by supplying a JSON Schema.
## Google Sheets Constraints
Apps Script runs on Google infrastructure, so its execution rules shape the
integration:
* Custom functions return a value to their formula cell but can't write elsewhere
in the spreadsheet.
* Custom functions never request authorization. They can use supported services,
including `UrlFetchApp`, but can't use services that require access to personal
Google data. Use a menu command when a workflow needs those services.
* Custom functions are stopped at 30 seconds, which is the tightest limit here.
Unsiloed document-processing endpoints return jobs that you then poll. In our
tests, a small job often takes around 20 seconds, so even one can approach the
limit and two sequential jobs are unlikely to finish in time. Keep custom
functions to one small job, and use a menu command, which gets six minutes, for
anything larger.
* A script can only read an image placed **in** a cell. An image floating over the
grid exposes no method that returns its bytes, so no script can reach it.
* Menu commands can update ranges, open dialogs, and run workflows that require
authorization.
* Repeated formulas can create one API request per cell. Prefer a function that
accepts a range or a menu-driven batch for larger datasets.
* Script properties belong to the Apps Script project, not to the spreadsheet.
Copying a spreadsheet copies its script but not its properties, so the copy
starts with no API key and its first run fails until you add one.
See the [Google custom functions guide](https://developers.google.com/apps-script/guides/sheets/functions)
for the current service and execution restrictions.
# MCP Server
Source: https://docs.unsiloed.ai/docs/integrations/mcp-server
Connect Claude (and any MCP client) to Unsiloed for parsing, classification, and structured extraction, all over OAuth.
Unsiloed exposes a **remote Model Context Protocol server** at
`https://mcp.unsiloed.ai/mcp`. Add it once in your MCP client
(Claude.ai, Claude Desktop, Claude Code, Cursor, Lovable, ChatGPT,
VS Code, and more) and your assistant can parse PDFs, classify
documents, and extract structured JSON on your behalf. No API key
pasting required.
## What is the Unsiloed MCP Server?
The Unsiloed MCP Server is a remote
[Model Context Protocol](https://modelcontextprotocol.io) server that gives
Claude (and any MCP-compatible client) direct, authenticated access to
Unsiloed's document-processing tools.
You connect once via OAuth. Claude then has six tools at its disposal:
Convert a PDF, DOCX, PPTX, XLSX, or image into clean, LLM-ready markdown.
Label a document against caller-defined categories (e.g. invoice vs contract vs receipt).
Pull structured JSON from a PDF using a caller-provided JSON Schema.
Poll long-running jobs (`get_parse_status`, `get_classify_status`, `get_extract_status`).
## How to connect
The walkthrough below uses Claude.ai; jump to
[Other MCP clients](#other-mcp-clients) for Claude Code, Cursor, Lovable,
ChatGPT, VS Code, n8n, and Windsurf.
Navigate to **Settings → Connectors → Add custom connector**.
Use the production endpoint:
```
https://mcp.unsiloed.ai/mcp
```
No API key is collected here. Authentication happens via OAuth in the next step.
Claude opens a popup to sign in via Unsiloed.
Use your normal Unsiloed account credentials. If you don't have an account
yet, create one at unsiloed.ai. The
free tier is enough to try it out.
You'll see "Claude wants access to your Unsiloed organization, with scopes:
`parse`, `classify`, `extract`, and `offline_access` (which lets the
connection stay signed in). Click **Allow**.
Back in Claude, the connector card should show **six tools** split into
read-only (the three status pollers) and read/write groups (parse, classify, extract).
Ask Claude in a new chat: *"List the tools available from the Unsiloed connector."*
You should see all six listed by name.
## Other MCP clients
Any MCP client connects the same way: paste the URL, sign in, approve.
```bash theme={null}
claude mcp add --transport http unsiloed https://mcp.unsiloed.ai/mcp
```
Then run `/mcp` inside Claude Code, select **unsiloed**, and choose
**Authenticate** to complete the browser sign-in.
**Settings → MCP & Integrations → New MCP Server**, or add to
`~/.cursor/mcp.json`:
```json theme={null}
{
"mcpServers": {
"unsiloed": { "url": "https://mcp.unsiloed.ai/mcp" }
}
}
```
Cursor triggers the OAuth flow on first use.
**Connectors → Chat connectors → New MCP server** (workspace
admin/owner required on Business and Enterprise plans):
* **Server name**: `Unsiloed`
* **Server URL**: `https://mcp.unsiloed.ai/mcp`
* **Authentication**: OAuth (default)
Click **Add & authorize** and complete the sign-in.
**Settings → Apps → Advanced → Developer mode → Create app**, paste
`https://mcp.unsiloed.ai/mcp`, and pick OAuth when prompted.
Add to `.vscode/mcp.json` (or your user-level MCP config):
```json theme={null}
{
"servers": {
"unsiloed": { "type": "http", "url": "https://mcp.unsiloed.ai/mcp" }
}
}
```
Add an **MCP Client Tool** node with:
* **Server Transport**: `HTTP Streamable` (not the deprecated SSE option)
* **Endpoint**: `https://mcp.unsiloed.ai/mcp`
* **Authentication**: OAuth2
n8n registers itself automatically and walks you through consent.
Clients without native remote-OAuth support can bridge through
[`mcp-remote`](https://github.com/geelen/mcp-remote). For Windsurf,
add to `~/.codeium/windsurf/mcp_config.json`:
```json theme={null}
{
"mcpServers": {
"unsiloed": {
"command": "npx",
"args": ["-y", "mcp-remote@latest", "https://mcp.unsiloed.ai/mcp"]
}
}
}
```
The same snippet works for any client that only launches local
(stdio) MCP servers.
## Signing in
You sign in once with your normal Unsiloed account. No API keys.
On the consent screen you'll be asked to approve four permissions.
**Approve all of them** unless you have a specific reason not to:
| Permission | What it enables |
| - | - |
| `parse` | `parse_document` + `get_parse_status` |
| `classify` | `classify_document` + `get_classify_status` |
| `extract` | `extract_data` + `get_extract_status` |
| `offline_access` | Staying signed in. Without it you'll re-authenticate every 15 minutes |
Good to know:
* An **actively used connection stays signed in indefinitely**. Your
client renews it automatically in the background.
* If a tool refuses with **"Missing required OAuth scope"**, you approved
fewer permissions than that tool needs. Disconnect, reconnect, and
approve everything.
* If your client ever shows **"reconnect required"** (for example after
signing out of Unsiloed), just click Reconnect and approve again.
## The six tools
### Parse
Accepts PDF, DOCX, DOC, PPTX, PPT, XLSX, XLS, PNG, JPEG, and TIFF.
Returns full markdown content plus per-chunk structure and job
metadata.
**Inputs**:
* `file_url` *(string, optional)*: publicly fetchable HTTPS URL. Presigned S3 URLs work.
* `file_base64` *(string, optional)*: base64-encoded file contents (provide either this OR `file_url`).
* `file_name` *(string, optional)*: defaults to `document.pdf`. Required for non-PDF formats when using `file_base64`.
* `mode` *(`fast` | `accurate` | `agentic`)*: see [Processing modes](#processing-modes).
* `page_range` *(string, optional)*: e.g. `"1-5"`, `"2,4,6"`, `"1-3,7,10-12"`. Omit for the whole document.
**Returns**: the merged markdown inline + job metadata (page count, total chunks, credit used, timestamps). For documents over \~30 pages, the call may exceed the tool timeout and return a `status: pending` envelope with a `job_id`. Poll with `get_parse_status`.
Use when `parse_document` returned `status: pending`.
**Inputs**:
* `job_id` *(UUID, required)*.
* `include_chunks` *(boolean, default true)*: set false for a cheap status-only poll.
**Returns**: full markdown + metadata once `status: Succeeded`. Otherwise just metadata.
#### Processing modes
All modes run through Unsiloed's parsing pipeline with smart layout
detection. Pick the one that matches your document complexity and latency
budget:
| Mode | Best for |
| - | - |
| `fast` | Clean born-digital PDFs (system-generated invoices, receipts). Lowest latency and lowest cost. |
| `accurate` | Most real-world documents: scanned PDFs, multi-column layouts, tables spanning pages. |
| `agentic` | Highest fidelity. Use for legal contracts, 10-K/10-Q filings, equations, handwriting. |
### Classify
**Inputs**:
* `file_url` or `file_base64` (one required).
* `categories` *(array, required)*: 1 to 20 objects shaped `{name: string, description?: string}`. Descriptions strongly improve accuracy by giving the classifier label hints.
**Returns**: predicted classification plus per-page confidence scores.
Same shape as `get_parse_status`. Use when `classify_document` returned `status: pending`.
### Extract
**Inputs**:
* `file_url` or `file_base64` (one required).
* `json_schema` *(object, required)*: JSON Schema (draft-07 compatible) defining the desired output. Per-field `description` values are passed to the underlying model as extraction hints.
* `model` *(`alpha` | `beta` | `gamma` | `delta`)*: see [Model tiers](#model-tiers).
* `enable_citations` *(boolean, default false)*: when true, includes bbox coordinates for each extracted value.
**Returns**: the extracted object, typed against your schema, with per-field confidence scores.
Same shape as `get_parse_status`. Use when `extract_data` returned `status: pending`.
#### Model tiers
| Tier | Pick when… |
| - | - |
| `alpha` | Simple key/value or shallow schemas; fastest and cheapest. |
| `beta` | Nested objects and arrays of structured items; mid-tier latency and cost. |
| `gamma` (default) | Strong balance of accuracy and latency. Recommended for production. |
| `delta` | Highest accuracy. Use for complex contracts, dense tables, and strict numerical extraction. |
## Example prompts
Once connected, just talk to Claude naturally. Some prompts to try:
*"Use the Unsiloed connector to parse the PDF at
[https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf](https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf)
and show me the markdown."*
*"Parse the MSA at `https://...` and extract `{parties: string[],
governing_law: string, effective_date: string, termination_clause: string}`
using the `delta` model."*
*"Classify these PDFs as `invoice`, `receipt`, or `purchase_order`, then for each
invoice extract `line_items: [{description, quantity, unit_price, total}]`."*
*"Parse this 10-K filing and extract the income-statement figures with
the schema `{revenue, cogs, gross_profit, operating_expenses, net_income}`."*
## Working with large documents
MCP clients have a limited context window, and many **truncate large tool
results**, so a 100-page parse can come back cut off in the chat even
though the server processed it fully. Work in slices instead of one big
dump:
If you only need specific fields, use `extract_data` with a JSON
Schema. It returns a compact JSON object instead of the full document
markdown, usually 100x smaller and immune to truncation.
Use `page_range` to process a big document in chunks the client can
hold: *"Parse pages 1-20 of this PDF"*, then *"now pages 21-40"*, and
so on. Ask your assistant to summarize or accumulate findings as it
goes rather than echoing full text.
When checking on a long-running job, ask for a status-only poll
(`include_chunks: false`) so the full content isn't re-sent on every
check.
Documents over \~30 pages may return `status: pending` with a
`job_id` instead of an immediate result. That's normal; the job keeps
running. Ask your assistant to poll `get_parse_status` after a few
seconds.
## Limits
* **Supported formats**: PDF, DOCX, DOC, PPTX, PPT, XLSX, XLS, PNG, JPEG,
TIFF.
* **File delivery**: via a publicly fetchable HTTPS URL (presigned S3
URLs work) or inline base64.
* **Long documents** run asynchronously. Expect a `job_id` + polling
instead of an instant answer for anything past \~30 pages.
* **Result size**: the server returns complete results, but your MCP
client may truncate what fits in the conversation. See
[Working with large documents](#working-with-large-documents).
* **Credits**: each processed page consumes credits from your
organization's plan. Track usage in
your dashboard.
## Troubleshooting
Browser pop-up blockers can intercept the OAuth window. Allow pop-ups for
`claude.ai` and try Connect again. If still blocked, try Claude Desktop
instead, which handles the redirect natively without relying on
`window.open`.
The OAuth tokens expire on signout and on admin revocation. Disconnect
the Unsiloed connector and re-add it to mint a fresh pair. No
re-registration is needed; Claude reuses the stored client\_id.
Your organization has been suspended by Unsiloed (usually for billing
reasons). Contact [support@unsiloed.ai](mailto:support@unsiloed.ai)
or visit your dashboard to
resolve. Once reactivated, the next tool call resumes within a minute,
with no need to reconnect.
During consent you approved fewer scopes than the tool needs. Disconnect
the connector and reconnect. When the consent screen appears, approve
all three of `parse`, `classify`, `extract`.
Your organization has exhausted its monthly parse credits. Top up at
your dashboard before
retrying.
Don't just retry the existing entry. Instead, **delete the connector entirely,
then add it again as a new one**. If it persists, verify your network
allows outbound HTTPS to `mcp.unsiloed.ai` and `hydra.unsiloed.ai`.
Disconnect the Unsiloed connector and reconnect once. Connections stay
signed in as long as they're in use; if yours doesn't, a fresh
reconnect resolves it.
Choose **Streamable HTTP** (sometimes labeled `streamable-http` or
just `HTTP`). Do not select SSE.
The server processed the full document; your MCP client truncated the
result to fit its context window. Re-run the request in smaller
slices using `page_range` (for example pages 1-20, then 21-40), or
switch to `extract_data` with a JSON Schema if you only need specific
fields. See [Working with large documents](#working-with-large-documents).
## Support
[support@unsiloed.ai](mailto:support@unsiloed.ai)
[hello@unsiloed.ai](mailto:hello@unsiloed.ai)
## See also
Drop-in Anthropic tool-use schemas for direct API usage without the MCP layer.
Call Unsiloed's parser, classifier, and extractor directly via REST.
# Integrate Unsiloed with Muse
Source: https://docs.unsiloed.ai/docs/integrations/muse
Connect Unsiloed to Meta's Muse agent with one prompt and an API key, then let Muse parse your PDFs and other documents.
[Muse](https://muse.ai) is Meta's personal AI agent. It can call the Unsiloed
API through a custom connector, which you set up from the chat: ask Muse to
connect Unsiloed, then paste your API key into Muse's Secure Vault. Muse
doesn't see the key in the conversation.
## Prerequisites
* **A Muse account.** Log in or create an account at [muse.ai](https://muse.ai).
Muse is also available as a mobile app, a [Mac app](https://ai.meta.com/muse/download/),
and in WhatsApp. The screenshots in this guide are from the Mac app.
* **An Unsiloed API key.** Create one in the
[Unsiloed dashboard](https://app.unsiloed.ai/developer).
## Connect Unsiloed to Muse
Start a chat with Muse and send this prompt:
```text wrap theme={null}
Connect Unsiloed so you can process PDFs and other documents. I need somewhere secure to store the API key.
```
Muse replies with an **Unsiloed** connector card.
Click **Connect** on the connector card. Muse shows what the custom
connector can do. Click **Continue**.
Paste your Unsiloed API key into the **API key** field and click **Add**.
Muse confirms that Unsiloed is connected.
## Process a Document
Attach a PDF to the chat and ask Muse to process it. The first time Muse
calls Unsiloed, it asks permission to share information with
`prod.visionapi.unsiloed.ai`. Choose **Allow once** or **Always allow this
site**.
Muse sends the document to Unsiloed and returns the parsed content as
Markdown, including the document's text, tables, and chart descriptions.
To disconnect Unsiloed, remove the connector in Muse's settings.
# Snowflake
Source: https://docs.unsiloed.ai/docs/integrations/snowflake
Set up Unsiloed in Snowflake: a UDF that extracts staged documents in SQL, plus a Cortex Agent that extracts them in natural language.
Snowflake can extract structured data from documents natively with
`AI_EXTRACT`. This guide routes that workload to Unsiloed by setting up two
pieces, in order:
1. **[A UDF](#extract-in-sql-with-a-udf)** that extracts documents from a
Snowflake stage in SQL. This is the foundation.
2. **[A Cortex Agent](#extract-with-a-cortex-agent)** that extracts those same
staged documents in natural language, by running the UDF through Snowflake's
Managed MCP server.
**Set up the UDF first.** Both paths stay inside Snowflake and run on that
same UDF.
## Why Call Unsiloed from Snowflake
Unsiloed adds schema-driven extraction with per-field confidence scores and
source citations, so you can route uncertain values to a review queue instead of
trusting every extracted value equally. Both pieces fit the Snowflake
architecture you already have and keep the API key in a Snowflake-managed
credential.
## Extract in SQL with a UDF
A Python user-defined function (UDF) over an
[External Access Integration](https://docs.snowflake.com/en/developer-guide/external-network-access/external-network-access-overview)
reads a staged file and sends it to Unsiloed's `/v2/extract`, returning typed
fields with confidence scores and citations. It works on any file in a stage.
For an internal stage, create the stage with
`ENCRYPTION = (TYPE = 'SNOWFLAKE_SSE')`. The UDF reads files through
`SnowflakeFile`, which can't open client-side-encrypted stages.
A ready-to-run setup (network rule, secret, integration, and the UDF) is in the
[Unsiloed cookbook](https://github.com/Unsiloed-AI/cookbook/tree/main/snowflake).
To install it:
In Snowsight, select **+ (Create)** in the top left, then **SQL File**.
Copy `setup.sql` from the cookbook into the worksheet. Set your database and
schema at the top, and paste your Unsiloed API key into the secret's
`SECRET_STRING`.
Open the **Run** dropdown next to the Run button and choose **Run All**. This
creates the network rule, secret, external access integration, and the
`unsiloed_extract` function.
Then call the function with a scoped file URL, the file's name, and a JSON
Schema:
```sql theme={null}
SELECT unsiloed_extract(
BUILD_SCOPED_FILE_URL(@MY_DB.MY_SCHEMA.PDF_STAGE, 'invoice.pdf'),
'invoice.pdf',
'{ "type": "object", "properties": {
"vendor": { "type": "string" }, "total": { "type": "string" } } }'
) AS result;
```
Pass the name as well as the URL. Unsiloed decides how to decode the file from
its extension, and a scoped file URL is encrypted, so the UDF can't recover the
name from it.
Each field comes back as an object with a value, a confidence score, and a
citation:
```json theme={null}
{
"total": {
"value": "4.11",
"score": { "grounding_score": 0.998, "extraction_score": 0.989 },
"citation": {
"page": 1,
"bbox": [548, 142, 573, 156],
"page_width": 612.0,
"page_height": 792.0
}
}
}
```
The `bbox` array holds `[x0, y0, x1, y1]` in points, measured against the
`page_width` and `page_height` in the same object, so you can scale it to
whatever size you render the page at.
To process a whole stage in one query, join to `DIRECTORY(@stage)`, which gives
you `relative_path` for both the URL and the name:
```sql theme={null}
SELECT relative_path,
unsiloed_extract(
BUILD_SCOPED_FILE_URL(@MY_DB.MY_SCHEMA.PDF_STAGE, relative_path),
relative_path,
'{ "type": "object", "properties": { "total": { "type": "string" } } }'
) AS result
FROM DIRECTORY(@MY_DB.MY_SCHEMA.PDF_STAGE);
```
Unsiloed also accepts an array-of-objects schema, so repeating rows (line items,
holdings, table rows) need no columnar rewrite.
## Extract with a Cortex Agent
With the UDF in place, a Snowflake Cortex Agent (the engine behind Snowflake
CoWork) can extract your staged documents in natural language. The agent runs
the UDF through Snowflake's **Managed MCP server**: its `execute_sql` tool
lets the agent call `unsiloed_extract` on any file in your stage, so it reaches
documents inside Snowflake without a URL.
### Prerequisites for the Cortex Agent
* A Snowflake account with Cortex Agents and Snowflake CoWork available.
* A role that can create MCP servers and agents (`ACCOUNTADMIN` works).
* The `unsiloed_extract` UDF from the previous section.
### Set Up the Agent
Expose a read-only `execute_sql` tool through Snowflake's Managed MCP. The
server runs queries under your default warehouse, so make sure one is set
(`ALTER USER SET DEFAULT_WAREHOUSE = ''`).
```sql theme={null}
CREATE OR REPLACE MCP SERVER unsiloed_sql_mcp FROM SPECIFICATION $$
tools:
- name: "execute_sql"
type: "SYSTEM_EXECUTE_SQL"
title: "Execute SQL"
description: "Run read-only SQL, including the unsiloed_extract UDF"
$$;
```
Reference the Managed MCP server in a top-level `mcp_servers` block, and tell
the agent how to call the UDF. Fully-qualify `unsiloed_extract`, since
`execute_sql` runs without a database or schema context. The orchestration
model must be on the agent allowlist (for example `claude-sonnet-4-5`,
`claude-sonnet-5`, or `auto`).
```sql theme={null}
CREATE OR REPLACE AGENT invoice_agent FROM SPECIFICATION $$
{
"models": { "orchestration": "claude-sonnet-4-5" },
"instructions": {
"response": "You extract structured data from documents in Snowflake stages, then report each value with its confidence score.",
"orchestration": "When the user names a document in a stage, call execute_sql with SELECT MY_DB.MY_SCHEMA.unsiloed_extract(BUILD_SCOPED_FILE_URL(@MY_DB.MY_SCHEMA.PDF_STAGE, ''), '', ''), passing the same filename twice and building from the requested fields."
},
"mcp_servers": [
{ "server_spec": { "name": "MY_DB.MY_SCHEMA.unsiloed_sql_mcp" } }
]
}
$$;
```
### Use the Agent
Open Snowflake CoWork at [ai.snowflake.com](https://ai.snowflake.com), start
a new chat, and pick your agent from the selector in the message box.
Ask in natural language, naming a file in your stage:
```
Extract the vendor, invoice number, and total from AmazonWebServices.pdf
in the PDF_STAGE.
```
The agent runs the UDF through `execute_sql` and reports the fields, each
with a confidence score:
## Troubleshoot the Cortex Agent
`execute_sql` runs without a database or schema context. Fully-qualify the
function in the agent's instructions: `MY_DB.MY_SCHEMA.unsiloed_extract(...)`.
Reference the MCP server in a top-level `mcp_servers` block, not under
`tools`/`tool_spec`. Also confirm the orchestration model is on the agent
allowlist (for example `claude-sonnet-4-5`, `claude-sonnet-5`, or `auto`;
`claude-4-sonnet` is not valid).
Make the orchestration instruction explicit: give the exact
`SELECT unsiloed_extract(BUILD_SCOPED_FILE_URL(@stage, ''), '', '')`
template and the stage name, so the agent maps the document name to a stage
path.
The agent dropped an argument, usually the file name. The function takes the
scoped URL, the file name, and the schema, so the filename appears twice in
the template. Restate it in the orchestration instruction.
The Managed MCP runs queries under your default warehouse. Set one with
`ALTER USER SET DEFAULT_WAREHOUSE = ''` and make sure it can
resume.
## See Also
Ready-to-run `setup.sql` for the in-SQL UDF path.
The `/v2/extract` endpoint, schema options, and response format.
# Quickstart
Source: https://docs.unsiloed.ai/docs/quickstart
Submit a document to the /parse endpoint and read back structured Markdown chunks.
This quickstart covers the **parsing** endpoint and is the fastest way to try Unsiloed AI. If you'd rather start with another capability, see the [Extraction quickstart](/docs/document-processing/extraction/quickstart), the [Classification guide](/docs/document-processing/classification/classification), or the [Splitting guide](/docs/document-processing/splitting/splitting).
By the end of this guide, you'll have a working script that uploads a PDF to the `/parse` endpoint, polls until parsing finishes, and saves the parsed result to disk as both JSON and Markdown. The full script is available in the dropdown below if you'd rather copy it and skip the walkthrough.
Set `UNSILOED_API_KEY` in your environment and save the document you want to parse as `document.pdf` in the same directory before running.
```python parse_document.py theme={null}
import json
import os
import time
import requests
API_KEY = os.environ["UNSILOED_API_KEY"]
BASE_URL = "https://prod.visionapi.unsiloed.ai"
with open("document.pdf", "rb") as f:
response = requests.post(
f"{BASE_URL}/parse",
headers={"api-key": API_KEY},
files={"file": ("document.pdf", f, "application/pdf")},
)
response.raise_for_status()
job_id = response.json()["job_id"]
print(f"Job submitted: {job_id}")
max_attempts = 60 # roughly 5 minutes at 5 seconds per poll
attempts = 0
while True:
result = requests.get(
f"{BASE_URL}/parse/{job_id}",
headers={"api-key": API_KEY},
).json()
print(f"Status: {result['status']}")
if result["status"] == "Succeeded":
break
if result["status"] in ("Failed", "Cancelled"):
raise RuntimeError(result.get("message", f"parse job {result['status'].lower()}"))
attempts += 1
if attempts >= max_attempts:
raise TimeoutError("Parse job did not finish within 5 minutes")
time.sleep(5)
with open("result.json", "w") as f:
json.dump(result, f, indent=2)
with open("output.md", "w") as f:
f.write("\n\n".join(chunk["embed"] for chunk in result["chunks"]))
print(f"Saved {result['total_chunks']} chunks to result.json and output.md")
```
Save this as `script.mjs` or set `"type": "module"` in your `package.json`. Requires Node.js 18 or newer for the global `fetch`, `FormData`, and `Blob`.
```javascript script.mjs theme={null}
import fs from "node:fs";
const API_KEY = process.env.UNSILOED_API_KEY;
const BASE_URL = "https://prod.visionapi.unsiloed.ai";
const form = new FormData();
form.append("file", new Blob([fs.readFileSync("document.pdf")]), "document.pdf");
const response = await fetch(`${BASE_URL}/parse`, {
method: "POST",
headers: { "api-key": API_KEY },
body: form,
});
if (!response.ok) throw new Error(`${response.status}: ${await response.text()}`);
const { job_id } = await response.json();
console.log(`Job submitted: ${job_id}`);
const maxAttempts = 60; // roughly 5 minutes at 5 seconds per poll
let attempts = 0;
let result;
while (true) {
const res = await fetch(`${BASE_URL}/parse/${job_id}`, {
headers: { "api-key": API_KEY },
});
result = await res.json();
console.log(`Status: ${result.status}`);
if (result.status === "Succeeded") break;
if (["Failed", "Cancelled"].includes(result.status)) {
throw new Error(result.message || `parse job ${result.status.toLowerCase()}`);
}
if (++attempts >= maxAttempts) throw new Error("Parse job did not finish within 5 minutes");
await new Promise((r) => setTimeout(r, 5000));
}
fs.writeFileSync("result.json", JSON.stringify(result, null, 2));
fs.writeFileSync("output.md", result.chunks.map((c) => c.embed).join("\n\n"));
console.log(`Saved ${result.total_chunks} chunks to result.json and output.md`);
```
```bash theme={null}
# Submit the document and capture the job_id from the response:
resp=$(curl -sS -X POST "https://prod.visionapi.unsiloed.ai/parse" \
-H "api-key: $UNSILOED_API_KEY" \
-F "file=@document.pdf")
JOB_ID=$(printf '%s\n' "$resp" | jq -er '.job_id') || {
printf 'Submission did not return a job_id:\n%s\n' "$resp" >&2
exit 1
}
echo "Job submitted: $JOB_ID"
# Poll until the job finishes, with a 5-minute timeout:
attempts=0
max_attempts=60
while true; do
resp=$(curl -sS -X GET "https://prod.visionapi.unsiloed.ai/parse/$JOB_ID" \
-H "api-key: $UNSILOED_API_KEY")
status=$(printf '%s\n' "$resp" | jq -er '.status') || {
printf 'Status response did not contain a status:\n%s\n' "$resp" >&2
exit 1
}
echo "Status: $status"
[ "$status" = "Succeeded" ] && break
{ [ "$status" = "Failed" ] || [ "$status" = "Cancelled" ]; } && { echo "Job $status"; exit 1; }
attempts=$((attempts + 1))
[ "$attempts" -ge "$max_attempts" ] && { echo "Parse job did not finish within 5 minutes"; exit 1; }
sleep 5
done
# Save the full response to disk:
printf '%s\n' "$resp" > result.json
```
## Step 1: Set Up Your Environment
Before writing any code, we need three things: an API key, a document, and the runtime for our chosen language.
### 1.1 Get an Unsiloed AI API Key
To get API access, [sign up on Unsiloed AI](https://cal.com/aman-mishra-p0ry57/15min). Export your key as an environment variable named `UNSILOED_API_KEY` so it stays out of source control:
```bash theme={null}
export UNSILOED_API_KEY="your-api-key"
```
### 1.2 Pick a Document to Parse
The `/parse` endpoint supports PDF, DOCX, PPTX, JPG, PNG, and other formats. The walkthrough below assumes a PDF saved as `document.pdf` in your working directory. To use a different format, update the filename and content type in the snippets to match your file.
If you don't have a document handy, download our [sample PDF](https://raw.githubusercontent.com/Unsiloed-AI/cookbook/c585446e46e4be2790c6c29fe2a7a3a1b346191d/sample-documents/sample.pdf) (a one-page Q1 2024 Sales Report) and save it as `document.pdf`.
### 1.3 Install Dependencies
You need Python 3.8 or newer. Install the `requests` package:
```bash theme={null}
pip install requests
```
You need Node.js 18 or newer for the global `fetch`, `FormData`, and `Blob`. No external packages needed.
You need cURL and `jq`. cURL is preinstalled on macOS and most Linux distributions. Install [`jq`](https://jqlang.org/download/) if it isn't already available.
## Step 2: Submit a Document
The `/parse` endpoint accepts a multipart upload and returns a `job_id` we can poll for results. All requests go to `https://prod.visionapi.unsiloed.ai` with the API key in the `api-key` header.
### 2.1 Set Up the Script
Create a file called `parse_document.py` and start with the imports and configuration:
```python parse_document.py theme={null}
import json
import os
import time
import requests
API_KEY = os.environ["UNSILOED_API_KEY"]
BASE_URL = "https://prod.visionapi.unsiloed.ai"
```
`API_KEY` reads your key from the environment so it doesn't get hard-coded into the file, and `BASE_URL` points at the Unsiloed AI production endpoint. We'll reuse both in every request below.
Create a file called `script.mjs` and start with the imports and configuration:
```javascript script.mjs theme={null}
import fs from "node:fs";
const API_KEY = process.env.UNSILOED_API_KEY;
const BASE_URL = "https://prod.visionapi.unsiloed.ai";
```
`API_KEY` reads your key from the environment so it doesn't get hard-coded into the file, and `BASE_URL` points at the Unsiloed AI production endpoint. We'll reuse both in every request below.
cURL doesn't need a setup step. Each command below inlines the API key and base URL directly.
### 2.2 Upload the Document
Send the file as a multipart upload to `/parse`. The endpoint expects the document under the form field name `file`.
Continue the file by uploading the document:
```python parse_document.py theme={null}
with open("document.pdf", "rb") as f:
response = requests.post(
f"{BASE_URL}/parse",
headers={"api-key": API_KEY},
files={"file": ("document.pdf", f, "application/pdf")},
)
response.raise_for_status()
```
The `raise_for_status()` call throws an `HTTPError` on any non-2xx response, so we don't need to check `.status_code` ourselves.
Continue the file by uploading the document:
```javascript script.mjs theme={null}
const form = new FormData();
form.append("file", new Blob([fs.readFileSync("document.pdf")]), "document.pdf");
const response = await fetch(`${BASE_URL}/parse`, {
method: "POST",
headers: { "api-key": API_KEY },
body: form,
});
if (!response.ok) throw new Error(`${response.status}: ${await response.text()}`);
```
`fetch` doesn't throw on non-2xx responses by default, so we check `response.ok` and raise the error ourselves.
Run:
```bash theme={null}
curl -X POST "https://prod.visionapi.unsiloed.ai/parse" \
-H "api-key: $UNSILOED_API_KEY" \
-F "file=@document.pdf"
```
The response prints to stdout. We need the `job_id` field for the next step.
### 2.3 Capture the Job ID
Next, read and print the `job_id`:
```python parse_document.py theme={null}
job_id = response.json()["job_id"]
print(f"Job submitted: {job_id}")
```
Run the script:
```bash theme={null}
python parse_document.py
```
The output should be a single line like `Job submitted: 1699d429-9c2e-464e-b311-d4b68a8444b8`.
Next, read and log the `job_id`:
```javascript script.mjs theme={null}
const { job_id } = await response.json();
console.log(`Job submitted: ${job_id}`);
```
Run the script:
```bash theme={null}
node script.mjs
```
The output should be a single line like `Job submitted: 1699d429-9c2e-464e-b311-d4b68a8444b8`.
The response body from the POST above looks like:
```json theme={null}
{
"job_id": "1699d429-9c2e-464e-b311-d4b68a8444b8",
"status": "Starting"
}
```
Copy the `job_id` value; you'll paste it into the polling command in the next step.
## Step 3: Poll for Results
The job runs asynchronously. We GET `/parse/{job_id}` repeatedly until the status is `Succeeded`, then save the parsed output to disk.
A status of `Succeeded` means the result is ready. `Failed` means the job errored, and `Cancelled` means processing stopped before completion. Other values such as `Starting`, `Queued`, and `Processing` mean the job is still running.
### 3.1 Write the Polling Loop
Then drop in a polling loop. The `max_attempts` cap stops the loop if the job hangs:
```python parse_document.py theme={null}
max_attempts = 60 # roughly 5 minutes at 5 seconds per poll
attempts = 0
while True:
result = requests.get(
f"{BASE_URL}/parse/{job_id}",
headers={"api-key": API_KEY},
).json()
print(f"Status: {result['status']}")
if result["status"] == "Succeeded":
break
if result["status"] in ("Failed", "Cancelled"):
raise RuntimeError(result.get("message", f"parse job {result['status'].lower()}"))
attempts += 1
if attempts >= max_attempts:
raise TimeoutError("Parse job did not finish within 5 minutes")
time.sleep(5)
```
Then drop in a polling loop. The `maxAttempts` cap stops the loop if the job hangs:
```javascript script.mjs theme={null}
const maxAttempts = 60; // roughly 5 minutes at 5 seconds per poll
let attempts = 0;
let result;
while (true) {
const res = await fetch(`${BASE_URL}/parse/${job_id}`, {
headers: { "api-key": API_KEY },
});
result = await res.json();
console.log(`Status: ${result.status}`);
if (result.status === "Succeeded") break;
if (["Failed", "Cancelled"].includes(result.status)) {
throw new Error(result.message || `parse job ${result.status.toLowerCase()}`);
}
if (++attempts >= maxAttempts) throw new Error("Parse job did not finish within 5 minutes");
await new Promise((r) => setTimeout(r, 5000));
}
```
Replace `JOB_ID` below with the value you captured from Step 2.3, then run this loop. It polls every 5 seconds and gives up after 5 minutes if the job hasn't completed:
```bash theme={null}
JOB_ID="paste-job-id-here"
attempts=0
max_attempts=60 # roughly 5 minutes at 5 seconds per poll
while true; do
resp=$(curl -sS -X GET "https://prod.visionapi.unsiloed.ai/parse/$JOB_ID" \
-H "api-key: $UNSILOED_API_KEY")
status=$(printf '%s\n' "$resp" | jq -er '.status') || {
printf 'Status response did not contain a status:\n%s\n' "$resp" >&2
exit 1
}
echo "Status: $status"
[ "$status" = "Succeeded" ] && break
{ [ "$status" = "Failed" ] || [ "$status" = "Cancelled" ]; } && { echo "Job $status"; exit 1; }
attempts=$((attempts + 1))
[ "$attempts" -ge "$max_attempts" ] && { echo "Parse job did not finish within 5 minutes"; exit 1; }
sleep 5
done
```
The loop keeps the latest response body in `$resp` for the next step.
### 3.2 Save the Parsed Output
Persist the result to disk so downstream code can read it. We'll write two files: `result.json` (the full response, including job metadata and segment-level layout) and `output.md` (the concatenated Markdown, suitable for previewing or feeding into a RAG pipeline).
Finally, write the result to disk:
```python parse_document.py theme={null}
with open("result.json", "w") as f:
json.dump(result, f, indent=2)
with open("output.md", "w") as f:
f.write("\n\n".join(chunk["embed"] for chunk in result["chunks"]))
print(f"Saved {result['total_chunks']} chunks to result.json and output.md")
```
Run the script:
```bash theme={null}
python parse_document.py
```
You should see a few `Status: Processing` lines, then `Status: Succeeded`, then a final summary line. The two output files appear in the working directory.
Finally, write the result to disk:
```javascript script.mjs theme={null}
fs.writeFileSync("result.json", JSON.stringify(result, null, 2));
fs.writeFileSync("output.md", result.chunks.map((c) => c.embed).join("\n\n"));
console.log(`Saved ${result.total_chunks} chunks to result.json and output.md`);
```
Run the script:
```bash theme={null}
node script.mjs
```
You should see a few `Status: Processing` lines, then `Status: Succeeded`, then a final summary line. The two output files appear in the working directory.
The polling loop in Step 3.1 left the full response in `$resp`. Write it to disk:
```bash theme={null}
printf '%s\n' "$resp" > result.json
```
The `result.json` file now holds the full response. The parsed Markdown for each chunk sits under `chunks[].embed`.
## Error Responses
Unsuccessful requests fall into three groups: HTTP errors raised before the job is queued, a `Failed` status on a job that couldn't complete, and a `Cancelled` status on a job stopped before completion.
### HTTP Errors
The `/parse` endpoint returns plain-text bodies on HTTP errors, not JSON. Calling `response.json()` on them raises, so check the status code before parsing. The common cases:
* **`401 Unauthorized`:** body is `Invalid API key: Invalid API key`. The `api-key` header is missing or wrong.
* **`400 Bad Request`:** body is `Content type error`. The upload form is malformed, usually because the file field isn't multipart.
* **`404 Not Found`:** body is `Task not found`. The `job_id` you polled doesn't exist.
### Failed Jobs
A job that was accepted but couldn't be processed comes back with `status: "Failed"`. The response shape matches a successful one, but `chunks` is empty and the `message` field describes what went wrong. For example, submitting a corrupt PDF returns:
```json theme={null}
{
"job_id": "7b31a7d7-e810-4a0b-931e-fbed0879bab2",
"status": "Failed",
"file_name": "bad.pdf",
"message": "Failed to initialize task",
"chunks": [],
"total_chunks": 0
}
```
## Response Shape
A successful response contains job metadata plus a `chunks[]` array. Each chunk has an `embed` Markdown string and an array of layout `segments` with their original positions on the page.
```json theme={null}
{
"job_id": "1699d429-9c2e-464e-b311-d4b68a8444b8",
"status": "Succeeded",
"file_name": "document.pdf",
"page_count": 1,
"total_chunks": 1,
"credit_used": 1,
"chunks": [
{
"chunk_id": "6b2eca3a-d14f-4164-ba9a-0a3a58fcaf45",
"embed": "## Q1 2024 Sales Report\nThe following table summarises...",
"segments": [
{
"segment_id": "c60d89b1-373e-428d-9950-544e7c903b61",
"segment_type": "SectionHeader",
"content": "Q1 2024 Sales Report",
"markdown": "## Q1 2024 Sales Report",
"html": "
Q1 2024 Sales Report
",
"bbox": { "left": 427.6, "top": 67.8, "width": 344.7, "height": 36.5 },
"page_number": 1,
"image": "https://s3.us-east-1.amazonaws.com/...",
"ocr": [ "..." ],
"confidence": 0.35
}
]
}
],
"pdf_url": "https://s3.us-east-1.amazonaws.com/..."
}
```
The fields you'll actually use depend on what you're building. They fall into three broad categories:
**For RAG and embeddings:**
* **`chunks[].embed`:** the chunk's content rolled up as Markdown, ready to pass to an embedder. This is the field the walkthrough writes to `output.md`.
**For layout, source highlighting, and visual overlays:**
* **`chunks[].segments[]`:** the layout primitives the chunk is built from
* **`segments[].segment_type`:** the region's type, for example `Text`, `Table`, or `SectionHeader` (see [Element Types](/docs/document-processing/parsing/element-types) for the full set)
* **`segments[].bbox`:** the segment's position on its page; pair it with `page_number` to identify which page
* **`segments[].markdown`, `html`, `content`:** the segment rendered as Markdown, HTML, or plain text
* **`segments[].image`:** signed URL to a cropped image of the segment
* **`segments[].ocr`:** word-level OCR boxes for highlighting matches in the source PDF
* **`segments[].confidence`:** the parser's confidence in this segment's classification, on a 0-1 scale. Lower values flag ambiguous regions but don't necessarily mean the content is wrong, so treat it as a debugging signal rather than a hard threshold.
**For job and usage tracking:**
* **`status`:** `Succeeded`, `Failed`, `Cancelled`, or one of the in-progress values (`Starting`, `Queued`, `Processing`, and so on)
* **`total_chunks`:** number of chunks in the result
* **`credit_used`:** credits consumed by this job
### Sample Markdown Output
Running the script with the [sample PDF](https://raw.githubusercontent.com/Unsiloed-AI/cookbook/c585446e46e4be2790c6c29fe2a7a3a1b346191d/sample-documents/sample.pdf) writes this to `output.md`:
```markdown theme={null}
## Q1 2024 Sales Report
The following table summarises regional sales performance for Q1
2024.
| Region | Sales Rep | Units Sold | Revenue ($) | Target ($) | % of Target |
| --- | --- | --- | --- | --- | --- |
| North | Alice Brown | 1,240 | 186,000 | 175,000 | 106% |
| South | Bob Smith | 980 | 147,000 | 160,000 | 92% |
| East | Carol Jones | 1,510 | 226,500 | 200,000 | 113% |
| West | David Lee | 870 | 130,500 | 150,000 | 87% |
| Central | Eve Martinez | 1,100 | 165,000 | 155,000 | 106% |
```
This is `chunks[].embed` joined with blank lines. The parser keeps headings, paragraphs, and tables as Markdown, so the output is ready to embed for RAG without further processing.
## Next Steps
For more on parsing, including element types, processing modes, response format, and presigned URLs, see the [Parsing overview](/docs/document-processing/parsing/parsing).
Configure chunking strategies, segment filters, and the OCR backend.
Pull typed fields out of a document using a JSON schema.
Browse the full request and response specs for every endpoint.
Check limits, supported formats, and answers to common questions.