> ## Documentation Index
> Fetch the complete documentation index at: https://docs.unsiloed.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Capacity and Backpressure

> What a 503 from the API means, how batch extraction counts against your limits, and how to poll jobs responsibly

## Overview

Unsiloed rejects requests for two different reasons, and the status code tells you which one applies:

* **`429 Too Many Requests`** comes from **your organization's policy**. You sent requests faster than your plan's per-second limit. See [Rate Limits](/docs/api-reference/limits/rate-limits).
* **`503 Service Unavailable`** comes from **platform capacity**. The platform is busy right now, whatever your own request rate. It is not caused by your organization's usage.

Both responses include a `Retry-After` header. In both cases the request was rejected before a job was created, so you can safely resubmit it once the wait is over.

<Note>
  Unsiloed does not queue excess requests on your behalf. A request is either accepted and processed, or it is rejected with a `429` or `503` that tells you when to try again. Accepted jobs run to completion.
</Note>

## Platform busy

Before a new job is accepted, the API checks how many jobs are in progress across the platform. If that number is at its limit, the submission is rejected with a `503` and a `Retry-After` between **5 and 60 seconds**. The busier the platform, the longer the suggested wait.

This check applies to job submissions: [`POST /v2/extract`](/docs/api-reference/extraction/extract-data), [`POST /classify`](/docs/api-reference/classification/classify-document), [`POST /splitter`](/docs/api-reference/splitting/split-document), and batch extraction. Jobs that were already accepted keep running and are not affected.

```http theme={null}
HTTP/1.1 503 Service Unavailable
Content-Type: application/json
Retry-After: 12
```

```json theme={null}
{
  "error": {
    "code": "service_unavailable",
    "message": "Server is overloaded. Retry after 12s.",
    "details": {
      "reason": "queue_full",
      "retry_after": 12
    }
  }
}
```

`error.details.reason` is `queue_full` and `error.details.retry_after` matches the `Retry-After` header.

## Platform-wide capacity limit

The API also has a platform-wide capacity limit on total request volume across all customers. It protects the service during unusual traffic spikes, and you are unlikely to see it in normal use. When it is reached, any request to the processing and job status endpoints returns a `503` with `Retry-After: 5`.

```http theme={null}
HTTP/1.1 503 Service Unavailable
Content-Type: application/json
Retry-After: 5
```

```json theme={null}
{
  "error": {
    "code": "service_unavailable",
    "message": "Service temporarily at capacity. Please retry."
  }
}
```

## Responding to a 503

1. Wait at least the number of seconds in `Retry-After`.
2. Resubmit the same request. No job was created, so this does not cause duplicates.
3. If you get another `503`, keep waiting for `Retry-After` and add backoff with jitter so many clients do not retry at the same moment.
4. Cap the number of retries and surface an error if the platform stays busy.

The retry helper in the [API FAQ](/docs/faq/api#how-do-i-handle-rate-limiting) handles both `429` and `503` this way.

<Warning>
  Do not retry a `503` immediately or in a tight loop. It delays recovery for everyone, including your own jobs that are already running.
</Warning>

## Batch extraction

`POST /batch/extract` submits several files in one request. Two rules apply:

* **At most 20 files per batch.** A larger batch is rejected with `400` and the message `At most 20 files per batch`.
* **Each file counts as one extraction request** against your organization's Extraction limit. A batch of 10 files uses the same allowance as 10 separate `POST /v2/extract` calls.

If a batch needs more requests than your remaining Extraction allowance, the whole batch is rejected with a `429` and **no jobs are created**. The files counted before the limit was reached still use up allowance, so wait for `Retry-After` before resubmitting.

If a batch is larger than your plan's Extraction allowance, keep batches at or below that allowance, or submit files individually and pace them.

If the platform is busy, a batch is rejected with the same `503` described in [Platform busy](#platform-busy) before any of its files are counted against your limit.

## Polling job status

Extraction, classification, and splitting run asynchronously: you submit a job, receive a `job_id`, and poll for the result.

* **Poll at a steady interval.** Every 5 seconds is a good default; jobs usually take seconds to minutes, so polling faster does not get you results sooner.
* **Set a time limit.** Stop polling after a maximum number of attempts and treat the job as timed out on your side.
* **Honor `Retry-After`.** Status endpoints have no per-organization rate limit, but they are covered by the platform-wide capacity limit. If a status request returns `503`, wait for `Retry-After` before polling again instead of counting it as a failed job.
* **Spread out many jobs.** When tracking a large number of jobs, stagger the polls instead of checking every job at the same instant.

<CodeGroup>
  ```python Python theme={null}
  import os
  import random
  import time

  import requests

  API_KEY = os.environ["UNSILOED_API_KEY"]
  BASE_URL = "https://prod.visionapi.unsiloed.ai"


  def wait_for_extraction(job_id, poll_seconds=5, max_attempts=60):
      """Poll an extraction job until it completes, fails, or times out (about 5 minutes)."""
      for _ in range(max_attempts):
          response = requests.get(f"{BASE_URL}/extract/{job_id}", headers={"api-key": API_KEY})

          if response.status_code in (429, 503):
              retry_after = response.headers.get("Retry-After", "")
              time.sleep(int(retry_after) if retry_after.isdigit() else poll_seconds)
              continue
          response.raise_for_status()

          job = response.json()
          if job["status"] in ("completed", "review"):
              return job
          if job["status"] in ("failed", "cancelled"):
              raise RuntimeError(job.get("error", f"extraction job {job['status']}"))
          time.sleep(poll_seconds + random.uniform(0, 1))

      raise TimeoutError(f"Job {job_id} did not finish in time")
  ```

  ```javascript JavaScript theme={null}
  const API_KEY = process.env.UNSILOED_API_KEY;
  const BASE_URL = "https://prod.visionapi.unsiloed.ai";

  const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms));

  // Poll an extraction job until it completes, fails, or times out (about 5 minutes).
  async function waitForExtraction(jobId, { pollMs = 5000, maxAttempts = 60 } = {}) {
    for (let attempt = 0; attempt < maxAttempts; attempt++) {
      const response = await fetch(`${BASE_URL}/extract/${jobId}`, {
        headers: { "api-key": API_KEY },
      });

      if (response.status === 429 || response.status === 503) {
        const retryAfter = Number.parseInt(response.headers.get("Retry-After") ?? "", 10);
        await sleep(Number.isNaN(retryAfter) ? pollMs : retryAfter * 1000);
        continue;
      }
      if (!response.ok) throw new Error(`${response.status}: ${await response.text()}`);

      const job = await response.json();
      if (job.status === "completed" || job.status === "review") return job;
      if (job.status === "failed" || job.status === "cancelled") {
        throw new Error(job.error || `extraction job ${job.status}`);
      }
      await sleep(pollMs + Math.random() * 1000);
    }
    throw new Error(`Job ${jobId} did not finish in time`);
  }
  ```
</CodeGroup>

## Related

<CardGroup cols={2}>
  <Card title="Rate Limits" icon="gauge" href="/docs/api-reference/limits/rate-limits">
    Per-plan limits, rate limit headers, and the 429 response
  </Card>

  <Card title="Checking API Health" icon="heart-pulse" href="/docs/api-reference/limits/checking-api-health">
    Confirm the API is up and your key works
  </Card>

  <Card title="Batch Extraction Cookbook" icon="layer-group" href="/docs/cookbooks/batch-extraction">
    Process many documents concurrently with a worker pool
  </Card>

  <Card title="Support" icon="envelope" href="mailto:support@unsiloed.ai">
    Contact the team if 503s persist
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.