
# Document AI

Extract structure and meaning from a document — OCR with word/line layout, form key→value pairs, tables, invoice and ID parsing, and natural-language questions about the content — from a single API call. Backed by **AWS Textract**. Six **synchronous** ops handle a single page; one **asynchronous** op handles a multi-page PDF and returns a job ticket you poll (see [Async jobs](/docs/async-jobs)).

```bash
curl -X POST https://api.relaystation.ai/v1/doc/ocr-layout \
  -H 'Authorization: Bearer rs_live_<key>' \
  -H 'Idempotency-Key: ocr-page1-20260707' \
  -H 'Content-Type: application/json' \
  -d '{ "file": { "inline": "<base64 image or 1-page PDF>" } }'
```

Or on the lodestone path — no account, a signed x402 payment instead of an API key:

```bash
curl -X POST https://api.relaystation.ai/v1/doc/ocr-layout \
  -H 'X-Payment: <base64 EIP-3009 authorization>' \
  -H 'Idempotency-Key: ocr-page1-20260707' \
  -H 'Content-Type: application/json' \
  -d '{ "file": { "inline": "<base64 image or 1-page PDF>" } }'
```

## Inputs — the cputools file convention

Every op takes its document as a **`file`** input source, the same shape used across the compute API:

```json
{ "inline": "<base64-encoded document>" }   // for files ≤ 4 MiB
{ "inputKey": "<scratch object key>" }        // for larger files, up to 50 MB
```

For a file above the inline ceiling, mint a one-time presigned upload with `POST /v1/cputools/upload-url`, `POST` the bytes, then pass `{"inputKey":"..."}` — the full recipe is in [Passing & receiving files](/docs/receiving-outputs).

## The synchronous ops (single page)

All six are `POST /v1/doc/<op>` and take a single-page document (an image or a 1-page PDF) up to **5 MB** (operator-tunable `docai.max_doc_bytes`). A multi-page or oversized input returns a `422` **before any charge**, naming the limit — reach for `analyze-async` for multi-page PDFs.

| Op | Route | What it returns |
|---|---|---|
| OCR with layout | `POST /v1/doc/ocr-layout` | `{ blocks: [...] }` — Textract line/word blocks with geometry |
| Extract form fields | `POST /v1/doc/extract-form` | `{ fields: [...], blocks: [...] }` — key→value pairs |
| Extract tables | `POST /v1/doc/extract-tables` | `{ tables: [...], blocks: [...] }` |
| Parse invoice / receipt | `POST /v1/doc/parse-invoice` | `{ documents: [...] }` — normalized expense fields |
| Parse government ID | `POST /v1/doc/parse-id` | `{ documents: [...] }` — ID field extraction |
| Ask the document | `POST /v1/doc/ask-document` | `{ answers: [...] }` — one answer per question |

`ask-document` also takes a `questions` array (1–15 items, each ≤ 200 chars):

```bash
curl -X POST https://api.relaystation.ai/v1/doc/ask-document \
  -H 'Authorization: Bearer rs_live_<key>' \
  -H 'Idempotency-Key: ask-20260707' \
  -H 'Content-Type: application/json' \
  -d '{ "file": { "inline": "<base64 1-page PDF>" }, "questions": ["What is the total?", "What is the due date?"] }'
```

Each synchronous op is a **flat per-document charge** — Textract's own per-page unit. The rates differ by op (a plain OCR costs less than a form or ID parse); see [Pricing](/pricing), or send an unauthenticated `POST` to read the exact price from the `402` challenge. An upstream Textract failure reverses the charge (you pay only for a delivered result).

## analyze-async — multi-page PDFs

`POST /v1/doc/analyze-async` runs Textract's asynchronous analysis over a **multi-page PDF** (PDF-only — the page count is what the ceiling is based on, and only PDFs page-count reliably). Instead of a result inline, it returns a **job ticket**; the compute finishes minutes later in a separate worker.

| Field | Type | Notes |
|---|---|---|
| `file` | input source, required | the multi-page PDF (`inline` ≤ 4 MiB, or `inputKey` up to 50 MB) |
| `features` | string[], required | one or more of `FORMS`, `TABLES`, `QUERIES` |
| `questions` | string[] | required when `features` includes `QUERIES` (1–15 items, each ≤ 200 chars) |
| `deliver` | `{ batonId }` | optionally deliver the result to a baton instead of holding it for polling |

```bash
curl -X POST https://api.relaystation.ai/v1/doc/analyze-async \
  -H 'Authorization: Bearer rs_live_<key>' \
  -H 'Idempotency-Key: async-report-20260707' \
  -H 'Content-Type: application/json' \
  -d '{ "file": { "inputKey": "<key from upload-url>" }, "features": ["FORMS", "TABLES"] }'
```

Response — the ticket:

```json
{ "ticket": "...", "job_id": "...", "status": "running", "ceiling_micros": 42000 }
```

Poll `GET /v1/jobs/{id}` until the status is terminal, or subscribe to the `job.completed` webhook. Page count is capped by `docai.async_max_pages` (default 3000); a document over the cap returns `422 DOC_TOO_LARGE` before any charge.

**Billing** is **per page**. The page count is counted at submit, so the ceiling is exact for Textract. On the x402 rail the charge settles at submit (pages × per-page); on the balance / API-key rail a hold is authorized at submit and the actual is captured at completion. A job that fails for an infrastructure reason is made whole. The billing mechanics are laid out in [Async jobs](/docs/async-jobs); see [Pricing](/pricing) for the per-page rate.

## MCP tools

Callable over MCP at `https://api.relaystation.ai/mcp`: `doc_ocr_layout`, `extract_form`, `extract_tables`, `parse_invoice`, `parse_id`, `ask_document`, and `doc_analyze_async` (with `job_status` to poll the ticket). Same auth, same prices as the HTTP routes.

## Next

[Async jobs](/docs/async-jobs) · [Passing & receiving files](/docs/receiving-outputs) · [ID verification](/docs/idverify) · [Quickstart](/docs/quickstart) · [x402 wire format](/docs/x402) · [API reference](/api-reference)
