Agent mode API

Agent mode endpoints

Structured extraction plus semantic reasoning with agent_definition parameters.

PDF-to-JSON · Agent mode

PDF to JSON with Agent mode

Endpoint

POSThttps://claix.dev/agent/pdf-json

Agent mode active

This endpoint runs structured extraction from the main schema first, then a reasoning phase with agent_definition. The response includes data[] (extraction), agent_data (typed inference), and log_id. The schema must have is_agent_mode enabled.

This endpoint accepts a PDF file and returns a single JSON object with the data extracted from the document, following exactly the structure you define via a schema. It works with selectable-text PDFs and scanned PDFs, because analysis is performed by an AI model with native document understanding, not a plain-text extractor.

It is intended for server-to-serverintegrations (backends, scripts, n8n/Zapier/Make). It must not be called from an end user's browser because it requires a secret API key.

Unlike Excel/CSV, the full PDF is treated as a single data source and the response contains exactly one record in data — ideal for invoices, contracts, forms, certificates, or reports.

1. Authentication

Every request must include your API key. It is a personal server credential and should be handled with the same care as a database password.

Option A — Dedicated header (recommended):

x-api-key: <YOUR_API_KEY>

Option B — Standard Authorization header:

Authorization: Bearer <YOUR_API_KEY>

Either one is sufficient. If you send both, x-api-key takes priority.

Before processing the PDF, the system validates that:

  • The API key exists and is active.
  • The associated account is active (not suspended).

If validation fails, the request is rejected with 401 without processing the file.

2. Request format

Method: POST · Content-Type: multipart/form-data (required)

FieldTypeRequiredDescription
fileBinary fileYesThe PDF to analyze. Must be the file itself, not a path or URL.
schema_idText (UUID)YesSchema previously created in your account, of type PDF → JSON.
space_idText (UUID)NoOptional. Knowledge space the stored document is attached to. Must belong to the same account as the API key. It only takes effect when the schema has the context window enabled, which is when the document is stored. You can then query the whole space with POST /space-context/{space_id}.

Field names must be exactly file and schema_id. Aliases such as pdf, documento, or upload are not supported.

The schema_id must correspond to a schema of type PDF → JSON. If you send one of Excel → JSON or another type, you will receive 400.

File requirements:

  • Format: .pdf only (validated by MIME type and extension).
  • Cannot be empty (0 bytes).
  • Maximum size: 15 MB (413 error if exceeded).
  • Compatible with text and scanned PDFs.

3. How to build the request

  1. Have your API key and the correct schema_id ready.
  2. Build a POST request to the endpoint URL.
  3. Add the authentication header.
  4. Send multipart/form-data with file (PDF) and schema_id.
  5. Check the HTTP status code: only 200 indicates success.

4. Request examples

See the panel on the right for examples in cURL, JavaScript, Node.js, Python, PHP, and n8n.

5. Successful response format

200 OK · Content-Type: application/json

{
  "success": true,
  "schema_utilizado": "Contratos",
  "total_registros": 1,
  "data": [
    {
      "puesto a ocupar": "Ing. Software Principal (Backend)",
      "Nombre contratante": "Tech Solutions S.L.",
      "fecha del contrato": "8 de Agosto de 2026",
      "persona contratada": "Gael Anaya"
    }
  ],
  "agent_data": {
    "salario": 55000,
    "es_parcial": false,
    "fecha_contrato": "después del 20/07/2026",
    "resumen_contrato": "Contrato indefinido: Gael Anaya como Backend en Tech Solutions S.L. Jornada completa, 55.000 €/año."
  },
  "log_id": "7c2e1a90-4b3d-4f8a-9e21-6d5c8b0a1f34"
}
FieldTypeDescription
successbooleanAlways true when HTTP is 200.
schema_utilizadostringName of the schema used for extraction.
total_registrosnumberNumber of records in data.
dataarrayObjects extracted from the main schema (same as extraction mode). If source verification is enabled on the schema, each property is { value, source }.
agent_dataobjectTyped Agent Mode answers per agent_definition (booleans, numbers, strings). If source verification is enabled on the schema, each field is { value, source }.
log_idstring (UUID)UUID of this call’s usage_logs row. Present on success and on most authenticated errors.

Every response includes log_id (the UUID of the usage_logs row) when the log could be stored. It also appears on most errors after the request is authenticated. Use it to find the call in the logs panel.

If source verification is enabled on the schema, each extracted property (and each agent_data field in Agent mode) becomes { "value": ..., "source": "..." } instead of a bare value. source is required: it cites the evidence (page, paragraph, cell, quoted snippet, image region, or the second / second range in audio). If there is no evidence, source is exactly requires_human_revision. If source verification is off, the format is unchanged.

Example with source verification enabled:

{
  "success": true,
  "schema_utilizado": "Contratos",
  "total_registros": 1,
  "log_id": "7c2e1a90-4b3d-4f8a-9e21-6d5c8b0a1f34",
  "data": [
    {
      "persona contratada": {
        "value": "Gael Anaya",
        "source": "página 1, párrafo 1"
      }
    }
  ],
  "agent_data": {
    "salario": {
      "value": 55000,
      "source": "página 2, cláusula retributiva"
    },
    "es_parcial": {
      "value": false,
      "source": "requires_human_revision"
    }
  }
}

6. Error codes

{
  "error": "Descripción legible del problema.",
  "detalle": "Información técnica adicional (solo presente en algunos casos).",
  "log_id": "7c2e1a90-4b3d-4f8a-9e21-6d5c8b0a1f34"
}

400 — Missing file or schema_id, invalid multipart, not a PDF, empty file, corrupt file, or schema of the wrong type.

401 — Authentication failed.

404 — schema_id does not exist or does not belong to your account.

413 — PDF exceeds 15 MB.

422 — PDF read but no extractable data according to the schema.

502 — AI service failure (transient).

405 — Method other than POST. · 500 — Internal error.

7. Code summary

CodeCategoryRetry?
200Success—
400Client error (malformed file or data)No — fix the request first
401Authentication errorNo — fix credentials first
404Resource not foundNo — fix schema_id first
405Incorrect HTTP methodNo — fix the method first
413PDF too largeNo — reduce file size first
422No extractable dataNo — review PDF/schema first
500Internal server errorYes, with caution
502AI service failureYes, recommended with backoff

8. Best practices

  • Validate the HTTP status code before reading data[0].
  • Check the PDF size on your client before sending it.
  • Remember: data always has exactly one element (one document = one object).
  • Retry automatically only on 500 and 502, never on 400, 401, 404, 413, or 422.
  • Clear descriptions on schema properties improve accuracy on documents with non-standard layouts.
  • Do not include your API key in frontend code or public repositories.

Request examples

curl -X POST "https://claix.dev/agent/pdf-json" \
  -H "x-api-key: <TU_API_KEY>" \
  -F "file=@./factura_marzo.pdf" \
  -F "schema_id=3c7a9f21-4b8e-4d1a-9c6f-2e0d8a5b7c4f"