Back to home

Resources

How Claix answers questions about your documents, without a vector database

A question we get asked a lot, usually by someone who's already built a RAG pipeline once and is tired of debugging it: "so where's the vector store?"

There isn't one. Claix doesn't chunk your documents, doesn't embed them, and doesn't do a nearest-neighbor search before answering a question. What it does instead is simpler to describe than it sounds, and the three endpoints involved map pretty directly onto three different things you'd want to do with a document once it's been processed.

After a document is stored

The three things you can do with a document after it's stored

Once a document is extracted (through any of the six format endpoints — PDF, Excel, Word, image, plain text, or audio) with its context window enabled, it gets stored as structured, persisted content tied to a document_id. From there, you have three options, and they're genuinely different operations, not three names for the same thing.

  • GET /get-document/{document_id}

    Just gives you back the content. No AI involved at this step — it's a retrieval of what's already there, in markdown. If you already know what you want (you're building your own pipeline downstream, or you just want the clean text without paying for interpretation again), this is the cheapest and fastest of the three, because nothing gets re-reasoned.

  • POST /document-context/{document_id}

    This is where you actually ask something. You send up to five typed questions — each one tagged with the format you expect back (string, int, boolean, timestamp, array) — and you get typed answers, 1:1, in the same order. Not a paragraph you have to parse yourself. If you ask for a boolean, you get true or false, not "yes, it appears that the contract does specify this."

  • POST /space-context/{space_id}

    The same request and response shape — this is intentional, so your code doesn't change when you switch from asking about one document to asking about a group of them — but it answers from every document inside a knowledge space at once. This is the one that actually does something a single-document lookup can't: it crosses files. "Which supplier bills the most across all invoices" only makes sense once there's more than one invoice to sum.

Grouping

How grouping actually happens

Spaces aren't a separate upload step. Any of the extraction endpoints accepts an optional space_id in the request. Pass the same space_id across as many extractions as you want, and each one joins that group. There's no "add document to space" call to remember separately — the grouping happens at the moment the document is processed, which is a small detail but it means you don't end up with documents that got extracted and then forgotten before anyone grouped them.

Scale

What happens when a space gets big

This is the part that's easy to gloss over and actually matters most. A knowledge space can grow without any limit on your side — but a call to space-context can't just hand all of it to a model, because that stops being accurate (and stops being affordable) somewhere well before "unlimited."

So there's a cap: the 50 most recently added documents, and a shared budget of 200,000 characters across whichever of those are selected. If a document gets trimmed to fit, the model is told explicitly that it didn't see the full thing — it doesn't quietly answer as if it had complete information when it didn't. Documents that have expired (if you've set a retention window) don't count toward the space at all, so an old invoice that should no longer be considered doesn't quietly skew a total.

None of this is retrieval by similarity. It's closer to "here's the data that's actually in scope, bounded by recency and size, and here's what got left out" — a budget problem, not a search problem.

Source Tracing

Where the verification comes in

If source verification is turned on for a space, the answer you get back isn't just the value — it's { value, source }. For a cross-document answer, the source field names which documents it actually pulled from, by document_id and filename, and says explicitly when more than one was crossed to produce the number. If there's nothing in the documents to support an answer, source comes back as requires_human_revision instead of a guess dressed up as a fact.

This is the same mechanism whether you're asking about one document or fifty of them — document-context and space-context share the response shape on purpose, so this isn't a special "space mode" feature, it's just what both endpoints do.

Why this shape

Why not embeddings, then

Short answer: because the questions people actually ask across a space tend to be things like "does this invoice match its contract" or "sum the totals across these three files" — and those aren't similarity questions. Nothing about cosine distance tells you whether 48,320€ is the right sum of three line items. A full-text match and a model that can read the relevant documents in full gets you there more reliably, and the "what did it actually look at" answer is a filename and a document ID, not a guess about which chunk the retriever happened to pick.

That said, this isn't a claim that nobody should ever use a vector database. If you're searching across a genuinely large, mostly static library where the question is "find me things conceptually related to X" rather than "reconcile these specific numbers," embeddings are solving a real problem that exact matching doesn't. Claix is built for the second kind of question, not as a replacement for the first.

Example

A small example

Two invoices and a contract, grouped into the same space at extraction time. One call, and the response that comes back. That's the whole mechanism. Three endpoints, one shared answer shape, and a budget that tells you honestly what it did and didn't look at.

Request
{
  "questions": [
    { "question": "Which supplier bills the most across all invoices?", "format": "string" },
    { "question": "Is there any contract whose amount does not match its invoice?", "format": "boolean" }
  ]
}

Response
{
  "ia_response": [
    { "value": "Suministros Omega S.A., 48,320€ across three invoices.", "source": "document_id a1b2... (invoice-feb.pdf) crossed with document_id b2c3... (invoice-mar.pdf): sum of amounts" },
    false
  ],
  "log_id": "7c2e1a90-4b3d-4f8a-9e21-6d5c8b0a1f34"
}

FAQ

Does Claix use a vector database?
No. Document and space queries use structured extraction plus full-text matching, not embeddings or nearest-neighbor search.
What's the difference between get-document and document-context?
get-document returns the stored content as-is, with no reasoning involved. document-context asks typed questions about that content and returns typed answers.
How many documents can a space-context call consider?
Up to the 50 most recently added live documents in the space, within a shared 200,000-character budget. Expired documents are excluded automatically.
What happens if a question can't be answered from the documents?
With source verification enabled, the answer comes back as requires_human_revision instead of a fabricated value.
Can I ask more than 5 questions in one call?
Not in a single request — split them across multiple calls to document-context or space-context.

Three endpoints, one answer shape, and an honest budget.

Ask a stored document, or every live document in a space, and get typed answers back. The call tells you what it looked at, and what it left out.