Your documents

What was read, what was not, and what you can re-check.

The contracts, filings, policies and mail your firm accumulated over a decade are a set nobody currently has an inventory of. Citenda builds the record of it: what is in scope, what was read, what could not be read (named, not averaged away), and how to re-check every part of that on your own machines.

The record

Citenda leaves every file in exactly one state

In scope and converted; in scope and flagged unreadable, listed by name; or declared out of scope, like spreadsheets. The record states which, for every file, so a total is always a total of something you can enumerate. That containment is the product: not a promise that everything works, but an account of what does, what does not, and where the boundary sits — signed off against your own file listing, not ours.

Every file ends in exactly one state Every file in your listing is split into three states with no remainder: converted, split and indexed (in the record, citable); could not be read (flagged and listed by name); or declared out of scope, like spreadsheets. The three lists sum to your own file listing. Your documents Every file in your listing State 1 — in the record Converted, split, indexed citable and re-checkable State 2 — named Could not be read flagged and listed by name State 3 — declared out Out of scope spreadsheets — declared, never dropped
Three states, no remainder — the lists sum to your own file listing

Once it is built

One corpus, and every desk asks it a different question

The record is not a report that lands on one person's desk. It sits on your infrastructure and answers whoever asks, in whatever they already work in — and each question stops only at the documents it actually needs.

Three desks, three questions, one corpus Three people in different teams each ask their own question of the same corpus. The client team asks what fee was quoted, and its route stops at Finance's fee schedule and Sales' client contract. Risk asks which funds changed fee since March, stopping at the fee schedule and Compliance's sign-off. Product asks whether factsheets and contracts agree, stopping at Marketing's factsheet and the client contract. Each answer returns to the person who asked, in the tool they already use: a chat, a Word document, a slide deck. Every answer quotes the documents it read. Anyone, any team Your documents Back where they work One corpus · on your machines Finance Fee schedule Marketing Factsheet Sales Client contract Compliance Sign-off Every team's documents, converted, split and indexed into one set. Each question stops only where it needs to, and quotes what it read. Client team What fee did we quote this client? Risk Which funds changed fee since March? Product Do our factsheets and contracts agree? In a chat Answered, with the two documents it quoted In a .docx Drafted in Word, every figure traced to a source In a .pptx Slides for the committee, citations in the notes
Three desks, three questions, one corpus — each route stops where it needs to, and comes home

How it is read

We build your corpus. You choose how privately it is read.

One corpus, two read paths Documents arrive through one of two doors — an edition we publish, or your own documents. Either way, what we deliver is the corpus our engine builds from them: the documents converted, split and indexed, built to be read and cited by a model, and handed over as plain files with stable citation keys, already on your machines. From there it forks into two read paths. Path 1, Sovereign mode, is an open-source model self-hosted on hardware you control: nothing leaves your network, and it runs with no connection at all. Path 2 is a hosted model — Claude Code, which reads from disk and sends what it reads; Cowork, which uploads the files you attach to the workspace; or any other stack. On those, what leaves your network is your question and the passages the tool reads. Citenda receives nothing on either path. Door A An edition we publish or Door B Your own documents either way What we deliver The corpus The documents, converted, split and indexed by our engine — built to be read and cited by a model. Plain files. Stable citation keys. Already on your machines. Path 1 · Sovereign mode An open-source model self-hosted, on hardware you control Leaves your network: nothing. Runs with no connection at all. Path 2 · you use a hosted model Claude Code best fit — reads from disk sends what it reads Cowork files you attach are uploaded to the workspace Any other stack open weights, in-house RAG, your own agent Leaves your network: your question, and the passages the tool reads.
One corpus, two read paths — the privacy decision is the reader you pick

Citenda receives nothing on either path — once the corpus is delivered we are not in the read path at all. Both doors land in the same box, and that is the load-bearing part: an edition we publish and your own documents come out as the same kind of artifact — a corpus our engine built, not a folder of documents handed back. So the read fork is identical for both, and the privacy decision is never a function of what you bought. Teal marks the recommended path, not the safe one — those are separate axes, and Path 1 is the one where nothing leaves.

The answer

Every question ends in one of three declared states

An answer, with the documents behind it. An evidenced stop that names what was covered and what could not be read. Or — if the underlying AI platform's capacity limits interrupt the session — a clear "not completed" with nothing lost, resumed when capacity returns. Never a silent failure, and never a confident guess.

We attach no number to that bound — no answer-time figure and no accuracy percentage for your archive — until a measurement of your archive produces one. That is what the Assessment is for.

Reproducibility

The record can be re-derived, not just re-read

What we hand over carries its own manifest: what went in, what came out, and the checksums that tie the two together. The included verify.py re-checks the delivered artifact on your machines, offline — so months later, when someone asks where an answer came from, the citation still points into a set you can enumerate and a record you can re-derive. If your auditor asks how you know the archive copy is intact, the answer is a command, not a meeting.

Where the build runs

Managed delivery leaves you one job — the one nobody can do for you

The Estate Assessment is always one script run in your environment. What can differ is where the corpus itself is built afterwards: on your infrastructure, or on ours. If we host it, there is no machine to provision and no on-site setup — you prepare your documents and transfer them once, and it is the faster of the two to get running.

What each way of running the build asks of you
On your infrastructure Managed — we host the build
A dedicated machine, provisioned by you —
Citenda on site to clone and run the engine —
Your documents, prepared Your documents, prepared
— One transfer, SFTP or equivalent

Managed delivery is priced as an addition to whichever floor applies, set in the offer alongside the estate measurement. There is no separate rate card for it — like everything after the Assessment, the amount is fixed once your archive has been measured.

Pricing

What it costs, starting from a fixed price

Every engagement starts with the Estate Assessment — a fixed price, credited in full against what comes next. After that, Foundation, Programme and Enterprise are priced as floors, fixed only once the Assessment has measured your archive.

See the pricing floors →

Citenda states its limits before you have to ask

  • We never repair a damaged identifier. If an ISIN, IBAN or VAT number fails its own check digit, we flag it and point you at the source page.
  • We name what we could not read. Unreadable or poorly-converting documents are flagged and listed by name — never dropped quietly out of a total.
  • We do not cover spreadsheets. We convert PDF, HTML/XML, Word, PowerPoint, email and images, and we would rather say so than quietly index tabular data badly.
  • You keep your documents. On your own infrastructure, conversion, splitting and indexing run on your machines and make no outbound call; the answering step runs under your own AI account and key, in your own logs. If you choose managed delivery, you transfer them to us once and we build there — that is the one case, and we name it.
  • We stop rather than guess. When the set does not contain the answer, Citenda says so and shows what it looked at.
  • We do not claim accuracy on your archive until we have measured it. Every published figure comes from a corpus we built and a benchmark we ran ourselves, and we label it that way.

The first step — paid, bounded

Start with the Estate Assessment

The pack is a set we maintain and you can check. This is your set — which nobody currently has an inventory of.

The Estate Assessment is one script run in your environment — nothing leaves the building — producing one measured profile of the contracts, filings, policies and mail your firm holds, its document estate: how much converts cleanly, what will not read and why, where the borderline cases sit, and what model access your team already has to work against. It is the first evidenced stop you buy, and it is the measurement every later fee is fixed against: a fixed fee, quoted after a measured profile of your archive.

Request an Estate Assessment

Or email us directly: boris.marchand@citenda.com

The evidence

The benchmark we ran ourselves, published with its limits

We do not claim accuracy on your archive until we have measured it — so the figures we do publish come from a controlled benchmark on a corpus we built, and they are labelled that way, limits first.

Read the benchmark: method, figures and limits →

Metadata enrichment is in development and will be priced when we can show you what it adds.