How it works

A document-review example, followed by recordings of three actual Studio tests. The documents are fictional; the responses and citation error are real.

Worked example · Document review

Start with the source. Keep the decision with a person.

Imagine reviewing a draft procedure. You need to find what it says, check whether the evidence supports it, and record what still needs an answer.

01

Check the extracted document

Read text directly where possible. Flag scans or diagrams that need different handling, and keep unreadable pages visible.

02

Compare evidence with agreed questions

Use the review criteria you supplied. Each finding should point to the passage that supports it.

03

Get a separate reading

A second pass reads the evidence before the draft. Code checks page references and quoted passages; disagreements remain visible.

04

Let the reviewer decide

A person accepts, rejects or investigates each finding. A citation check cannot establish whether the conclusion is sound.

This explains a design approach. The checks and acceptance criteria for a build depend on your documents and the decisions you need to make.

How I would compare models+

Agree the expected answers with a qualified reviewer before running the test. Include missing evidence, incorrect findings and cases where the model should leave the question unresolved.

Measure missed problems, false alarms, human review time and API cost. A cheaper model is useful only if the complete review still meets the agreed standard.

Choosing how to search documents+

Exact identifiers and conceptual questions need different tests. I would compare lexical, semantic and combined search against queries representative of your work, then inspect the retrieved passages.

Recorded in CPLT Studio · 12 September 2026

Three questions. Real answers. Sources checked.

These clips show actual responses from the Document Desk — Luna agent using ten fictional documents. Questions and retrieved excerpts were processed by OpenAI. This is a recorded demonstration, not a live chat or evidence from a customer.

1. Normal support hours

Question: According to the Support guide, what are the normal support hours? Show the source.

The answer gives Monday–Thursday, 09:00–16:00 UTC and preserves the closure-date exception. The cited Support guide supports that answer.

Silent screen recording. Fictional documents. Actual response shown; clip trimmed after the result.

2. A closure date — and a citation error

Question: Will support be open on Monday 14 September 2026 at 10:00 UTC? Check both the Support guide and Team calendar.

The conclusion is correct: support is closed on 14 September 2026. The two citation links are attached to the opposite source statements. The error is retained in this clip; a correct conclusion does not make this a clean pass.

Silent screen recording. Fictional documents. Typing and setup omitted; the answer and citation check are retained.

3. A phone number that is not in the documents

Question: What telephone number should I call for support?

The answer says no telephone number is provided in the supplied documents. It does not invent one. This is one observed result, not a guarantee that missing information will always be recognised.

Silent screen recording. Fictional documents. Actual response shown; clip trimmed after the result.

I would assess a larger, representative set of questions before recommending a build. These clips show what happened in three tests, including a source-link mistake; they do not establish a reliability rate, time saving or production-readiness claim.

Choosing the model

Nobody on your team should have to pick a model

Most tools hand people a dropdown of model names and let them guess. The names change every few months, and the guess is wrong in one of two directions — a one-line question sent to the most expensive model on the list, or a genuinely hard one answered in a hurry by the cheapest. Here the platform sizes the question first and picks for them.

01

It reads the question before it answers it

A quick assessment of what the ask actually needs. Its cost and response time depend on the routing implementation and should be measured.

02

Ordinary questions go to a small, fast model

Reformat this, summarise that, what does this term mean. Most of the day's traffic never needs more.

03

Analysis moves up a tier

Longer reasoning, comparisons across documents, anything where a fast answer would be a shallow one.

04

Hard problems buy thinking time before they buy a bigger invoice

The first escalation is letting a model think for longer, not reaching for a costlier one. Price is the last lever, not the first.

05

A conversation keeps the model it started with

The choice is made once, at the top of a thread, so an answer's character doesn't shift halfway through a discussion.

Open models and frontier models sit in the same pool. Which of them a question is allowed to reach is your decision, not the router's — you enable the endpoints, using your own provider keys. Work you have marked sensitive is answered by models running on hardware you control, and those bytes never leave the building. Anything routed outside goes to a provider you approved, sending the minimum needed, and every call is logged where you can read it.

None of this is a black box we invented and licensed to you. The routing policy is a piece of configuration on your own system: you can read it, change the boundary, or pin a team to local models entirely. The routing behaviour and data boundaries are agreed and tested for each build.

This is a demonstration of the platform. CPLT — the association behind it, and the builds it takes on when a team asks — lives on cplt.tech, along with what a build costs. It starts with a free forty-five-minute scoping call that ends in a one-page written note on what your team can run.

Book a scoping call ↗ Back to cplt.tech ↗