Keyline API for Agents: The Runbook
This page is written to be read by an AI agent as well as a person. It is the whole integration on one page: how to authenticate, the three retrieval calls, what comes back, and the failure modes worth knowing before you hit them.
If you are a person deciding whether to bother, read Keyline for developers instead.
Status: the Public Preview opens soon. Everyone on the signup list gets a key, the command-line tool, and the request and response samples the day it opens. This page documents the contract itself, which is true now and will not change when the samples land.
What Keyline is, for an agent
Keyline is a retrieval tool. You call it with a natural-language question. It searches a document library that was indexed earlier, reads the pages worth reading, and returns an answer with the page and the cropped image region each claim came from.
Reach for it when:
- The answer is inside a figure, a diagram, or a table, where text extraction loses it.
- The question spans a library and you do not know which file holds the answer.
- The user will ask where the answer came from, and a page number alone is not enough.
Do not reach for it when:
- You already have the text and just need a model to reason over it.
- You need structured field extraction from one known document. That is a parsing job.
- The documents are mostly prose. Ordinary text search is fine and cheaper.
Core concepts, in dependency order
| Term | What it means |
|---|---|
| Organization | Billing and membership boundary |
| Project | The tenancy unit that owns documents. Every data call is scoped to one project. A project is the corpus |
| File | One uploaded document, addressed by a content hash |
| Snippet | One cropped semantic block of a page: text, figure, table, equation. The atom of a citation |
| Virtual page | A logical group of snippets that can span a physical page break, so a figure and the paragraph explaining it stay together |
| Bundle | Several files treated as one target for a deep read |
| Info | A saved investigation session you can resume later |
The two that matter for citations are snippet and virtual page. A snippet is what gets cropped and returned as an image. A virtual page is what the retrieval loop decides to read.
Authentication
One header. Server to server only.
- Credential: a project-scoped bearer token issued to a service account.
- Format: the prefix
plms_followed by 43 URL-safe characters. - The token carries no permissions of its own. Access is read from the live project grant on every request, so a revoked grant takes effect immediately.
- The token is shown once, at creation and at reset. Store it as a secret.
- Regenerating a token invalidates the old one immediately. There is no overlapping validity window, so plan the swap.
Two behaviors to code against:
- Bearer never falls back to a session. If a token is malformed, unknown, or revoked, the request fails. It does not degrade into an anonymous request.
- Every failure returns an identical unauthorized response. Invalid, revoked, never-issued, and not-permitted are deliberately indistinguishable. Do not write logic that tries to tell them apart, and do not retry hoping for a different error.
No CORS. Do not call this from a browser. There is no browser SDK and browser usage is out of scope by design.
Configuration your integration needs
Supply all three out of band. A token cannot discover them for you.
| Value | Where it comes from |
|---|---|
| Base URL | Provided with your Public Preview access |
| Project ID | Copied from the project settings page in the Keyline UI |
| Token | Issued with your Public Preview access |
The project ID is a real onboarding step. A token cannot enumerate the organizations or projects it belongs to, which is a deliberate isolation property.
The three retrieval calls
Pick by what you know before you ask.
| You know | Use | Shape |
|---|---|---|
| Nothing. You have a library and a question | Agentic search | Streamed. A reasoning loop that searches, resolves, reads, and decides when it has enough |
| Which file or bundle holds the answer | Deep read | Streamed. A fixed pipeline that reads inside that scope in depth |
| You only want to know which files mention something | Vector search | One response, no reasoning, cheapest by a wide margin |
Agentic search
The default choice for an agent. One request carries the question, optional prior conversation turns, and the last reference number you used so numbering continues across turns. The project is the corpus and there is no per-call file filter.
Internally it runs a tool loop: search the topic index, resolve promising hits into real content, report progress, and stop when the evidence is sufficient. You do not drive the loop. You consume its output.
Cost scales with library size and you cannot cap it from the request. There is no result-count or depth parameter. On a large library, prefer a project scoped to the documents that matter over one project holding everything.
Deep read
Use when your agent has already identified the file, or when the user named it. Pass the question plus the file, or a bundle plus the specific files within it. Follow-ups pass the root conversation identifier and the server loads the prior turns as context.
Two things to handle:
- The progress stream analyzes several units in parallel, so per-unit output interleaves. Buffer by unit identifier before displaying anything.
- Deep read does not save the conversation for you. Inference and persistence are separate calls today. If you want the exchange kept, accumulate the final text and save it in a second request.
Vector search
One call, one response, no model reasoning. It returns hits grouped by file, where each hit is a topic path through that document's hierarchy with a distance score. Good for "which of these files talk about this" and for narrowing before a deep read.
It does not return page numbers. Resolving a hit to actual page content is not available over a token in this preview, so do not design your own retrieval loop on top of vector search yet. Use agentic search for that and let the server run the loop.
Consuming a streamed answer
Both agentic search and deep read stream. Three kinds of event arrive:
| Event | What to do with it |
|---|---|
| Progress | Surface it. A long query should not be silent. This is the loop narrating its own steps |
| Working output | Intermediate analysis, useful for a live view. Not the answer. Buffer by unit identifier because units interleave |
| Final answer | The synthesized answer, streamed token by token, carrying inline numbered citations |
| Completion | Ends the stream and carries the structured reference list |
| Error | Ends the stream. Surface the failure, do not silently return a partial answer as if it were complete |
Disconnecting aborts the work server side, so abandoning a stream does not keep burning cost.
When nothing relevant is found, the answer says so. The pipeline does not fabricate an answer from an empty result. Treat a no-information response as a valid outcome and pass it through honestly rather than re-asking with a looser question.
Citations, and the one rule that matters
Answers carry numbered citations that resolve, through the completion event's reference list, to a file and a unit of content. From there:
- Page number is available for every citation.
- The cropped image of the cited block is retrievable as a PNG. This is the highest-value thing the API returns and the reason to prefer it over a text snippet.
- Region coordinates inside the page are not available over a token in this preview. They arrive with the Public Preview. Until then you can show the crop, but you cannot draw your own highlight box on the full page.
Never fabricate a citation
Only cite identifiers the server actually returned. Do not construct, guess, adjust, or interpolate a citation identifier, and do not carry one over from a previous answer. A citation that does not resolve is worse than no citation, because the whole value of this API is that a claim can be checked.
Rendering guidance that holds up in practice
- Download images to local paths rather than embedding remote URLs. You cannot control which renderer the user sees, and a local path renders everywhere. Make the download idempotent: if the file is already there, reuse it.
- Put the crop immediately next to the sentence that cites it. The crop is the argument. Placed at the bottom in a reference list, it is decoration.
- Show the retrieval loop's progress before the answer lands. Surface which pages are being read as soon as you know, so the user can follow along instead of waiting.
- Do not expose internal vocabulary to end users. Words like snippet and virtual page are for you. Say "the table on page 212" to the person reading the answer.
Ingesting documents
Upload returns immediately. Everything after it is asynchronous and your integration has to handle that.
- Request an upload. You get either a direct upload target to send the bytes to, or an instruction to fall back to a multipart upload.
- Send the bytes. If you were given a direct target with required headers, send exactly those headers. They are part of the signature and the upload fails without them. The target expires, so do not hold it.
- Poll the file's status until it is ready. There are no webhooks. Polling every few seconds is the expected pattern.
- Note that a document becomes deep-readable slightly before the library-wide search index finishes. Two distinct statuses both mean usable, and treating only the final one as ready leaves the document unavailable longer than it needs to be.
Practical numbers to plan against:
| Accepted formats | PDF, PNG, JPEG, DOCX, XLSX |
| Maximum file size | 260 MB |
| Maximum pages per file | 500 |
| Indexing speed | Roughly 30 seconds per page |
| Concurrency | Capped per project, around three files at a time, so one tenant cannot starve others |
| Region | US region available now |
A large first import is genuinely slow. Index a meaningful subset, verify it answers a real question, then widen.
What you are billed for
Worth knowing if you are planning a large ingest, because indexing is the cost that scales with your library rather than with your traffic.
| Unit | Rate |
|---|---|
| Indexing | $0.06 per page |
| Query | From $0.05 per question |
| Storage | $0.05 per GB per month |
Indexing is charged per page on first registration, and again if you update a document. It is a one-time cost per page: querying the same library repeatedly does not re-index anything. Storage is prorated daily and includes the index alongside the file.
Public Preview access comes with $50 in free credit, roughly 800 pages indexed with room to ask a few hundred questions.
Practical consequence for an agent: prefer indexing a scoped subset that answers a real question over indexing everything available. A fifty-thousand-page library is a meaningful first invoice, and you cannot un-index it for free.
What a token cannot do
The bearer token reaches an enumerated set of endpoints. Anything outside it is refused even when the account's role would otherwise allow it. This is transport hardening layered on top of normal permissions, and it is expected to relax as rate limiting and audit logging land.
In this preview a token cannot:
- Delete anything. Not files, not bundles, not sessions.
- Rename a file, or save a conversation attached to a file. Conversations attached to bundles and to saved sessions can be saved.
- Read text-extraction outputs, page dimensions, or region coordinates.
- Resolve a vector-search hit into page content.
- Manage members, or grant anyone any access.
- Discover which organization or project it belongs to.
Design around all of these rather than probing for them. Every one returns the same opaque unauthorized response, so probing teaches you nothing.
Failure modes, in the order you will meet them
- Everything returns unauthorized. Check the token prefix and length first, then the project ID, then whether the endpoint is one a token may reach at all. The response is identical for all three causes.
- A file never becomes ready. Confirm the format is one of PDF, PNG, JPEG, DOCX, or XLSX, and that it is under the size and page limits. A rejected format fails at upload, before any processing starts.
- Uploading the same document twice. Duplicate names in the processing queue are rejected, and a file already being processed is locked. Treat both as "already in flight" rather than as errors to retry.
- Interleaved progress output looks like nonsense. It is parallel analysis. Buffer by unit identifier.
- The conversation was not saved. Deep read does not persist. That is a second call.
- An answer says nothing relevant was found. Often correct, especially on architectural, electrical, and wiring drawings, which are a confirmed failure case. Do not paper over it with a vaguer question.
- Error text arrives in Japanese. Some server messages are not fully translated yet. Match on status codes.
Known gaps in this preview
Stated so you can plan rather than discover:
- No self-serve key issuance until the preview opens. Keys come from the Keyline team until then.
- The command-line tool ships with the Public Preview. Until then, drive the API directly over HTTP.
- No published machine-readable API schema.
- No webhooks.
- No rate limiting or quotas. Be a considerate caller.
- No region coordinates over a token. Crops yes, rectangles not yet.
Getting access
Sign up for the Public Preview from the developer page. Everyone on the list gets a key, the command-line tool, and the request and response samples the day it opens.