Exhibit 001 · The Corpus · 1998 to today

Your machine holds everything you ever wrote.
Now it can answer.

docenta is a local-first content retrieval engine built for LLM agents. It turns your files, mail, and chat history into one askable collection. Answers come with receipts: exact file, exact span. Nothing leaves your machine.

agent session · docenta over MCP
when did we decide to switch the importer to streaming, and why?
docenta · semantic_search → timeline channel · 6 ms
1. docs/design/importer-notes.md · "2024-03-08: streaming import decided, batch peak RAM 14 GB was the trigger"
2. Mail · "Re: importer memory ceiling" · 2024-03-07 · the thread that forced the call
3. agent session · 2024-03-08 · the refactor itself, step by step
every claim cited · token budget enforced · gaps disclosed, never papered over

Real behavior, not a mockup: document search, mail, and past agent sessions are one corpus. The docent answers from all of it.

Exhibit 002 · Why nothing else does this

File search finds names. grep reads bytes. Cloud RAG reads your life on someone else's disk.

docenta answers "which of my files says this" and proves it. A docent is the museum guide you ask about a collection. Documentum is Latin for "a thing that teaches". Your corpus gets a docent.

Local-first. Actually.

Ingest, index, embeddings, answers: all on your hardware. Zero uploads, zero telemetry, zero cloud dependency. Originals are never copied; only searchable derivatives are stored.

Agent-native. Not a search box.

The primary user is your AI agent. Token-disciplined content_search, entity_card, timeline, and evidence_pack tools over MCP; your agent already speaks the protocol.

Receipts. Every time.

Answers cite exact spans in exact files. Token budgets are enforced and disclosed. When the corpus cannot answer, docenta says so instead of improvising.

Collection statistics · one real machine
815,405content objects from 1.7M file instances
246,000searchable messages out of one 32 GB mailbox
4-6 msentity and timeline answers
7.5 minfull ingest of a 463k-file home directory
0 bytessent to any cloud, ever
Measured on the maker's own corpus, Apple M4 Max. Every number traces to a committed field run.
Exhibit 003 · How the docent works

A spine that ingests, three channels that answer, one artifact that proves it.

ingest walk → hash → extract → index · containers explode into members · re-scan costs O(changes)
channels exact BM25 · paragraph lexical · dense vectors, fused and re-ranked · entities, timelines, themes on top
answer evidence_pack: route plan, cited spans, conflicts, disclosed gaps, enforced token budget
live mail, chat exports, and agent sessions stream in continuously · your corpus feeds itself

In a 48-task agent harness, the evidence pack answered 44 of 48 questions against 39 for iterative search, in 1 to 2.5 tool calls per task. Deterministic layers, zero learned models in the loop, bit-stable re-runs.

Exhibit 004 · Who asks

One engine, your kind of collection.

The same docent, pointed at what your work actually holds. Each page shows the shape of an answer for one kind of desk.

Exhibit 005 · The velvet rope

Early access is small on purpose.

docenta is a commercial product of SKY, LLC, now in private preview. Tell us what your agent needs to remember and we will tell you when the door opens.

request early access or check back here: the public beta will be announced on docenta.ai