Bring it in.
PDF, Office, CSV, email, DXF drawings and scans, from a file share, S3 or a URL.
Drop in PDFs, scans and drawings. indx looks at every page, decides how to read it, and hands your search and agents paragraphs, tables and diagrams that keep their page and section. You never pick a parser.
An agent cannot answer from half a sentence. It needs the whole clause, the table row it points to, and the step that comes before an inspection.
Before anything expensive runs, indx scans each page and picks the reader it needs: the text layer when there is one, OCR when the page is a scan, a diagram parser when it recognizes a chart. Nobody configures this per file.
native-extractionScan, read by generic-ocrChart recognized process-chart-parserThree pages indx received. Each one shows how indx chose to read it. Move over any block on the page and the output is what your AI receives for it.
On the table, each cell is its own target. On the chart, try a symbol or a connecting line, then the lists in the output.
indx looks at every page before it reads it, then keeps the document's own shape in what it hands over.
PDF, Office, CSV, email, DXF drawings and scans, from a file share, S3 or a URL.
A cheap preflight picks the reader for each page: text layer, OCR or a drawing parser.
Lines are joined into paragraphs, cells into tables, shapes and connectors into diagrams.
Sections, defined terms, cross-references, entities and embeddings.
JSON with page, box, section and reader on every block, for search and agents.
Deployed into your AWS, Azure or Google Cloud account. No document leaves unless you switch on a hosted model.
Page, box, section and the reader that produced it, so an answer can always be checked on the page.
Text layers, OCR, drawings and embeddings run on CPU, with engineers alongside you from the first interview to rollout.
Bring a document. Hover the result.