Text
PDF, Word, Excel, PowerPoint and plain text, with the structure of pages and sections.
Theka's artificial intelligence reads text, scans, tables, charts and images, fills in fields and makes every document searchable by meaning. For each category you decide what to do, with which model and how deeply.
| Category | Extraction | Fields | Summary | OCR | Semantic | Knowledge |
|---|---|---|---|---|---|---|
| Technical data sheets | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Signed contracts | ✓ | ✓ | – | ✓ | ✓ | – |
| Product photos | ✓ | – | – | – | ✓ | – |
| Internal documents | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
A lot of business information sits inside tables, charts, photos and scans. Theka reads all of it, so search and knowledge don't stop at the text.
PDF, Word, Excel, PowerPoint and plain text, with the structure of pages and sections.
OCR reads scanned pages and photos of paper documents, only where needed.
Rows and columns become data: the numbers inside a table can be found and cited.
Photos are described, charts translated into their data, repeated images (logos, letterheads) ignored.
Automatically on upload, or on request for a single document or a selection.
A summary of the document and the labels to find it again.
Dates, amounts, codes, parties; the category's custom fields filled in wherever there is evidence.
If the category or type is missing at upload, the AI suggests them by reading the document.
Every document becomes searchable even with words different from the ones it contains.
Not every document deserves the same attention: choose the depth for the whole archive and change it where needed.
Text and structure, without OCR or image analysis. The fastest and the cheapest.
OCR where needed, embedded images and photos described. The default level.
All images, charts and visual analysis of every page. For documents where the graphics matter.
Every search combines words, meaning, customers, projects, topics and knowledge facts, and ranks documents by relevance. Every result carries up to three passages with their page.
AI providers are connectors: add the key, choose the model for each task and change it whenever you like, without touching documents or workflows.
An inexpensive model for invoice tags, a more capable one for contracts, another for search by meaning. Category settings take precedence over the general ones.
A few lines of “AI context” per document type tell the AI what matters: they guide summaries, fields, image descriptions and knowledge.
AI consumes tokens from your provider: Theka counts them and gives you the tools to decide.
For bulk processing you see documents, timings and tokens before confirming.
Limits for knowledge and for images and tables; everything else keeps working.
If the provider says stop, the work is postponed and administrators are notified.
When a model you use is about to be retired, Theka flags it and suggests its successor.
OpenAI, Anthropic, Google Gemini and Mistral, with your own key. You can use different models for different tasks and for different categories.
Only for the tasks you enable, and only for the categories where you enable them. With no AI tasks enabled, no content leaves for an external provider.
You choose the reading level, the model per category and the daily caps; before bulk processing you see the estimate. If the provider's budget runs out, the work is postponed and administrators are notified.
Yes: on a document you can regenerate a single task (for example the tags) or a single stage, without rereading everything.
A 30-minute call for a guided demo and an initial assessment of how well Theka fits the way you work.