Skip to content
Search and AI

An AI that reads the whole document. And you configure it.

Theka's artificial intelligence reads text, scans, tables, charts and images, fills in fields and makes every document searchable by meaning. For each category you decide what to do, with which model and how deeply.

AI configuration · per category
AI tasks enabled per category
CategoryExtractionFieldsSummaryOCRSemanticKnowledge
Technical data sheets✓✓✓✓✓✓
Signed contracts✓✓–✓✓–
Product photos✓–––✓–
Internal documents✓✓✓✓✓✓
What it reads

Not just the text.

A lot of business information sits inside tables, charts, photos and scans. Theka reads all of it, so search and knowledge don't stop at the text.

Text

PDF, Word, Excel, PowerPoint and plain text, with the structure of pages and sections.

Scans

OCR reads scanned pages and photos of paper documents, only where needed.

Tables

Rows and columns become data: the numbers inside a table can be found and cited.

Images and charts

Photos are described, charts translated into their data, repeated images (logos, letterheads) ignored.

What it produces

Documents already in order when they arrive.

Automatically on upload, or on request for a single document or a selection.

Summary and tags

A summary of the document and the labels to find it again.

Metadata and fields

Dates, amounts, codes, parties; the category's custom fields filled in wherever there is evidence.

Classification

If the category or type is missing at upload, the AI suggests them by reading the document.

Search by meaning

Every document becomes searchable even with words different from the ones it contains.

Depth

Three reading levels, per category.

Not every document deserves the same attention: choose the depth for the whole archive and change it where needed.

01

Basic

Text and structure, without OCR or image analysis. The fastest and the cheapest.

02

Standard

OCR where needed, embedded images and photos described. The default level.

03

Full

All images, charts and visual analysis of every page. For documents where the graphics matter.

Search

Seven signals, one list of results.

Every search combines words, meaning, customers, projects, topics and knowledge facts, and ranks documents by relevance. Every result carries up to three passages with their page.

  • Finds documents even when you don't use the same words they do
  • Tells apart the version where it found the text from the current one
  • Always filters by permissions: nobody finds what they could not open
01Document metadata
02Document text
03Meaning
04Customers and suppliers
05Projects
06Knowledge topics
07Knowledge facts
Models

The right model for every task.

AI providers are connectors: add the key, choose the model for each task and change it whenever you like, without touching documents or workflows.

Per task and per category

An inexpensive model for invoice tags, a more capable one for contracts, another for search by meaning. Category settings take precedence over the general ones.

OpenAIAnthropicGoogle GeminiMistral

Instructions per document type

A few lines of “AI context” per document type tell the AI what matters: they guide summaries, fields, image descriptions and knowledge.

Costs

Know what you spend, before you spend it.

AI consumes tokens from your provider: Theka counts them and gives you the tools to decide.

Estimate before you start

For bulk processing you see documents, timings and tokens before confirming.

Daily caps

Limits for knowledge and for images and tables; everything else keeps working.

Budget exhausted

If the provider says stop, the work is postponed and administrators are notified.

Models being retired

When a model you use is about to be retired, Theka flags it and suggests its successor.

Frequently asked questions

About artificial intelligence.

Which AI models can I use?

OpenAI, Anthropic, Google Gemini and Mistral, with your own key. You can use different models for different tasks and for different categories.

Are documents sent to the AI provider?

Only for the tasks you enable, and only for the categories where you enable them. With no AI tasks enabled, no content leaves for an external provider.

How do I keep spending under control?

You choose the reading level, the model per category and the daily caps; before bulk processing you see the estimate. If the provider's budget runs out, the work is postponed and administrators are notified.

Can I redo just part of the processing?

Yes: on a document you can regenerate a single task (for example the tags) or a single stage, without rereading everything.

// contact

Let's try Theka
on your archive.

A 30-minute call for a guided demo and an initial assessment of how well Theka fits the way you work.