Initial ingestion

€ —

One-time cost

Hybrid search

€ —

Per month, including cloud rental

Search with assistant

€ —

Per month, including hybrid search and cloud rental

Total first-year cost

€ —

First 12 months · search only

Phase 0 · Cloud rental

Monthly total€ —

Server rental for the self-hosted hybrid search service, including vector and keyword index storage and backups.

Sizing hint from Phase 1 inputs

Enter the document count, pages and an embedding model in Phase 1 to estimate the index size.

Phase 1 · Initial ingestion

One-time total€ —

One-time total = documents × pages/document × tokens/page ÷ (1 − overlap/100) × embedding rate ÷ 1,000,000 + additional compute

Used only for the index sizing hint.

Advanced ingestion assumptions

Affects the index size hint, not the embedding cost.

Phase 2 · Hybrid search

Monthly usage€ —

Monthly usage = searches/month × (query embedding cost/search + reranking cost/search) + document update embedding cost/month

Monthly search usage breakdown

Query embeddings€ —
New or updated documents€ —
Reranking€ —

Searches made directly, excluding assistant conversations, which are counted in Phase 3.

Each one is fully re-embedded with the Phase 1 assumptions.

Estimated searches per month: —

Phase 3 · Search with assistant

Monthly usage€ —

Monthly usage = Phase 2 usage + assistant turns/month × [retrieval cost/turn + (input tokens/turn × blended input rate + output tokens/turn × output rate) ÷ 1,000,000]

Every turn runs one hybrid retrieval plus one generation.

Estimated assistant turns per month: —

Complete model input: system prompt, retrieved passages and conversation history.

Include billable reasoning tokens.

0 % is the conservative default. A stable system prompt and long conversations can reach 50 % or more; cache write premiums are ignored.

Model pricing sources · checked
Model Input / request rate Cached input Output Unit