Why Search-as-a-Service Is Faster Than In-House Solutions
We often see companies spending months building AI search from scratch for each product: writing their own vectorization, setting up reranking, configuring infrastructure. Search as a Service is a ready-made semantic search platform with an AI layer that we connect via a unified search API and search SDK. Our engineers with 10+ years of MLOps experience handle all infrastructure: from choosing an embedding model to configuring vector search, hybrid search, and reranking.
With over 20 successful deployments and 5+ years in the search space, we have refined our approach. A typical scenario: a company has 5 products, each needing search across catalogs, documents, and content. Without a platform, that means 5 independent implementations, 5 times configuring indexes, 5 times paying for GPU for the embedding model. With our platform, it's one shared service, different indexes (tenants), a single API. Infrastructure costs drop by 30–50% thanks to shared GPU, and search deployment time shrinks from months to weeks.
For a typical deployment of 5 products, our clients report infrastructure cost savings of $20K–$50K per year.
How Multi-Tenancy Works in the Search Platform
We provide multi-tenant search with isolated environments. Each client gets an isolated environment: their own collection in Qdrant (or another vector DB), their own indexing pipeline, and flexible limits. We guarantee that data from different customers never mixes, and load is evenly distributed through horizontal scaling.
from fastapi import FastAPI, Header, HTTPException, Depends
from pydantic import BaseModel
from typing import Optional
import uuid
app = FastAPI(title="Search as a Service")
class IndexConfig(BaseModel):
name: str
embedding_model: str = "intfloat/multilingual-e5-large"
chunk_size: int = 512
chunk_overlap: int = 64
language: str = "ru"
reranker_enabled: bool = True
class SearchRequest(BaseModel):
query: str
index_name: str
filters: Optional[dict] = None
top_k: int = 10
mode: str = "hybrid" # "vector" | "keyword" | "hybrid"
generate_answer: bool = False
class SearchService:
def __init__(self):
self.tenant_indexes = {} # tenant_id → {index_name → index}
self.embedding_models = {} # model_name → loaded_model
self.reranker = self._load_reranker()
async def create_index(self, tenant_id: str, config: IndexConfig):
"""Creates an isolated index for a tenant"""
collection_name = f"{tenant_id}_{config.name}"
# Each tenant is a separate collection in Qdrant
# with its own payload filters
self.qdrant.create_collection(
collection_name=collection_name,
vectors_config=VectorParams(
size=self._get_vector_size(config.embedding_model),
distance=Distance.COSINE
)
)
return {"index_id": collection_name, "status": "created"}
async def search(
self,
tenant_id: str,
request: SearchRequest
) -> dict:
collection = f"{tenant_id}_{request.index_name}"
if request.mode == "hybrid":
results = await self._hybrid_search(
collection, request.query,
request.filters, request.top_k
)
elif request.mode == "vector":
results = await self._vector_search(
collection, request.query, request.top_k
)
else:
results = await self._keyword_search(
collection, request.query, request.top_k
)
if request.generate_answer and results:
answer = await self._generate_answer(request.query, results)
return {"results": results, "answer": answer}
return {"results": results}
SDK for Consumer Teams
We provide a Python SDK with minimal dependencies. Teams connect to the platform in 10 minutes—no need to dive into vector DB or LLM details.
# pip install search-platform-sdk
from search_platform import SearchClient
client = SearchClient(
api_key="sk-...",
base_url="https://search.internal.company.com"
)
# Index documents
client.index.upload(
index_name="product-catalog",
documents=[
{"id": "p001", "title": "Dell XPS Laptop", "description": "...",
"price": 89999, "category": "laptops"},
# ...
]
)
# Search
results = client.search(
index_name="product-catalog",
query="thin laptop for video editing",
filters={"price": {"lte": 100000}, "category": "laptops"},
top_k=5,
generate_answer=True
)
print(results.answer) # "Based on your query, I recommend..."
print(results.items) # list of documents with relevance scores
Why Choose Search-as-a-Service Instead of a Homegrown Solution?
Compare: building it yourself requires hiring a team of ML engineers, choosing and training an embedding model, setting up vector search, reranker, load balancer, monitoring—6–12 months of work. Our platform delivers the same functionality in 6–8 weeks, 4× faster, with guaranteed 99.9% SLA and P99 latency under 1 second. We've already stress-tested the architecture with 2M documents and 12 products—result: 35% infrastructure savings, 4 days to onboard a team via SDK.
| Feature | In-House Development | Search-as-a-Service |
|---|---|---|
| Time to deploy | 6–12 months | 6–8 weeks |
| Infrastructure cost | High (separate GPUs, multiple teams) | Up to 35% savings via shared GPU |
| SLA | Lower (no monitoring, no backups) | 99.9% |
| Support for new data sources | Requires custom work | Plugins and SDK |
Case study: a SaaS company migrated from scattered Elasticsearch instances to a unified platform. Three teams connected via SDK in 1 day (without understanding vector databases or embedding models). Infrastructure costs dropped by 35%—a significant saving. Average search response time: 280 ms P50, 650 ms P99 on a 2M document corpus.
Rate Limiting and Monitoring
We control load with dynamic limits per plan and log every request to TimescaleDB.
from fastapi_limiter import FastAPILimiter
from fastapi_limiter.depends import RateLimiter
import redis.asyncio as redis
# Per-tenant limits
TENANT_LIMITS = {
"free": "100/minute",
"pro": "1000/minute",
"enterprise": "unlimited"
}
@app.post("/search")
@limiter.limit(get_tenant_limit) # dynamic limit by plan
async def search_endpoint(
request: SearchRequest,
x_api_key: str = Header(...),
tenant = Depends(authenticate_tenant)
):
return await search_service.search(tenant.id, request)
Billing and Usage Tracking
Each search and each embedding request is logged to TimescaleDB:
CREATE TABLE search_usage (
id BIGSERIAL PRIMARY KEY,
tenant_id TEXT NOT NULL,
index_name TEXT NOT NULL,
query_hash TEXT, -- hash for anonymization
mode TEXT,
latency_ms INTEGER,
result_count INTEGER,
answer_generated BOOLEAN,
tokens_used INTEGER, -- for LLM answer
created_at TIMESTAMPTZ DEFAULT NOW()
);
-- TimescaleDB hypertable for efficient time-range queries
SELECT create_hypertable('search_usage', 'created_at');
Technical Details of Indexing Architecture
Documents go through chunking (splitting into 512-token chunks with 64-token overlap), vectorization via an embedding model, then are saved in Qdrant with full payload. Each tenant gets a separate collection. During hybrid search, results are combined via weighted sum with reranking from a cross-encoder.What's Included
- Audit of current search infrastructure (if any): index analysis, latency measurements, bottleneck identification.
- Design of multi-tenant architecture: choose a vector DB (Qdrant, Pinecone, Weaviate), configure isolation, plan capacity.
- API and SDK development: RESTful API (FastAPI) and Python SDK with support for hybrid search, reranking, and RAG search with answer generation based on RAG with hallucination control via few-shot prompting.
- Integration with existing systems: custom connectors for CMS, ERP, DMS.
- Monitoring and billing: Grafana dashboards, TimescaleDB logs, rate limiting system per plan.
- Documentation and training: fully documented API schema, code examples, 2-hour online training for teams.
- Support and SLA: 99.9% uptime guarantee, incident response within 1 hour.
Process
- Analysis — discuss requirements, load profile, data volume, and desired metrics (P50/P99 latency).
- Design — create architecture, choose stack (embedding model, vector DB, reranker), design tenant schema.
- Implementation — develop API, SDK, indexing pipeline, integration tests.
- Load testing — on your data (or synthetic) measure latency, throughput, identify failure points.
- Deployment — deploy on your infrastructure (AWS/GCP/on-prem) or our cloud. Configure monitoring.
- Acceptance and training — demo, handover documentation, train your team.
SLA and Platform Parameters
| Parameter | Value |
|---|---|
| Latency P50 | < 300 ms |
| Latency P99 | < 1 sec |
| Availability | 99.9% |
| Max document size | 5 MB |
| Supported formats | PDF, DOCX, TXT, HTML, JSON |
| Languages | RU, EN, DE, FR, ES (multilingual-E5) |
| Max tenants | Unlimited (horizontal scaling) |
Timelines: basic platform (API + indexing + hybrid search) — 6–8 weeks; with answer generation, SDK, and billing — 3–4 months. Contact us for a precise estimate of your project — we'll calculate cost and timeline for your task. Get a consultation from our AI engineers.







