Why Search-as-a-Service Is Faster and Cheaper

We design and deploy artificial intelligence systems: from prototype to production-ready solutions. Our team combines expertise in machine learning, data engineering and MLOps to make AI work not in the lab, but in real business.
Showing 1 of 1All 1564 services
Why Search-as-a-Service Is Faster and Cheaper
Complex
from 1 week to 3 months
Frequently Asked Questions

AI Development Areas

AI Solution Development Stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1357
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1250
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    956
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1188
  • image_logo-advance_0.webp
    B2B Advance company logo design
    646
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    929

Why Search-as-a-Service Is Faster Than In-House Solutions

We often see companies spending months building AI search from scratch for each product: writing their own vectorization, setting up reranking, configuring infrastructure. Search as a Service is a ready-made semantic search platform with an AI layer that we connect via a unified search API and search SDK. Our engineers with 10+ years of MLOps experience handle all infrastructure: from choosing an embedding model to configuring vector search, hybrid search, and reranking.

With over 20 successful deployments and 5+ years in the search space, we have refined our approach. A typical scenario: a company has 5 products, each needing search across catalogs, documents, and content. Without a platform, that means 5 independent implementations, 5 times configuring indexes, 5 times paying for GPU for the embedding model. With our platform, it's one shared service, different indexes (tenants), a single API. Infrastructure costs drop by 30–50% thanks to shared GPU, and search deployment time shrinks from months to weeks.

For a typical deployment of 5 products, our clients report infrastructure cost savings of $20K–$50K per year.

How Multi-Tenancy Works in the Search Platform

We provide multi-tenant search with isolated environments. Each client gets an isolated environment: their own collection in Qdrant (or another vector DB), their own indexing pipeline, and flexible limits. We guarantee that data from different customers never mixes, and load is evenly distributed through horizontal scaling.

from fastapi import FastAPI, Header, HTTPException, Depends
from pydantic import BaseModel
from typing import Optional
import uuid

app = FastAPI(title="Search as a Service")

class IndexConfig(BaseModel):
    name: str
    embedding_model: str = "intfloat/multilingual-e5-large"
    chunk_size: int = 512
    chunk_overlap: int = 64
    language: str = "ru"
    reranker_enabled: bool = True

class SearchRequest(BaseModel):
    query: str
    index_name: str
    filters: Optional[dict] = None
    top_k: int = 10
    mode: str = "hybrid"  # "vector" | "keyword" | "hybrid"
    generate_answer: bool = False

class SearchService:
    def __init__(self):
        self.tenant_indexes = {}  # tenant_id → {index_name → index}
        self.embedding_models = {}  # model_name → loaded_model
        self.reranker = self._load_reranker()

    async def create_index(self, tenant_id: str, config: IndexConfig):
        """Creates an isolated index for a tenant"""
        collection_name = f"{tenant_id}_{config.name}"
        # Each tenant is a separate collection in Qdrant
        # with its own payload filters
        self.qdrant.create_collection(
            collection_name=collection_name,
            vectors_config=VectorParams(
                size=self._get_vector_size(config.embedding_model),
                distance=Distance.COSINE
            )
        )
        return {"index_id": collection_name, "status": "created"}

    async def search(
        self,
        tenant_id: str,
        request: SearchRequest
    ) -> dict:
        collection = f"{tenant_id}_{request.index_name}"

        if request.mode == "hybrid":
            results = await self._hybrid_search(
                collection, request.query,
                request.filters, request.top_k
            )
        elif request.mode == "vector":
            results = await self._vector_search(
                collection, request.query, request.top_k
            )
        else:
            results = await self._keyword_search(
                collection, request.query, request.top_k
            )

        if request.generate_answer and results:
            answer = await self._generate_answer(request.query, results)
            return {"results": results, "answer": answer}

        return {"results": results}

SDK for Consumer Teams

We provide a Python SDK with minimal dependencies. Teams connect to the platform in 10 minutes—no need to dive into vector DB or LLM details.

# pip install search-platform-sdk
from search_platform import SearchClient

client = SearchClient(
    api_key="sk-...",
    base_url="https://search.internal.company.com"
)

# Index documents
client.index.upload(
    index_name="product-catalog",
    documents=[
        {"id": "p001", "title": "Dell XPS Laptop", "description": "...",
         "price": 89999, "category": "laptops"},
        # ...
    ]
)

# Search
results = client.search(
    index_name="product-catalog",
    query="thin laptop for video editing",
    filters={"price": {"lte": 100000}, "category": "laptops"},
    top_k=5,
    generate_answer=True
)

print(results.answer)   # "Based on your query, I recommend..."
print(results.items)    # list of documents with relevance scores

Why Choose Search-as-a-Service Instead of a Homegrown Solution?

Compare: building it yourself requires hiring a team of ML engineers, choosing and training an embedding model, setting up vector search, reranker, load balancer, monitoring—6–12 months of work. Our platform delivers the same functionality in 6–8 weeks, 4× faster, with guaranteed 99.9% SLA and P99 latency under 1 second. We've already stress-tested the architecture with 2M documents and 12 products—result: 35% infrastructure savings, 4 days to onboard a team via SDK.

Feature In-House Development Search-as-a-Service
Time to deploy 6–12 months 6–8 weeks
Infrastructure cost High (separate GPUs, multiple teams) Up to 35% savings via shared GPU
SLA Lower (no monitoring, no backups) 99.9%
Support for new data sources Requires custom work Plugins and SDK

Case study: a SaaS company migrated from scattered Elasticsearch instances to a unified platform. Three teams connected via SDK in 1 day (without understanding vector databases or embedding models). Infrastructure costs dropped by 35%—a significant saving. Average search response time: 280 ms P50, 650 ms P99 on a 2M document corpus.

Rate Limiting and Monitoring

We control load with dynamic limits per plan and log every request to TimescaleDB.

from fastapi_limiter import FastAPILimiter
from fastapi_limiter.depends import RateLimiter
import redis.asyncio as redis

# Per-tenant limits
TENANT_LIMITS = {
    "free": "100/minute",
    "pro": "1000/minute",
    "enterprise": "unlimited"
}

@app.post("/search")
@limiter.limit(get_tenant_limit)  # dynamic limit by plan
async def search_endpoint(
    request: SearchRequest,
    x_api_key: str = Header(...),
    tenant = Depends(authenticate_tenant)
):
    return await search_service.search(tenant.id, request)

Billing and Usage Tracking

Each search and each embedding request is logged to TimescaleDB:

CREATE TABLE search_usage (
    id          BIGSERIAL PRIMARY KEY,
    tenant_id   TEXT NOT NULL,
    index_name  TEXT NOT NULL,
    query_hash  TEXT,           -- hash for anonymization
    mode        TEXT,
    latency_ms  INTEGER,
    result_count INTEGER,
    answer_generated BOOLEAN,
    tokens_used  INTEGER,       -- for LLM answer
    created_at  TIMESTAMPTZ DEFAULT NOW()
);

-- TimescaleDB hypertable for efficient time-range queries
SELECT create_hypertable('search_usage', 'created_at');
Technical Details of Indexing Architecture Documents go through chunking (splitting into 512-token chunks with 64-token overlap), vectorization via an embedding model, then are saved in Qdrant with full payload. Each tenant gets a separate collection. During hybrid search, results are combined via weighted sum with reranking from a cross-encoder.

What's Included

  • Audit of current search infrastructure (if any): index analysis, latency measurements, bottleneck identification.
  • Design of multi-tenant architecture: choose a vector DB (Qdrant, Pinecone, Weaviate), configure isolation, plan capacity.
  • API and SDK development: RESTful API (FastAPI) and Python SDK with support for hybrid search, reranking, and RAG search with answer generation based on RAG with hallucination control via few-shot prompting.
  • Integration with existing systems: custom connectors for CMS, ERP, DMS.
  • Monitoring and billing: Grafana dashboards, TimescaleDB logs, rate limiting system per plan.
  • Documentation and training: fully documented API schema, code examples, 2-hour online training for teams.
  • Support and SLA: 99.9% uptime guarantee, incident response within 1 hour.

Process

  1. Analysis — discuss requirements, load profile, data volume, and desired metrics (P50/P99 latency).
  2. Design — create architecture, choose stack (embedding model, vector DB, reranker), design tenant schema.
  3. Implementation — develop API, SDK, indexing pipeline, integration tests.
  4. Load testing — on your data (or synthetic) measure latency, throughput, identify failure points.
  5. Deployment — deploy on your infrastructure (AWS/GCP/on-prem) or our cloud. Configure monitoring.
  6. Acceptance and training — demo, handover documentation, train your team.

SLA and Platform Parameters

Parameter Value
Latency P50 < 300 ms
Latency P99 < 1 sec
Availability 99.9%
Max document size 5 MB
Supported formats PDF, DOCX, TXT, HTML, JSON
Languages RU, EN, DE, FR, ES (multilingual-E5)
Max tenants Unlimited (horizontal scaling)

Timelines: basic platform (API + indexing + hybrid search) — 6–8 weeks; with answer generation, SDK, and billing — 3–4 months. Contact us for a precise estimate of your project — we'll calculate cost and timeline for your task. Get a consultation from our AI engineers.

LLM Development: Fine-Tuning, RAG, Agents, and Production Deployment

Using GPT‑4 or Claude 3.5 Sonnet through a public API is not a solution — it's just a tool. When the requirement is to "make it like ChatGPT, but on our data," there is a real engineering challenge behind it: from prompt engineering to training a 70B model on your own infrastructure. End-to-end LLM solution development is a complex stack, and we have been doing it for over 5 years. During this time, we have completed over 20 projects in generative AI: from RAG systems for legal departments to custom support agents. Where exactly your task falls depends on data, latency requirements, budget, and how critical confidentiality is.

A typical situation: the client has already tried ChatGPT, but results are unstable — sometimes accurate, sometimes hallucinating. Or they need integration into a corporate portal while complying with security policies. Let's break down each layer of the stack in detail — from RAG to production deployment.

Why Do RAG Systems Break and How to Fix It?

RAG (Retrieval-Augmented Generation) looks simple: find relevant documents, put them in context, get an answer. In practice, it fails in several places.

Chunking without overlap. Classic mistake: chunk_size=512, overlap=0. If the answer lies across two chunks, retrieval won't find either with sufficient confidence. Solution: overlap 15–25% of chunk_size, or better yet, sentence-aware splitting with spaCy or NLTK instead of naive character splitting.

Poor embedder. text-embedding-ada-002 is good for general use, but on legal or medical texts, specialized models like E5-large-v2, BGE-M3, or fine-tuned sentence-transformers on domain data outperform it. Recall@5 differences can be 15–25%.

No re-ranking. Vector search optimizes for speed, not relevance. A cross-encoder re-ranker (ms-marco-MiniLM-L-6-v2, bge-reranker-large) after initial retrieval improves top-3 accuracy with acceptable latency (+50–150ms). This is often more impactful than improving the embedding model.

Hybrid search. Dense vectors alone work poorly on exact queries: names, SKUs, codes. BM25 (sparse) finds exact matches but misses semantics. Hybrid via RRF (Reciprocal Rank Fusion) is the optimal compromise. Qdrant, Weaviate, and pgvector 0.7+ support hybrid search natively.

Typical production architecture for a corporate knowledge base
  1. Documents → preprocessing (PyMuPDF, Unstructured)
  2. Chunking → embedding (BGE-M3)
  3. Qdrant (hybrid dense+sparse)
  4. Cross-encoder re-ranking
  5. Context → LLM (vLLM or OpenAI API)
  6. Answer with sources (RAGAS for quality evaluation)

When to Fine-Tune Instead of Prompt Engineering?

Prompt engineering solves ~70% of LLM adaptation tasks for a domain. The remaining 30% require fine-tuning. Three indicators: the model ignores a specific output format even with detailed prompting; the task requires deep knowledge of specialized vocabulary (medicine, law); you need to significantly reduce token costs by replacing a large model with a smaller specialized one.

LoRA and QLoRA are the standard for SFT. LoRA adds trainable low-rank matrices to attention layers. A typical configuration for Llama-3 8B: r=64, lora_alpha=128, target_modules=["q_proj","v_proj","k_proj","o_proj"] yields ~0.8% trainable parameters, training on one A100 40GB. QLoRA adds 4-bit quantization (NF4) and allows fine-tuning 70B models on two A100 40GB, though speed drops by half compared to bf16.

DPO instead of RLHF. Direct Preference Optimization requires only (chosen, rejected) pairs, not scalar reward signals. DPOTrainer from the trl library (Hugging Face) implements it in a few dozen lines.

Common mistake. A dataset of 500 examples, 5 epochs, validation loss 0.8 — seems fine. But on test, the model degrades on general instructions. Cause: catastrophic forgetting. Solution: add 10–20% general instruction-following examples (Alpaca, FLAN) to the training set to preserve original capabilities.

How to Choose a Base Model: 8B or 70B?

Model Parameters Strengths Context
Llama-3.1 8B 8B Quality/speed balance 128k
Llama-3.1 70B 70B Complex reasoning 128k
Mistral 7B / Mixtral 8x7B 7B / 47B Efficiency for size 32k
Qwen2.5 72B 72B Code, multilingual 128k
Gemma 2 27B 27B Open license 8k

For most tasks, fine-tuning an 8B model is sufficient. 70B is needed when deep reasoning is required or the 8B baseline does not reach the required quality even after fine-tuning. Inference cost for Llama-3 8B via vLLM on A100 is efficient; the exact cost depends on volume.

What Does PagedAttention Bring to Production?

vLLM is the first choice for serving open-source models. PagedAttention is the key technical innovation: KV-cache is managed like virtual memory in an OS, without fragmentation. This yields 2–4x higher throughput compared to naive HuggingFace Transformers inference. The vLLM documentation confirms that continuous batching and PagedAttention are the standard for high-load LLM services.

Typical numbers on A100 80GB for Llama-3 8B (bf16): 400–600 req/s, P50 latency 200–400ms, P99 latency 600–900ms at concurrency 64. For 70B on two A100 with tensor parallelism: 80–120 req/s, P99 latency 1.5–2.5s. AWQ or GPTQ quantization reduces memory consumption by 2x with quality loss within 1–3%.

Multi-Agent Systems

Agents are LLMs with access to tools: search, code execution, API calls, database interaction. Common patterns:

  • ReAct (Reason + Act): the model reasons → chooses a tool → observes the result → reasons again. LangChain and LlamaIndex implement it out of the box.
  • Multi-agent orchestration: multiple specialized agents with a coordinator on top. Example: coordinator → researcher (search + summarization) → coder (code generation and execution) → critic (verification). Tools: AutoGen (Microsoft), CrewAI, custom implementation on LangGraph.

In production, agent systems are non-deterministic. Essential: guardrails, step limits, logging of each step, human-in-the-loop for critical actions.

How We Work: Stages, Timeline, Deliverables

Stage Duration What You Get
Audit and data collection 1–2 weeks Eval dataset of 100+ examples, task formalization
Baseline (prompt + RAG) 1–2 weeks Working prototype, quality metrics
Fine-tuning (if needed) 2–4 weeks Trained model, LoRA weights, model card
Deployment and monitoring 1–2 weeks vLLM server, Grafana + Prometheus
Documentation and training 1 week API documentation, team training

What Is Included

We deliver:

  • Technical documentation (model card, configs, deployment instructions)
  • Access to infrastructure (code repository, trained weights)
  • 1 month of post-deployment support (consultations, bug fixes)
  • Customer team training (2–3 sessions on system operation)

Timeline: basic RAG prototype — 1–2 weeks. Fine-tuning with customer data — 3–6 weeks (including data preparation). Production system with monitoring and retraining — 2–4 months. Cost is calculated individually based on data volume, model complexity, and infrastructure requirements.

We guarantee the quality of the final model with performance benchmarks and ongoing monitoring. Our engineers have hands‑on experience with dozens of production LLM systems.

Want to evaluate your project? Leave a request — we will prepare a preliminary summary within 1–2 business days. Or get a consultation on choosing the approach: RAG, fine-tuning, or hybrid — we will tell you what works best for you. Contact us to discuss your LLM development needs. Schedule a free consultation today.