AI Enterprise Search: Intelligent Corporate Search

We design and deploy artificial intelligence systems: from prototype to production-ready solutions. Our team combines expertise in machine learning, data engineering and MLOps to make AI work not in the lab, but in real business.
Showing 1 of 1All 1564 services
AI Enterprise Search: Intelligent Corporate Search
Medium
~2-4 weeks
Frequently Asked Questions

AI Development Areas

AI Solution Development Stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1357
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1250
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    956
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1188
  • image_logo-advance_0.webp
    B2B Advance company logo design
    646
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    929

In a company of 500 employees, knowledge is stored across 7–12 systems simultaneously: Confluence, SharePoint, Jira, corporate email, 1C, internal CRMs, file servers. An employee spends an average of 2.5 hours per day searching for information — this is not a metaphor, it's data from McKinsey Global Institute. The key problem is not that data exists, but that finding it through ordinary keyword search is impossible: a document is named 'Regulation_v3_final_FINAL2', and the query is 'how to book a business trip'. We develop AI document search — a system that solves this problem through semantic understanding of the query and hybrid search across all sources simultaneously.

Technology Approach

Problems We Solve

Data silos are not the only complexity. Even if a document is found, it may be outdated or irrelevant to the context. Typical scenarios:

  • Low keyword search precision: the part number 'A-123' is searched as an ordinary word, without understanding it is a part number.
  • Lack of context: the query 'marketing budget' should consider that the user is a department head, not an intern.
  • Security breaches: an employee sees documents they should not have access to if ACL is not implemented.

Why AI Enterprise Search Is More Effective Than Regular Search

Pure vector search works well on semantically similar queries but poorly on exact matches: part numbers, names, dates. The query 'contract with Romashka LLC dated March 12' yields better results via BM25, while 'procedure for data breach incident' works better via vector search. The hybrid approach combines the strengths of both methods.

Search Type Strengths Weaknesses
Keyword (BM25) Exact matches, part numbers, dates Does not understand synonyms, context
Vector (Embeddings) Semantics, synonyms, generalizations Poor on rare terms, cold start
Hybrid (ours) Best of both worlds, CrossEncoder reranking More complex implementation, requires fine-tuning
from qdrant_client import QdrantClient
from rank_bm25 import BM25Okapi
from sentence_transformers import SentenceTransformer, CrossEncoder
import numpy as np

class HybridSearchEngine:
    def __init__(self):
        self.dense_model = SentenceTransformer("intfloat/multilingual-e5-large")
        self.reranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2")
        self.qdrant = QdrantClient(url="http://localhost:6333")
        self.bm25_index = None
        self.doc_store = {}

    def search(self, query: str, top_k: int = 20, final_k: int = 5) -> list[dict]:
        # Dense search
        query_emb = self.dense_model.encode(
            f"query: {query}",  # E5 format
            normalize_embeddings=True
        )
        dense_results = self.qdrant.search(
            collection_name="enterprise_docs",
            query_vector=query_emb.tolist(),
            limit=top_k
        )

        # Sparse search (BM25)
        bm25_scores = self.bm25_index.get_scores(query.lower().split())
        top_bm25_ids = np.argsort(bm25_scores)[-top_k:][::-1]

        # Merge results (RRF - Reciprocal Rank Fusion)
        merged = self._reciprocal_rank_fusion(dense_results, top_bm25_ids)

        # Cross-encoder reranking
        pairs = [(query, self.doc_store[doc_id]["text"]) for doc_id in merged[:top_k]]
        rerank_scores = self.reranker.predict(pairs)
        reranked = sorted(
            zip(merged[:top_k], rerank_scores),
            key=lambda x: x[1],
            reverse=True
        )

        return [self.doc_store[doc_id] for doc_id, _ in reranked[:final_k]]

    def _reciprocal_rank_fusion(self, dense: list, sparse: list, k: int = 60) -> list:
        scores = {}
        for rank, item in enumerate(dense):
            doc_id = item.id
            scores[doc_id] = scores.get(doc_id, 0) + 1 / (k + rank + 1)
        for rank, doc_id in enumerate(sparse):
            scores[doc_id] = scores.get(doc_id, 0) + 1 / (k + rank + 1)
        return sorted(scores, key=scores.get, reverse=True)

Query Understanding and Answer Generation

Our RAG system combines the hybrid search engine with a language model to generate precise answers from company documents.

from langchain_openai import ChatOpenAI
from langchain_core.prompts import ChatPromptTemplate

class EnterpriseSearchAssistant:
    def __init__(self, search_engine: HybridSearchEngine):
        self.search = search_engine
        self.llm = ChatOpenAI(model="gpt-4o-mini", temperature=0.1)

    def answer(self, query: str, user_context: dict) -> dict:
        # Expand query for better search
        expanded_query = self._expand_query(query, user_context)

        # Search across all sources
        results = self.search.search(expanded_query, final_k=7)

        if not results:
            return {"answer": "No documents found for your query.", "sources": []}

        # Generate answer with citations
        context = "\n\n".join([
            f"[{i+1}] {r['title']} ({r['source']})\n{r['text'][:500]}"
            for i, r in enumerate(results)
        ])

        prompt = f"""Answer the employee's question based on company documents.
Use only the provided documents. Cite sources [1], [2], etc.
If the answer is incomplete, say so.

Question: {query}
Documents:
{context}"""

        answer = self.llm.invoke(prompt).content
        return {
            "answer": answer,
            "sources": [{"id": i+1, "title": r["title"], "url": r["url"]}
                        for i, r in enumerate(results)]
        }

Implementation and Security

Source Indexing

Each source has a separate connector with incremental updates:

Source Connector Indexing Frequency Peculiarities
Confluence REST API (spaces, pages) Every 30 min Confluence markup processing
SharePoint Microsoft Graph API Every hour Word/PDF/PPTX support
Jira REST API Every 15 min Tickets + comments
Email IMAP/Exchange Realtime (webhooks) Only sent/received
File server Inotify / polling On change PDF, DOCX, XLSX, TXT

How We Ensure Corporate Data Security

A critical detail of access-rights search — users must see only documents they are authorized to access. Implemented via metadata filters in Qdrant:

# When searching, filter by user permissions
results = qdrant.search(
    collection_name="enterprise_docs",
    query_vector=query_emb,
    query_filter=Filter(
        must=[
            FieldCondition(
                key="allowed_groups",
                match=MatchAny(any=user_groups)
            )
        ]
    ),
    limit=20
)
Example of access control configurationACL is configured via document metadata: groups, roles, permissions. The Qdrant vector database supports filtering at the point (document) level, ensuring security.

Our Process

  1. Analytics and source audit: identify where data resides, estimate volume and change frequency.
  2. Architecture design: choose embedding model, vector DB, configure pipeline.
  3. Data indexing: connect connectors, perform trial indexing on 10% of data.
  4. Search and reranking tuning: optimize RRF parameters, set thresholds for CrossEncoder.
  5. UI and bot integration: embed search in web interface, Slack bot, API.
  6. Testing and debugging: test on real employee queries, iteratively improve.
  7. Deployment and monitoring: deploy to production, set up logs and metrics (latency p99, recall).

What Is Included

  • Architecture documentation and operational instructions.
  • Configured connectors to all sources.
  • Web search interface and integration with corporate messenger.
  • Team training (2-3 workshops).
  • 3-month warranty support after launch.
  • Access control setup documentation.

Timelines and Cost

  • Pilot (1-2 sources, ~50,000 documents): 4–6 weeks.
  • Full system (5–8 sources + UI): 3–4 months.
  • ACL setup: +2–3 weeks.

Cost is calculated individually after audit. On average, investments pay back in 6–12 months due to reduced search time. For a 500-employee company, we estimate annual savings of $200,000 based on reduced search time. Order a pilot project — the first stage takes no more than 2 days for analysis.

Results and Case Study

One of our clients, a manufacturing holding company with 1200 employees. We indexed 340,000 documents from Confluence, SharePoint, and a custom document management system. Index loading time was 18 hours (one-time). Regular updates — delta of 30 minutes, ~200 documents, taking 4 minutes. Average top-3 accuracy (evaluated by HR team on a sample of 500 queries): 78% relevant results compared to 41% with the previous Elasticsearch keyword search (1.9 times better). By automating search, the company achieved significant budget savings — approximately $150,000 annually.

How Search Accuracy Is Evaluated

Evaluation is performed on real employee queries. We form a sample of 300–500 typical questions, label relevant documents, and compute recall@k and precision@k. The key tool is CrossEncoder, which ranks candidates and improves top-3 accuracy to 78%. Validation is done by the client on their data.

Why Choose Us

  • 5+ years in enterprise AI solutions.
  • 20+ projects implementing semantic search in companies from 200 to 5000 employees.
  • Guarantee of SLA 99.9% and search accuracy no less than 75% according to your criteria.
  • Certified MLOps and NLP engineers.

Contact us to get a consultation on your project and a preliminary estimate within 2 days.

LLM Development: Fine-Tuning, RAG, Agents, and Production Deployment

Using GPT‑4 or Claude 3.5 Sonnet through a public API is not a solution — it's just a tool. When the requirement is to "make it like ChatGPT, but on our data," there is a real engineering challenge behind it: from prompt engineering to training a 70B model on your own infrastructure. End-to-end LLM solution development is a complex stack, and we have been doing it for over 5 years. During this time, we have completed over 20 projects in generative AI: from RAG systems for legal departments to custom support agents. Where exactly your task falls depends on data, latency requirements, budget, and how critical confidentiality is.

A typical situation: the client has already tried ChatGPT, but results are unstable — sometimes accurate, sometimes hallucinating. Or they need integration into a corporate portal while complying with security policies. Let's break down each layer of the stack in detail — from RAG to production deployment.

Why Do RAG Systems Break and How to Fix It?

RAG (Retrieval-Augmented Generation) looks simple: find relevant documents, put them in context, get an answer. In practice, it fails in several places.

Chunking without overlap. Classic mistake: chunk_size=512, overlap=0. If the answer lies across two chunks, retrieval won't find either with sufficient confidence. Solution: overlap 15–25% of chunk_size, or better yet, sentence-aware splitting with spaCy or NLTK instead of naive character splitting.

Poor embedder. text-embedding-ada-002 is good for general use, but on legal or medical texts, specialized models like E5-large-v2, BGE-M3, or fine-tuned sentence-transformers on domain data outperform it. Recall@5 differences can be 15–25%.

No re-ranking. Vector search optimizes for speed, not relevance. A cross-encoder re-ranker (ms-marco-MiniLM-L-6-v2, bge-reranker-large) after initial retrieval improves top-3 accuracy with acceptable latency (+50–150ms). This is often more impactful than improving the embedding model.

Hybrid search. Dense vectors alone work poorly on exact queries: names, SKUs, codes. BM25 (sparse) finds exact matches but misses semantics. Hybrid via RRF (Reciprocal Rank Fusion) is the optimal compromise. Qdrant, Weaviate, and pgvector 0.7+ support hybrid search natively.

Typical production architecture for a corporate knowledge base
  1. Documents → preprocessing (PyMuPDF, Unstructured)
  2. Chunking → embedding (BGE-M3)
  3. Qdrant (hybrid dense+sparse)
  4. Cross-encoder re-ranking
  5. Context → LLM (vLLM or OpenAI API)
  6. Answer with sources (RAGAS for quality evaluation)

When to Fine-Tune Instead of Prompt Engineering?

Prompt engineering solves ~70% of LLM adaptation tasks for a domain. The remaining 30% require fine-tuning. Three indicators: the model ignores a specific output format even with detailed prompting; the task requires deep knowledge of specialized vocabulary (medicine, law); you need to significantly reduce token costs by replacing a large model with a smaller specialized one.

LoRA and QLoRA are the standard for SFT. LoRA adds trainable low-rank matrices to attention layers. A typical configuration for Llama-3 8B: r=64, lora_alpha=128, target_modules=["q_proj","v_proj","k_proj","o_proj"] yields ~0.8% trainable parameters, training on one A100 40GB. QLoRA adds 4-bit quantization (NF4) and allows fine-tuning 70B models on two A100 40GB, though speed drops by half compared to bf16.

DPO instead of RLHF. Direct Preference Optimization requires only (chosen, rejected) pairs, not scalar reward signals. DPOTrainer from the trl library (Hugging Face) implements it in a few dozen lines.

Common mistake. A dataset of 500 examples, 5 epochs, validation loss 0.8 — seems fine. But on test, the model degrades on general instructions. Cause: catastrophic forgetting. Solution: add 10–20% general instruction-following examples (Alpaca, FLAN) to the training set to preserve original capabilities.

How to Choose a Base Model: 8B or 70B?

Model Parameters Strengths Context
Llama-3.1 8B 8B Quality/speed balance 128k
Llama-3.1 70B 70B Complex reasoning 128k
Mistral 7B / Mixtral 8x7B 7B / 47B Efficiency for size 32k
Qwen2.5 72B 72B Code, multilingual 128k
Gemma 2 27B 27B Open license 8k

For most tasks, fine-tuning an 8B model is sufficient. 70B is needed when deep reasoning is required or the 8B baseline does not reach the required quality even after fine-tuning. Inference cost for Llama-3 8B via vLLM on A100 is efficient; the exact cost depends on volume.

What Does PagedAttention Bring to Production?

vLLM is the first choice for serving open-source models. PagedAttention is the key technical innovation: KV-cache is managed like virtual memory in an OS, without fragmentation. This yields 2–4x higher throughput compared to naive HuggingFace Transformers inference. The vLLM documentation confirms that continuous batching and PagedAttention are the standard for high-load LLM services.

Typical numbers on A100 80GB for Llama-3 8B (bf16): 400–600 req/s, P50 latency 200–400ms, P99 latency 600–900ms at concurrency 64. For 70B on two A100 with tensor parallelism: 80–120 req/s, P99 latency 1.5–2.5s. AWQ or GPTQ quantization reduces memory consumption by 2x with quality loss within 1–3%.

Multi-Agent Systems

Agents are LLMs with access to tools: search, code execution, API calls, database interaction. Common patterns:

  • ReAct (Reason + Act): the model reasons → chooses a tool → observes the result → reasons again. LangChain and LlamaIndex implement it out of the box.
  • Multi-agent orchestration: multiple specialized agents with a coordinator on top. Example: coordinator → researcher (search + summarization) → coder (code generation and execution) → critic (verification). Tools: AutoGen (Microsoft), CrewAI, custom implementation on LangGraph.

In production, agent systems are non-deterministic. Essential: guardrails, step limits, logging of each step, human-in-the-loop for critical actions.

How We Work: Stages, Timeline, Deliverables

Stage Duration What You Get
Audit and data collection 1–2 weeks Eval dataset of 100+ examples, task formalization
Baseline (prompt + RAG) 1–2 weeks Working prototype, quality metrics
Fine-tuning (if needed) 2–4 weeks Trained model, LoRA weights, model card
Deployment and monitoring 1–2 weeks vLLM server, Grafana + Prometheus
Documentation and training 1 week API documentation, team training

What Is Included

We deliver:

  • Technical documentation (model card, configs, deployment instructions)
  • Access to infrastructure (code repository, trained weights)
  • 1 month of post-deployment support (consultations, bug fixes)
  • Customer team training (2–3 sessions on system operation)

Timeline: basic RAG prototype — 1–2 weeks. Fine-tuning with customer data — 3–6 weeks (including data preparation). Production system with monitoring and retraining — 2–4 months. Cost is calculated individually based on data volume, model complexity, and infrastructure requirements.

We guarantee the quality of the final model with performance benchmarks and ongoing monitoring. Our engineers have hands‑on experience with dozens of production LLM systems.

Want to evaluate your project? Leave a request — we will prepare a preliminary summary within 1–2 business days. Or get a consultation on choosing the approach: RAG, fine-tuning, or hybrid — we will tell you what works best for you. Contact us to discuss your LLM development needs. Schedule a free consultation today.