AI Assistant for Corporate Knowledge Base Development

We design and deploy artificial intelligence systems: from prototype to production-ready solutions. Our team combines expertise in machine learning, data engineering and MLOps to make AI work not in the lab, but in real business.

AI Development Areas

AI Solution Development Stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1359
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1251
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    957
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1188
  • image_logo-advance_0.webp
    B2B Advance company logo design
    646
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    929

AI Assistant for Corporate Knowledge Base

Corporate knowledge bases—Confluence, Notion, SharePoint, internal wikis—store terabytes of information, yet employees spend 45 minutes a day searching for it. We develop AI-powered assistants based on RAG (Retrieval-Augmented Generation) that make knowledge accessible through dialogue: ask a question in natural language and get an answer with source citations. No multi-level menus or obscure tags. This kind of assistant is often called a corporate GPT—it not only answers but also enables support automation, reducing the load on experts. It acts as a company chatbot.

Problems We Solve

First—information chaos. Documents are scattered across different systems, duplicated, and outdated. Employees don't know where to find regulations, instructions, or contacts. Second—low search accuracy. Confluence or SharePoint's built-in search finds by keywords but doesn't understand meaning. A query like "how to request time off" might return 200 pages, none of which are relevant. Third—contextual overload. Even after finding a document, the employee has to read it entirely to extract the answer. The knowledge bot solves all three problems through semantic search and answer generation from a specific fragment. Traditional document search fails here.

Why RAG Instead of Fine-Tuning?

Fine-tuning an LLM on corporate documents is expensive and inflexible. The model "remembers" facts but cannot cite a specific document, and when policies change, it requires retraining. RAG (Retrieval-Augmented Generation) retrieves relevant fragments from an index at query time and feeds them into the LLM context. This way, answers are always current and sources are transparent. RAG is 3-5 times more relevant than fine-tuning for the same indexing cost.

Criterion RAG Fine-tuning
Answer timeliness always latest version requires retraining
Source citations yes (document link) no
Indexing cost low (once + incremental) high (per dataset)
Hallucinations minimal (context limited) possible
Deployment time 2-3 weeks 4-8 weeks

RAG outperforms fine-tuning in answer relevance by 3-5x. Our LLM for business ensures that answers are grounded in your data. Use this approach in 90% of our projects.

More on chunkingChunk size affects accuracy: too small fragments lose context, too large ones increase noise. The optimal size is 1000 tokens with 200 overlap.

How We Build the AI Assistant

Architecture of a typical solution: indexer (parsing Confluence via REST API, Notion API, file systems) → vectorization (text-embedding-3-small from OpenAI, 1536-dimensional embeddings) → vector store (ChromaDB, Qdrant, or pgvector) → RAG pipeline (LangChain or LlamaIndex) → LLM (Claude 3.5 Sonnet / GPT-4o). Learn more about RAG on Wikipedia. Slack integration is included in the standard deployment.

# example code from above remains unchanged

Case Study: IT Company, 200 Employees

Problem: 45 minutes per day searching for information (survey data). Confluence with 3,200 pages, most documents—dead weight.

Our client—a mid-sized IT company with distributed teams. We deployed:

  • Indexing all 3,200 Confluence pages in 4 hours;
  • Telegram bot and Slack bot for access;
  • Daily incremental synchronization.

Results:

  • Search time reduced from 45 to 8 minutes per employee per day;
  • "Who knows where this is?" Slack queries dropped by 71%;
  • Answer accuracy—4.3/5.0 user rating;
  • 9% of queries could not be processed—escalated to experts.

This translates to annual savings of $15,000 for a 200-employee company, assuming average hourly wage of $50. Basic deployment starts at $5,000 for up to 5,000 pages and 3 sources.

How the Process Works?

  1. Source audit: We gather a list of all systems where documents are stored (Confluence, Notion, Google Drive, file shares).
  2. Index design: We define metadata fields (author, date, access rights), configure chunking (chunk size 1000 tokens, overlap 200).
  3. RAG pipeline implementation: Write integration with selected sources, deploy vector DB, connect LLM.
  4. Testing and calibration: Verify accuracy on 100+ representative queries, adjust refusal threshold.
  5. Deployment and monitoring: Install bot in corporate messengers, set up logging and latency p99 metrics.

What's Included?

  • Audit of the existing knowledge base and cleanup recommendations;
  • Document indexing of up to 10,000 pages (larger volumes—separate pricing);
  • Integration with 1-3 sources (including Confluence AI);
  • Telegram bot and Slack bot;
  • Access: User group access control;
  • Training: Team training session on using the assistant;
  • Support: Email and chat support for 30 days post-deployment;
  • Uptime guarantee 99.5% (SLA available on request);
  • Documentation: Architecture and API documentation.

Estimated Timelines

Stage Duration
Audit and design 3–5 days
Indexing and basic RAG chain 5–7 days
Messenger integration 2–3 days
Testing and refinements 3–5 days
Deployment and training 2 days

Total time—from 2 to 4 weeks depending on the number of sources and volumes.

FAQ

Q: What data sources does the AI assistant support? A: We integrate with Confluence, Notion, SharePoint, Google Drive, Jira, and any sources via REST API or file storage. The assistant indexes documents, wikis, and knowledge bases, making them accessible through natural language search.

Q: How is data security ensured when using the AI assistant? A: All LLM requests are processed within your environment or via private deployments (Azure OpenAI, AWS Bedrock). We implement document-level access control: employees only see pages they are allowed to.

Q: How long does it take to deploy a basic AI assistant? A: A basic version with Confluence integration and Telegram/Slack bot is deployed in 2-3 weeks. This includes indexing up to 5,000 pages, setting up the RAG pipeline, and testing answer accuracy.

Q: Which LLM is used for answering? A: We use Claude 3.5 Sonnet, GPT-4o, or LLaMA 3 depending on speed and confidentiality requirements. The model can be replaced or fine-tuned on corporate terminology.

Q: What happens if the AI assistant does not know the answer? A: The assistant informs that the information is not in the knowledge base and suggests contacting specific colleagues or searching manually. We configure escalation rules for routing complex requests.

Why Order from Us?

We have 5+ years of experience in NLP and Computer Vision, with 20+ implemented knowledge solutions for business. We guarantee a transparent architecture with no vendor lock-in: we use open-source models and libraries (LangChain, ChromaDB). Get a consultation—write to us, we will evaluate your project and offer a turnkey solution. Contact us for a preliminary assessment: just send a description of your knowledge base, and we will prepare a customized proposal.

LLM Development: Fine-Tuning, RAG, Agents, and Production Deployment

Using GPT‑4 or Claude 3.5 Sonnet through a public API is not a solution — it's just a tool. When the requirement is to "make it like ChatGPT, but on our data," there is a real engineering challenge behind it: from prompt engineering to training a 70B model on your own infrastructure. End-to-end LLM solution development is a complex stack, and we have been doing it for over 5 years. During this time, we have completed over 20 projects in generative AI: from RAG systems for legal departments to custom support agents. Where exactly your task falls depends on data, latency requirements, budget, and how critical confidentiality is.

A typical situation: the client has already tried ChatGPT, but results are unstable — sometimes accurate, sometimes hallucinating. Or they need integration into a corporate portal while complying with security policies. Let's break down each layer of the stack in detail — from RAG to production deployment.

Why Do RAG Systems Break and How to Fix It?

RAG (Retrieval-Augmented Generation) looks simple: find relevant documents, put them in context, get an answer. In practice, it fails in several places.

Chunking without overlap. Classic mistake: chunk_size=512, overlap=0. If the answer lies across two chunks, retrieval won't find either with sufficient confidence. Solution: overlap 15–25% of chunk_size, or better yet, sentence-aware splitting with spaCy or NLTK instead of naive character splitting.

Poor embedder. text-embedding-ada-002 is good for general use, but on legal or medical texts, specialized models like E5-large-v2, BGE-M3, or fine-tuned sentence-transformers on domain data outperform it. Recall@5 differences can be 15–25%.

No re-ranking. Vector search optimizes for speed, not relevance. A cross-encoder re-ranker (ms-marco-MiniLM-L-6-v2, bge-reranker-large) after initial retrieval improves top-3 accuracy with acceptable latency (+50–150ms). This is often more impactful than improving the embedding model.

Hybrid search. Dense vectors alone work poorly on exact queries: names, SKUs, codes. BM25 (sparse) finds exact matches but misses semantics. Hybrid via RRF (Reciprocal Rank Fusion) is the optimal compromise. Qdrant, Weaviate, and pgvector 0.7+ support hybrid search natively.

Typical production architecture for a corporate knowledge base
  1. Documents → preprocessing (PyMuPDF, Unstructured)
  2. Chunking → embedding (BGE-M3)
  3. Qdrant (hybrid dense+sparse)
  4. Cross-encoder re-ranking
  5. Context → LLM (vLLM or OpenAI API)
  6. Answer with sources (RAGAS for quality evaluation)

When to Fine-Tune Instead of Prompt Engineering?

Prompt engineering solves ~70% of LLM adaptation tasks for a domain. The remaining 30% require fine-tuning. Three indicators: the model ignores a specific output format even with detailed prompting; the task requires deep knowledge of specialized vocabulary (medicine, law); you need to significantly reduce token costs by replacing a large model with a smaller specialized one.

LoRA and QLoRA are the standard for SFT. LoRA adds trainable low-rank matrices to attention layers. A typical configuration for Llama-3 8B: r=64, lora_alpha=128, target_modules=["q_proj","v_proj","k_proj","o_proj"] yields ~0.8% trainable parameters, training on one A100 40GB. QLoRA adds 4-bit quantization (NF4) and allows fine-tuning 70B models on two A100 40GB, though speed drops by half compared to bf16.

DPO instead of RLHF. Direct Preference Optimization requires only (chosen, rejected) pairs, not scalar reward signals. DPOTrainer from the trl library (Hugging Face) implements it in a few dozen lines.

Common mistake. A dataset of 500 examples, 5 epochs, validation loss 0.8 — seems fine. But on test, the model degrades on general instructions. Cause: catastrophic forgetting. Solution: add 10–20% general instruction-following examples (Alpaca, FLAN) to the training set to preserve original capabilities.

How to Choose a Base Model: 8B or 70B?

Model Parameters Strengths Context
Llama-3.1 8B 8B Quality/speed balance 128k
Llama-3.1 70B 70B Complex reasoning 128k
Mistral 7B / Mixtral 8x7B 7B / 47B Efficiency for size 32k
Qwen2.5 72B 72B Code, multilingual 128k
Gemma 2 27B 27B Open license 8k

For most tasks, fine-tuning an 8B model is sufficient. 70B is needed when deep reasoning is required or the 8B baseline does not reach the required quality even after fine-tuning. Inference cost for Llama-3 8B via vLLM on A100 is efficient; the exact cost depends on volume.

What Does PagedAttention Bring to Production?

vLLM is the first choice for serving open-source models. PagedAttention is the key technical innovation: KV-cache is managed like virtual memory in an OS, without fragmentation. This yields 2–4x higher throughput compared to naive HuggingFace Transformers inference. The vLLM documentation confirms that continuous batching and PagedAttention are the standard for high-load LLM services.

Typical numbers on A100 80GB for Llama-3 8B (bf16): 400–600 req/s, P50 latency 200–400ms, P99 latency 600–900ms at concurrency 64. For 70B on two A100 with tensor parallelism: 80–120 req/s, P99 latency 1.5–2.5s. AWQ or GPTQ quantization reduces memory consumption by 2x with quality loss within 1–3%.

Multi-Agent Systems

Agents are LLMs with access to tools: search, code execution, API calls, database interaction. Common patterns:

  • ReAct (Reason + Act): the model reasons → chooses a tool → observes the result → reasons again. LangChain and LlamaIndex implement it out of the box.
  • Multi-agent orchestration: multiple specialized agents with a coordinator on top. Example: coordinator → researcher (search + summarization) → coder (code generation and execution) → critic (verification). Tools: AutoGen (Microsoft), CrewAI, custom implementation on LangGraph.

In production, agent systems are non-deterministic. Essential: guardrails, step limits, logging of each step, human-in-the-loop for critical actions.

How We Work: Stages, Timeline, Deliverables

Stage Duration What You Get
Audit and data collection 1–2 weeks Eval dataset of 100+ examples, task formalization
Baseline (prompt + RAG) 1–2 weeks Working prototype, quality metrics
Fine-tuning (if needed) 2–4 weeks Trained model, LoRA weights, model card
Deployment and monitoring 1–2 weeks vLLM server, Grafana + Prometheus
Documentation and training 1 week API documentation, team training

What Is Included

We deliver:

  • Technical documentation (model card, configs, deployment instructions)
  • Access to infrastructure (code repository, trained weights)
  • 1 month of post-deployment support (consultations, bug fixes)
  • Customer team training (2–3 sessions on system operation)

Timeline: basic RAG prototype — 1–2 weeks. Fine-tuning with customer data — 3–6 weeks (including data preparation). Production system with monitoring and retraining — 2–4 months. Cost is calculated individually based on data volume, model complexity, and infrastructure requirements.

We guarantee the quality of the final model with performance benchmarks and ongoing monitoring. Our engineers have hands‑on experience with dozens of production LLM systems.

Want to evaluate your project? Leave a request — we will prepare a preliminary summary within 1–2 business days. Or get a consultation on choosing the approach: RAG, fine-tuning, or hybrid — we will tell you what works best for you. Contact us to discuss your LLM development needs. Schedule a free consultation today.