AI Assistant for Product Documentation on RAG

We design and deploy artificial intelligence systems: from prototype to production-ready solutions. Our team combines expertise in machine learning, data engineering and MLOps to make AI work not in the lab, but in real business.
Showing 1 of 1All 1564 services
AI Assistant for Product Documentation on RAG
Medium
~1-2 weeks
Frequently Asked Questions

AI Development Areas

AI Solution Development Stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1358
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1251
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    957
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1188
  • image_logo-advance_0.webp
    B2B Advance company logo design
    646
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    929

AI Assistant for Product Documentation on RAG

Users spend up to 30 minutes searching for answers in scattered documentation — opening dozens of pages but not finding what they need. Support teams drown in repetitive questions like 'how to configure X' and 'where to find Y'. We build an AI assistant that extracts precise answers from the docs site and delivers them in chat. No hallucinations: every response is backed by a citation from the documentation.

According to a Forrester study, the average employee spends 22% of their time searching for internal company information. For product documentation — with 300+ pages and multiple versions — this figure reaches 30%. An assistant based on RAG (Retrieval-Augmented Generation) reduces search time to seconds and support load by up to 60%.

What Problems Does the AI Assistant Solve?

Information overload. Product documentation may contain 300+ pages, multiple versions, and languages. The user doesn't know which section to open. The assistant finds the relevant fragment in seconds and presents it in the answer.

Outdated answers. When documentation updates, search indices become stale. Our assistant works on a fresh vector database — reindexing runs automatically on every CI/CD deploy.

Support load. Based on our data, after implementing the assistant, the number of tickets with 'how-to' questions drops by 50–60%. This frees up support engineers for complex tasks.

RAG Mechanism: How to Eliminate Hallucinations

The key component is Retrieval-Augmented Generation (RAG). We don't let the LLM answer from its memory; instead, we supply it with context from the documentation. The vector database (ChromaDB, Qdrant, or pgvector) stores text chunks with metadata: version, title, URL. On a query, semantic search retrieves up to 5 most relevant chunks, and the model formulates the answer strictly based on this data. This approach is 5 times more efficient than standard keyword search and almost completely eliminates hallucinations.

from anthropic import Anthropic
from langchain_openai import OpenAIEmbeddings
from langchain_community.vectorstores import Chroma
import json
from typing import Optional

client = Anthropic()
embeddings_model = OpenAIEmbeddings(model="text-embedding-3-small")

class DocAssistant:

    def __init__(self, product_name: str, db_path: str):
        self.product_name = product_name
        self.vectorstore = Chroma(
            collection_name=f"docs_{product_name}",
            embedding_function=embeddings_model,
            persist_directory=db_path,
        )

    def answer(
        self,
        question: str,
        product_version: Optional[str] = None,
        conversation_id: Optional[str] = None,
    ) -> dict:
        """Answers a question based on documentation"""

        # Filter by version if specified
        where_filter = {"version": product_version} if product_version else None

        results = self.vectorstore.similarity_search_with_score(
            question, k=5, filter=where_filter
        )

        if not results:
            return {
                "answer": f"No information found in {self.product_name} documentation for your question.",
                "sources": [],
                "confidence": "low",
                "suggest_support": True,
            }

        context = "\n\n".join([
            f"[{doc.metadata.get('title', 'Document')}, {doc.metadata.get('section', '')}]:\n{doc.page_content}"
            for doc, _ in results[:4]
        ])

        response = client.messages.create(
            model="claude-sonnet-4-5",
            max_tokens=2048,
            system=f"""You are a support specialist for {self.product_name}.

STRICT RULES:
1. Answer ONLY based on the provided documentation
2. Cite specific sections when necessary
3. If the answer is not in the documentation, say "This information is not described in the documentation"
4. Do not invent functionality
5. For steps, use numbered lists
6. Always end with: "Need further help? Contact support: [email protected]" """,
            messages=[{
                "role": "user",
                "content": f"""Question: {question}
{f"Product version: {product_version}" if product_version else ""}

Documentation:
{context}"""
            }]
        )

        answer_text = response.content[0].text

        # Determine confidence based on specific citations
        confidence = "high" if any(
            r[1] < 0.3 for r in results[:2]  # Low distance = high similarity
        ) else "medium"

        return {
            "answer": answer_text,
            "sources": [
                {
                    "title": doc.metadata.get("title"),
                    "section": doc.metadata.get("section"),
                    "url": doc.metadata.get("url"),
                    "version": doc.metadata.get("version"),
                }
                for doc, _ in results[:3]
            ],
            "confidence": confidence,
            "suggest_support": confidence == "low",
        }

Indexing Documentation from Different Sources

We support import from GitBook, Confluence, local Markdown files, and any static HTML sites. We write an adapter for each source. Example — indexing GitBook via sitemap:

import aiohttp
from bs4 import BeautifulSoup
from langchain.text_splitter import MarkdownHeaderTextSplitter

class DocIndexer:

    def __init__(self, vectorstore: Chroma):
        self.vectorstore = vectorstore
        self.md_splitter = MarkdownHeaderTextSplitter(
            headers_to_split_on=[("##", "section"), ("###", "subsection")]
        )

    async def index_gitbook(self, base_url: str, version: str = "latest"):
        """Index GitBook documentation"""
        async with aiohttp.ClientSession() as session:
            # Get sitemap
            async with session.get(f"{base_url}/sitemap.xml") as resp:
                sitemap = await resp.text()

            import re
            urls = re.findall(r'<loc>(.*?)</loc>', sitemap)

            for url in urls[:100]:  # Limit
                async with session.get(url) as page_resp:
                    html = await page_resp.text()

                soup = BeautifulSoup(html, "html.parser")
                title = soup.find("h1")
                content = soup.find("article") or soup.find("main")

                if not content:
                    continue

                text = content.get_text(separator="\n", strip=True)
                chunks = self.md_splitter.split_text(text)

                self.vectorstore.add_texts(
                    texts=[c.page_content for c in chunks],
                    metadatas=[{
                        "title": title.get_text() if title else "Unknown",
                        "url": url,
                        "version": version,
                        "section": c.metadata.get("section", ""),
                    } for c in chunks]
                )

    def index_markdown_files(self, docs_dir: str, version: str = "latest"):
        """Index local .md documentation files"""
        for md_file in Path(docs_dir).rglob("*.md"):
            content = md_file.read_text()
            chunks = self.md_splitter.split_text(content)

            # Extract title from first H1 line
            title = md_file.stem.replace("-", " ").title()
            for line in content.splitlines():
                if line.startswith("# "):
                    title = line[2:].strip()
                    break

            self.vectorstore.add_texts(
                texts=[c.page_content for c in chunks],
                metadatas=[{
                    "title": title,
                    "file": str(md_file.relative_to(docs_dir)),
                    "version": version,
                    "section": c.metadata.get("section", ""),
                } for c in chunks]
            )

Widget for the Docs Site

The user should be able to ask a question without leaving the documentation page. We embed a chat widget that connects to the assistant's API. Here's a minimal implementation in pure JavaScript:

// docs-chat-widget.js
class DocsChatWidget {
    constructor(config) {
        this.apiUrl = config.apiUrl;
        this.productVersion = config.version || 'latest';
        this.container = this.createWidget();
        document.body.appendChild(this.container);
    }

    createWidget() {
        const container = document.createElement('div');
        container.innerHTML = `
            <div id="docs-chat-btn" style="position:fixed;bottom:24px;right:24px;cursor:pointer;
                background:#5865F2;color:white;padding:12px 20px;border-radius:24px;
                box-shadow:0 4px 12px rgba(0,0,0,0.2);">
                💬 Ask AI
            </div>
            <div id="docs-chat-panel" style="display:none;position:fixed;bottom:80px;right:24px;
                width:380px;height:520px;background:white;border-radius:12px;
                box-shadow:0 8px 32px rgba(0,0,0,0.15);overflow:hidden;">
                <div style="padding:16px;background:#5865F2;color:white;">
                    <strong>AI Documentation</strong>
                    <span onclick="this.closest('#docs-chat-panel').style.display='none'"
                          style="float:right;cursor:pointer">✕</span>
                </div>
                <div id="chat-messages" style="height:380px;overflow-y:auto;padding:16px;"></div>
                <div style="padding:12px;border-top:1px solid #eee;display:flex;gap:8px;">
                    <input id="chat-input" type="text" placeholder="Ask a question..."
                           style="flex:1;padding:8px;border:1px solid #ddd;border-radius:6px;">
                    <button onclick="window.docsChat.send()" style="padding:8px 16px;
                        background:#5865F2;color:white;border:none;border-radius:6px;cursor:pointer;">→</button>
                </div>
            </div>
        `;
        return container;
    }

    async send() {
        const input = document.getElementById('chat-input');
        const question = input.value.trim();
        if (!question) return;

        input.value = '';
        this.addMessage('user', question);

        const response = await fetch(this.apiUrl + '/ask', {
            method: 'POST',
            headers: { 'Content-Type': 'application/json' },
            body: JSON.stringify({ question, version: this.productVersion })
        });

        const data = await response.json();
        this.addMessage('assistant', data.answer, data.sources);
    }
}

window.docsChat = new DocsChatWidget({
    apiUrl: 'https://api.myproduct.com/docs-ai',
    version: document.querySelector('meta[name="docs-version"]')?.content
});

Why Is Response Versioning Critical?

If the product is actively developed, a user on an older version will get an incorrect answer if the assistant relies on the latest documentation. We store a version meta-field for each chunk and filter the search by the version the client provides (or determine it via user-agent). Answer accuracy increases to 97% versus 82% without filtering.

What's Included?

Component Description Duration
Docs indexing Parse all documentation pages, split into chunks, generate embeddings, load into vector DB 2–3 days
RAG backend FastAPI service with LangChain, integration with LLM (Claude, GPT-4), version filtering, error handling 3–5 days
Site widget Ready JS widget with customizable styles, dark theme support, analytics 2–3 days
Helpdesk integration Escalation to Zendesk / Freshdesk / Intercom, transfer of dialog history 1–2 weeks
Documentation and training Instructions for content updates, deploy new version, metrics dashboard 2 days

Comparison of Approaches to Building an Assistant

Approach Accuracy Time to Implement Hallucinations
RAG (ours) 95–97% 1–2 weeks Almost none
Fine-tuning LLM 80–85% 2–4 weeks Possible
Pure LLM without context 70–75% Low Frequent

The RAG approach offers the best combination of accuracy and implementation speed, and most importantly, it almost completely eliminates hallucinations because the model relies on real documents.

Practical Case: SaaS Product with 8,000 Users

Documentation: 320 GitBook pages, 5 product versions. Our client — a project management SaaS platform — faced 40% of tickets being basic questions. We implemented the AI assistant in 10 days. Results:

  • Support tickets like 'how to configure X' dropped by 58%.
  • TTFR (time to first response) decreased from 4 hours to 2 seconds — users get instant answers.
  • Documentation satisfaction (CSAT) rose from 3.2 to 4.4 out of 5.
  • Support savings: the client estimated that the ticket reduction saved $10,000 per month by reducing the number of operators.

How Does the Assistant Interact with a Live Support Agent?

If the assistant is not confident in its answer (confidence "low"), it suggests the user contact support. The escalation button transfers the dialog history to the helpdesk (Zendesk, Freshdesk, Intercom). The agent sees the context: the user's question, the found documentation fragments, and the generated answer. This eliminates repeated questioning and speeds up resolution.

Timeline and Pricing

Estimated timeline — from 3 days to 2 weeks depending on complexity. Pricing is calculated individually — contact us to discuss your case. We guarantee: the assistant will not hallucinate, supports versioning, and is easy to update.

Contact us — get an architecture consultation and a demo for your docs site. It can be launched in a week, and within a month you can measure the reduction in support load.

LLM Development: Fine-Tuning, RAG, Agents, and Production Deployment

Using GPT‑4 or Claude 3.5 Sonnet through a public API is not a solution — it's just a tool. When the requirement is to "make it like ChatGPT, but on our data," there is a real engineering challenge behind it: from prompt engineering to training a 70B model on your own infrastructure. End-to-end LLM solution development is a complex stack, and we have been doing it for over 5 years. During this time, we have completed over 20 projects in generative AI: from RAG systems for legal departments to custom support agents. Where exactly your task falls depends on data, latency requirements, budget, and how critical confidentiality is.

A typical situation: the client has already tried ChatGPT, but results are unstable — sometimes accurate, sometimes hallucinating. Or they need integration into a corporate portal while complying with security policies. Let's break down each layer of the stack in detail — from RAG to production deployment.

Why Do RAG Systems Break and How to Fix It?

RAG (Retrieval-Augmented Generation) looks simple: find relevant documents, put them in context, get an answer. In practice, it fails in several places.

Chunking without overlap. Classic mistake: chunk_size=512, overlap=0. If the answer lies across two chunks, retrieval won't find either with sufficient confidence. Solution: overlap 15–25% of chunk_size, or better yet, sentence-aware splitting with spaCy or NLTK instead of naive character splitting.

Poor embedder. text-embedding-ada-002 is good for general use, but on legal or medical texts, specialized models like E5-large-v2, BGE-M3, or fine-tuned sentence-transformers on domain data outperform it. Recall@5 differences can be 15–25%.

No re-ranking. Vector search optimizes for speed, not relevance. A cross-encoder re-ranker (ms-marco-MiniLM-L-6-v2, bge-reranker-large) after initial retrieval improves top-3 accuracy with acceptable latency (+50–150ms). This is often more impactful than improving the embedding model.

Hybrid search. Dense vectors alone work poorly on exact queries: names, SKUs, codes. BM25 (sparse) finds exact matches but misses semantics. Hybrid via RRF (Reciprocal Rank Fusion) is the optimal compromise. Qdrant, Weaviate, and pgvector 0.7+ support hybrid search natively.

Typical production architecture for a corporate knowledge base
  1. Documents → preprocessing (PyMuPDF, Unstructured)
  2. Chunking → embedding (BGE-M3)
  3. Qdrant (hybrid dense+sparse)
  4. Cross-encoder re-ranking
  5. Context → LLM (vLLM or OpenAI API)
  6. Answer with sources (RAGAS for quality evaluation)

When to Fine-Tune Instead of Prompt Engineering?

Prompt engineering solves ~70% of LLM adaptation tasks for a domain. The remaining 30% require fine-tuning. Three indicators: the model ignores a specific output format even with detailed prompting; the task requires deep knowledge of specialized vocabulary (medicine, law); you need to significantly reduce token costs by replacing a large model with a smaller specialized one.

LoRA and QLoRA are the standard for SFT. LoRA adds trainable low-rank matrices to attention layers. A typical configuration for Llama-3 8B: r=64, lora_alpha=128, target_modules=["q_proj","v_proj","k_proj","o_proj"] yields ~0.8% trainable parameters, training on one A100 40GB. QLoRA adds 4-bit quantization (NF4) and allows fine-tuning 70B models on two A100 40GB, though speed drops by half compared to bf16.

DPO instead of RLHF. Direct Preference Optimization requires only (chosen, rejected) pairs, not scalar reward signals. DPOTrainer from the trl library (Hugging Face) implements it in a few dozen lines.

Common mistake. A dataset of 500 examples, 5 epochs, validation loss 0.8 — seems fine. But on test, the model degrades on general instructions. Cause: catastrophic forgetting. Solution: add 10–20% general instruction-following examples (Alpaca, FLAN) to the training set to preserve original capabilities.

How to Choose a Base Model: 8B or 70B?

Model Parameters Strengths Context
Llama-3.1 8B 8B Quality/speed balance 128k
Llama-3.1 70B 70B Complex reasoning 128k
Mistral 7B / Mixtral 8x7B 7B / 47B Efficiency for size 32k
Qwen2.5 72B 72B Code, multilingual 128k
Gemma 2 27B 27B Open license 8k

For most tasks, fine-tuning an 8B model is sufficient. 70B is needed when deep reasoning is required or the 8B baseline does not reach the required quality even after fine-tuning. Inference cost for Llama-3 8B via vLLM on A100 is efficient; the exact cost depends on volume.

What Does PagedAttention Bring to Production?

vLLM is the first choice for serving open-source models. PagedAttention is the key technical innovation: KV-cache is managed like virtual memory in an OS, without fragmentation. This yields 2–4x higher throughput compared to naive HuggingFace Transformers inference. The vLLM documentation confirms that continuous batching and PagedAttention are the standard for high-load LLM services.

Typical numbers on A100 80GB for Llama-3 8B (bf16): 400–600 req/s, P50 latency 200–400ms, P99 latency 600–900ms at concurrency 64. For 70B on two A100 with tensor parallelism: 80–120 req/s, P99 latency 1.5–2.5s. AWQ or GPTQ quantization reduces memory consumption by 2x with quality loss within 1–3%.

Multi-Agent Systems

Agents are LLMs with access to tools: search, code execution, API calls, database interaction. Common patterns:

  • ReAct (Reason + Act): the model reasons → chooses a tool → observes the result → reasons again. LangChain and LlamaIndex implement it out of the box.
  • Multi-agent orchestration: multiple specialized agents with a coordinator on top. Example: coordinator → researcher (search + summarization) → coder (code generation and execution) → critic (verification). Tools: AutoGen (Microsoft), CrewAI, custom implementation on LangGraph.

In production, agent systems are non-deterministic. Essential: guardrails, step limits, logging of each step, human-in-the-loop for critical actions.

How We Work: Stages, Timeline, Deliverables

Stage Duration What You Get
Audit and data collection 1–2 weeks Eval dataset of 100+ examples, task formalization
Baseline (prompt + RAG) 1–2 weeks Working prototype, quality metrics
Fine-tuning (if needed) 2–4 weeks Trained model, LoRA weights, model card
Deployment and monitoring 1–2 weeks vLLM server, Grafana + Prometheus
Documentation and training 1 week API documentation, team training

What Is Included

We deliver:

  • Technical documentation (model card, configs, deployment instructions)
  • Access to infrastructure (code repository, trained weights)
  • 1 month of post-deployment support (consultations, bug fixes)
  • Customer team training (2–3 sessions on system operation)

Timeline: basic RAG prototype — 1–2 weeks. Fine-tuning with customer data — 3–6 weeks (including data preparation). Production system with monitoring and retraining — 2–4 months. Cost is calculated individually based on data volume, model complexity, and infrastructure requirements.

We guarantee the quality of the final model with performance benchmarks and ongoing monitoring. Our engineers have hands‑on experience with dozens of production LLM systems.

Want to evaluate your project? Leave a request — we will prepare a preliminary summary within 1–2 business days. Or get a consultation on choosing the approach: RAG, fine-tuning, or hybrid — we will tell you what works best for you. Contact us to discuss your LLM development needs. Schedule a free consultation today.