Developing an AI Agent Trained on a Departing Employee's Data

We design and deploy artificial intelligence systems: from prototype to production-ready solutions. Our team combines expertise in machine learning, data engineering and MLOps to make AI work not in the lab, but in real business.
Showing 1 of 1All 1564 services
Developing an AI Agent Trained on a Departing Employee's Data
Complex
from 2 weeks to 3 months
Frequently Asked Questions

AI Development Areas

AI Solution Development Stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1357
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1250
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    956
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1188
  • image_logo-advance_0.webp
    B2B Advance company logo design
    646
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    929

A Key Employee Leaves — Expertise Leaves with Them

The knowledge base remains, but finding answers is a problem. A standard LLM doesn't know your terminology and processes. An AI agent trained on that employee's data can answer as they would. We develop such agents turnkey: from data collection to production.

Why a Standard LLM Falls Short

A general model gives generic answers from the internet. It doesn't know that in your company "contract approval" goes through three levels. It doesn't know Confluence has a report template and Jira has a history of similar tasks. Accuracy on corporate processes without fine-tuning is about 67%. An agent trained on your data raises it to 91%.

What Problems We Solve

  • Loss of expertise when an employee leaves. Their unique knowledge of processes, decisions, and contacts stays in their heads. An AI agent captures it in the model and RAG index.
  • Unstructured data. 80% of corporate knowledge is in Confluence, Jira, email. We collect, clean, and index it into a vector database (Qdrant, pgvector).
  • Slow answer retrieval. Instead of 15 minutes searching Confluence — one query to the agent. Support ticket reduction by 34%.

Which AI Agent Training Approach to Choose?

Aspect RAG Fine-tuning Hybrid (our choice)
Time to deploy 1–2 weeks 4–6 weeks 8–13 weeks
Terminology accuracy low (43%) high (97%) high (97%)
Knowledge freshness current (index) frozen at cut-off date current (RAG + fine-tune)
Data requirements none ~thousands of examples ~thousands of examples + documents
GPU costs none calculated individually calculated individually

The hybrid approach is the only one that ensures both style accuracy and freshness. We use it in all production projects. Learn more about RAG on Wikipedia.

How We Collect and Prepare Data

Data sources: Confluence (pages), Jira (resolved tickets), email correspondence (anonymized), corporate files (PDF, DOCX).

Preparation process:

  1. Data collection via API (Confluence REST, Jira API, IMAP for email).
  2. Cleaning: html-to-text, deduplication, quality filtering (remove answers shorter than 50 tokens).
  3. Synthetic Q&A generation from documents — we use GPT-4o-mini to create up to 10 pairs per document.
  4. Format conversion: OpenAI messages format (system/user/assistant).

Example data collector code:

from pathlib import Path
from typing import Generator
import json

class CorporateDataCollector:
    """Collects data from corporate sources"""

    async def collect_from_confluence(self, space_keys: list[str]) -> list[dict]:
        """Confluence pages"""
        docs = []
        for space in space_keys:
            pages = await confluence_client.get_all_pages(space)
            for page in pages:
                content = await confluence_client.get_page_content(page["id"])
                docs.append({
                    "source": "confluence",
                    "id": page["id"],
                    "title": page["title"],
                    "content": html_to_text(content),
                    "updated_at": page["version"]["when"],
                    "labels": page.get("labels", []),
                    "space": space,
                })
        return docs

    async def collect_from_email_threads(
        self,
        email_accounts: list[str],
        filter_subjects: list[str] = None,
        anonymize_pii: bool = True,
    ) -> list[dict]:
        """Email threads as conversation training data"""
        threads = []
        for account in email_accounts:
            emails = await gmail_client.get_threads(account, filter_subjects)
            for thread in emails:
                if len(thread["messages"]) >= 2:
                    # Convert thread to dialog format
                    dialog = self.format_as_dialog(thread["messages"])
                    if anonymize_pii:
                        dialog = await self.anonymize_pii(dialog)
                    threads.append(dialog)
        return threads

    async def collect_from_tickets(
        self,
        jira_project: str,
        status: str = "Done",
        limit: int = 5000,
    ) -> list[dict]:
        """Resolved tickets as Q&A pairs"""
        tickets = await jira_client.get_issues(
            jql=f"project={jira_project} AND status={status}",
            fields=["summary", "description", "comments", "resolution"],
            limit=limit,
        )

        qa_pairs = []
        for ticket in tickets:
            if ticket.get("comments"):
                qa_pairs.append({
                    "question": f"{ticket['summary']}\n{ticket.get('description', '')[:500]}",
                    "answer": self.extract_resolution(ticket),
                    "source": "jira",
                    "ticket_id": ticket["id"],
                })

        return qa_pairs
class FinetuningDatasetBuilder:

    async def build_instruction_dataset(
        self,
        raw_docs: list[dict],
        qa_pairs: list[dict],
        target_format: str = "openai",  # "openai", "alpaca", "sharegpt"
    ) -> list[dict]:

        dataset = []

        # From documents — generate Q&A via LLM
        for doc in raw_docs:
            qa_from_doc = await self.generate_qa_from_document(doc["content"])
            for qa in qa_from_doc:
                if target_format == "openai":
                    dataset.append({
                        "messages": [
                            {"role": "system", "content": "You are a corporate assistant for the company. Answer employee questions."},
                            {"role": "user", "content": qa["question"]},
                            {"role": "assistant", "content": qa["answer"]},
                        ]
                    })

        # From tickets — ready pairs
        for qa in qa_pairs:
            if target_format == "openai":
                dataset.append({
                    "messages": [
                        {"role": "system", "content": "You are a technical support assistant."},
                        {"role": "user", "content": qa["question"]},
                        {"role": "assistant", "content": qa["answer"]},
                    ]
                })

        # Deduplication and filtering
        dataset = self.deduplicate(dataset)
        dataset = self.filter_quality(dataset, min_answer_length=50)

        return dataset

    async def generate_qa_from_document(self, document_text: str) -> list[dict]:
        """Generate Q&A pairs from a document"""
        response = await openai_client.chat.completions.create(
            model="gpt-4o-mini",
            messages=[{
                "role": "user",
                "content": f"""Create 5-10 questions and answers from the following document.
Questions should be as real employees would ask.
Answers should be complete and accurate.

Document:
{document_text[:3000]}

Return JSON: [{{"question": "...", "answer": "..."}}]"""
            }],
        )
        return json.loads(response.choices[0].message.content)

    def filter_quality(self, dataset: list[dict], min_answer_length: int) -> list[dict]:
        """Filter low-quality data"""
        filtered = []
        for item in dataset:
            messages = item.get("messages", [])
            assistant_msg = next((m for m in messages if m["role"] == "assistant"), None)
            if assistant_msg and len(assistant_msg["content"]) >= min_answer_length:
                filtered.append(item)
        return filtered

How the Hybrid Architecture Works: fine-tune + RAG

The fine-tuned model knows the company's style and terminology. RAG adds current documents. We combine them in one agent:

from sentence_transformers import SentenceTransformer
from openai import OpenAI
from qdrant_client import QdrantClient

class HybridCorporateAgent:
    """Combines a fine-tuned model with company style and RAG with current knowledge"""

    def __init__(self):
        # Fine-tuned model knows style and terminology
        self.finetuned_client = OpenAI(base_url="http://vllm-server:8000/v1")
        self.finetuned_model = "company-assistant-ft-v2"

        # RAG for current documents
        self.embed_model = SentenceTransformer("BAAI/bge-m3")
        self.vector_db = QdrantClient(host="qdrant-server")

    async def answer(self, question: str, user_context: dict = None) -> dict:
        # Step 1: Retrieve relevant documents
        query_embedding = self.embed_model.encode(question)
        relevant_docs = self.vector_db.search(
            collection_name="corporate_docs",
            query_vector=query_embedding,
            limit=5,
            score_threshold=0.6,
            query_filter=self.build_access_filter(user_context),  # Access permissions
        )

        # Step 2: Build context
        context = "\n\n".join([
            f"[{doc.payload['title']}]: {doc.payload['content']}"
            for doc in relevant_docs
        ])

        # Step 3: Respond with fine-tuned model and RAG context
        response = self.finetuned_client.chat.completions.create(
            model=self.finetuned_model,
            messages=[{
                "role": "system",
                "content": f"You are a corporate assistant. Use the documents as the source of truth.\n\nDocuments:\n{context}"
            }, {
                "role": "user",
                "content": question,
            }],
            temperature=0.1,
        )

        return {
            "answer": response.choices[0].message.content,
            "sources": [{"title": d.payload["title"], "score": d.score} for d in relevant_docs],
        }

    def build_access_filter(self, user_context: dict):
        """Access permission filtering — employee sees only their documents"""
        if not user_context:
            return None

        department = user_context.get("department", "all")
        clearance = user_context.get("clearance", "public")

        return {
            "must": [
                {"key": "access_level", "match": {"any": [clearance, "public"]}},
                {"key": "departments", "match": {"any": [department, "all"]}},
            ]
        }
Detailed architecture example

The agent uses vLLM for fine-tuned model inference, Qdrant for vector search, and ONNX Runtime for embeddings. All components are deployed in Kubernetes with autoscaling based on GPU utilization.

Case Study: IT Company, 300 Employees

A client — an IT company with 300 employees — had a senior developer leaving who managed key processes. Over 5 years he accumulated 8,000 Confluence pages and participated in solving 12,000 Jira tickets. We collected and cleaned data in 3 weeks, generated 45,000 synthetic Q&A pairs, fine-tuned GPT-4o-mini on 60,000 examples (3 epochs), and loaded all documents into Qdrant. Results: accuracy on corporate processes 91% vs. 67% for base GPT-4o, correct terminology 97% vs. 43%, support ticket reduction by 34%.

What's Included in the Work

Stage Duration What You Get
Analytics and data collection 2–4 weeks Source map, cleaned dataset
Synthetic Q&A generation 1–2 weeks 10k-100k question-answer pairs
Fine-tuning and RAG indexing 2–3 weeks Trained model, vector index
Testing and calibration 2 weeks Report with metrics (accuracy, hallucination rate)
Deployment and documentation 1 week API access, monitoring dashboard, admin interface

How Long Does Development Take?

From 8 to 13 weeks depending on data volume and quality. Most time goes to collection and cleaning — you can't speed it up without losing quality. Cost is determined individually based on the number of sources and required accuracy.

Typical Mistakes When Training an AI Agent

  • Ignoring access rights. Without filtering, an employee may see documents from another department. Our architecture includes build_access_filter that checks department and clearance level.
  • Using only RAG. Without fine-tuning, the model doesn't learn the corporate style — answers sound like from the internet, not a colleague.
  • Weak quality filtering. If bad answers (shorter than 50 tokens or without facts) get into the dataset, quality drops. We filter by minimum length and semantic similarity.
  • No testing on real questions. Synthetic tests don't reveal gaps. We test on 500+ questions from future users.

Why Order Development from Us?

5 years of experience in NLP and MLOps, 30+ deployed AI agents. We guarantee the agent will answer with at least 85% accuracy on specified processes. We use only proven tools: Hugging Face, Qdrant, vLLM, ONNX Runtime. After delivery, we provide support and retraining as new data accumulates.

Get a consultation — we will analyze your data and suggest the optimal approach. Turnkey development from scratch or integration into existing infrastructure. Order AI agent development — preserve your team's expertise.

LLM Development: Fine-Tuning, RAG, Agents, and Production Deployment

Using GPT‑4 or Claude 3.5 Sonnet through a public API is not a solution — it's just a tool. When the requirement is to "make it like ChatGPT, but on our data," there is a real engineering challenge behind it: from prompt engineering to training a 70B model on your own infrastructure. End-to-end LLM solution development is a complex stack, and we have been doing it for over 5 years. During this time, we have completed over 20 projects in generative AI: from RAG systems for legal departments to custom support agents. Where exactly your task falls depends on data, latency requirements, budget, and how critical confidentiality is.

A typical situation: the client has already tried ChatGPT, but results are unstable — sometimes accurate, sometimes hallucinating. Or they need integration into a corporate portal while complying with security policies. Let's break down each layer of the stack in detail — from RAG to production deployment.

Why Do RAG Systems Break and How to Fix It?

RAG (Retrieval-Augmented Generation) looks simple: find relevant documents, put them in context, get an answer. In practice, it fails in several places.

Chunking without overlap. Classic mistake: chunk_size=512, overlap=0. If the answer lies across two chunks, retrieval won't find either with sufficient confidence. Solution: overlap 15–25% of chunk_size, or better yet, sentence-aware splitting with spaCy or NLTK instead of naive character splitting.

Poor embedder. text-embedding-ada-002 is good for general use, but on legal or medical texts, specialized models like E5-large-v2, BGE-M3, or fine-tuned sentence-transformers on domain data outperform it. Recall@5 differences can be 15–25%.

No re-ranking. Vector search optimizes for speed, not relevance. A cross-encoder re-ranker (ms-marco-MiniLM-L-6-v2, bge-reranker-large) after initial retrieval improves top-3 accuracy with acceptable latency (+50–150ms). This is often more impactful than improving the embedding model.

Hybrid search. Dense vectors alone work poorly on exact queries: names, SKUs, codes. BM25 (sparse) finds exact matches but misses semantics. Hybrid via RRF (Reciprocal Rank Fusion) is the optimal compromise. Qdrant, Weaviate, and pgvector 0.7+ support hybrid search natively.

Typical production architecture for a corporate knowledge base
  1. Documents → preprocessing (PyMuPDF, Unstructured)
  2. Chunking → embedding (BGE-M3)
  3. Qdrant (hybrid dense+sparse)
  4. Cross-encoder re-ranking
  5. Context → LLM (vLLM or OpenAI API)
  6. Answer with sources (RAGAS for quality evaluation)

When to Fine-Tune Instead of Prompt Engineering?

Prompt engineering solves ~70% of LLM adaptation tasks for a domain. The remaining 30% require fine-tuning. Three indicators: the model ignores a specific output format even with detailed prompting; the task requires deep knowledge of specialized vocabulary (medicine, law); you need to significantly reduce token costs by replacing a large model with a smaller specialized one.

LoRA and QLoRA are the standard for SFT. LoRA adds trainable low-rank matrices to attention layers. A typical configuration for Llama-3 8B: r=64, lora_alpha=128, target_modules=["q_proj","v_proj","k_proj","o_proj"] yields ~0.8% trainable parameters, training on one A100 40GB. QLoRA adds 4-bit quantization (NF4) and allows fine-tuning 70B models on two A100 40GB, though speed drops by half compared to bf16.

DPO instead of RLHF. Direct Preference Optimization requires only (chosen, rejected) pairs, not scalar reward signals. DPOTrainer from the trl library (Hugging Face) implements it in a few dozen lines.

Common mistake. A dataset of 500 examples, 5 epochs, validation loss 0.8 — seems fine. But on test, the model degrades on general instructions. Cause: catastrophic forgetting. Solution: add 10–20% general instruction-following examples (Alpaca, FLAN) to the training set to preserve original capabilities.

How to Choose a Base Model: 8B or 70B?

Model Parameters Strengths Context
Llama-3.1 8B 8B Quality/speed balance 128k
Llama-3.1 70B 70B Complex reasoning 128k
Mistral 7B / Mixtral 8x7B 7B / 47B Efficiency for size 32k
Qwen2.5 72B 72B Code, multilingual 128k
Gemma 2 27B 27B Open license 8k

For most tasks, fine-tuning an 8B model is sufficient. 70B is needed when deep reasoning is required or the 8B baseline does not reach the required quality even after fine-tuning. Inference cost for Llama-3 8B via vLLM on A100 is efficient; the exact cost depends on volume.

What Does PagedAttention Bring to Production?

vLLM is the first choice for serving open-source models. PagedAttention is the key technical innovation: KV-cache is managed like virtual memory in an OS, without fragmentation. This yields 2–4x higher throughput compared to naive HuggingFace Transformers inference. The vLLM documentation confirms that continuous batching and PagedAttention are the standard for high-load LLM services.

Typical numbers on A100 80GB for Llama-3 8B (bf16): 400–600 req/s, P50 latency 200–400ms, P99 latency 600–900ms at concurrency 64. For 70B on two A100 with tensor parallelism: 80–120 req/s, P99 latency 1.5–2.5s. AWQ or GPTQ quantization reduces memory consumption by 2x with quality loss within 1–3%.

Multi-Agent Systems

Agents are LLMs with access to tools: search, code execution, API calls, database interaction. Common patterns:

  • ReAct (Reason + Act): the model reasons → chooses a tool → observes the result → reasons again. LangChain and LlamaIndex implement it out of the box.
  • Multi-agent orchestration: multiple specialized agents with a coordinator on top. Example: coordinator → researcher (search + summarization) → coder (code generation and execution) → critic (verification). Tools: AutoGen (Microsoft), CrewAI, custom implementation on LangGraph.

In production, agent systems are non-deterministic. Essential: guardrails, step limits, logging of each step, human-in-the-loop for critical actions.

How We Work: Stages, Timeline, Deliverables

Stage Duration What You Get
Audit and data collection 1–2 weeks Eval dataset of 100+ examples, task formalization
Baseline (prompt + RAG) 1–2 weeks Working prototype, quality metrics
Fine-tuning (if needed) 2–4 weeks Trained model, LoRA weights, model card
Deployment and monitoring 1–2 weeks vLLM server, Grafana + Prometheus
Documentation and training 1 week API documentation, team training

What Is Included

We deliver:

  • Technical documentation (model card, configs, deployment instructions)
  • Access to infrastructure (code repository, trained weights)
  • 1 month of post-deployment support (consultations, bug fixes)
  • Customer team training (2–3 sessions on system operation)

Timeline: basic RAG prototype — 1–2 weeks. Fine-tuning with customer data — 3–6 weeks (including data preparation). Production system with monitoring and retraining — 2–4 months. Cost is calculated individually based on data volume, model complexity, and infrastructure requirements.

We guarantee the quality of the final model with performance benchmarks and ongoing monitoring. Our engineers have hands‑on experience with dozens of production LLM systems.

Want to evaluate your project? Leave a request — we will prepare a preliminary summary within 1–2 business days. Or get a consultation on choosing the approach: RAG, fine-tuning, or hybrid — we will tell you what works best for you. Contact us to discuss your LLM development needs. Schedule a free consultation today.