How an AI-powered VK bot with RAG solves business problems
A typical VK bot built on handlers and regex breaks when a user writes "I want red for 5000" or "send payment link" — it crashes with an unhandled exception. An LLM with an RAG pipeline solves this: GPT-4 understands intent, ChromaDB returns relevant products, and vkbottle handles events without losses. For over 5 years, we have deployed such bots for 30+ projects — from ticket sales to technical support. Natural Language Processing (NLP) using Natasha and spaCy libraries accurately extracts entities and intents.
How to design an RAG pipeline for a VK bot
We collect the community's dialog history, identify the top 10 intents (order, return, status). We design the RAG pipeline: which documents to index, how to chunk text, which embedding model to use (e.g., intfloat/multilingual-e5-large). We use chunking techniques — fixed size of 512 tokens with 128 overlap. According to VK API documentation, Callback events are transmitted in JSON format.
Stack selection: from pilot to production
The primary VK API is Callback (webhook) with the vkbottle library. For AI — LangChain with OpenAI or Mistral provider. Vector DB — ChromaDB (in-memory for pilot) or Qdrant (for production with millions of vectors). We fine-tune the model via LoRA if answer accuracy is below 85%. For NLP, we use Natasha (NER) and spaCy libraries.
Developing handlers and integrating with VK
We set up a FastAPI server with endpoints for VK Callback, integrate LangChain into handlers. Below is a minimal template:
from vkbottle import Bot, Message
from vkbottle.bot import BotLabeler
bot = Bot(token=VK_TOKEN)
labeler = BotLabeler()
@labeler.message()
async def handle_message(message: Message):
user_response = await ai_handler.process(message.text, user_id=message.from_id)
await message.answer(user_response)
bot.labeler.load(labeler)
bot.run_forever()
Testing and deployment
We target p95 latency < 2 sec, load testing up to 100 RPS. Deploy on Kubernetes with autoscaling based on GPU utilization. After launch, we monitor with W&B, fine-tune prompts based on user feedback. Automatic A/B test on 10% of traffic allows comparing accuracy and satisfaction metrics.
Why stack choice affects total cost of ownership?
Compare two approaches:
| Component |
Budget (initial) |
Production (scalable) |
| Model |
GPT-4o mini (cheaper) |
Fine-tuned Mistral on custom data |
| Vector DB |
ChromaDB (in-process) |
Qdrant with sharding |
| Infrastructure |
1 server 16 vCPU, 64 GB RAM |
Kubernetes + GPU nodes (T4) |
| RAG pipeline |
LangChain default |
LangChain + self-querying retriever |
For up to 10,000 users per month, the budget option suffices. When growing to 100,000, the production stack pays off due to lower latency and fewer escalations. For example, fine-tuned Mistral is 40% more accurate than base GPT-4o mini on specialized queries, and Qdrant is 3x faster than ChromaDB at 100k vectors.
Metrics comparison before and after implementation
Real results from one project (financial consulting)
| Metric |
Before |
After RAG bot |
Change |
| Operator response time |
15 min |
5 sec |
-97% |
| Escalations to human |
100% |
12% |
-88% |
| Conversion to lead |
18% |
42% |
+133% |
What's included in the turnkey solution
- Architecture documentation (ER diagrams, RAG sequence diagrams)
- Repository with CI/CD (GitHub Actions, Docker, Helm)
- VK community setup: Callback API, keyboards, carousels
- Operator training: how to change prompts, add new documents to knowledge base
- 3-month warranty: bug fixes, model fine-tuning
Contact us — we'll estimate your project in 1 day. We'll send a demo bot with your data. Get a consultation on your scenario — we'll tell you if an AI chatbot is right for your business.
What does VK Mini Apps integration bring?
If you need an interface with an order form, appointment calendar, or payment — we use VK Mini Apps (React + VK Bridge). The bot invokes the Mini App via a button, the Mini App returns the result via postMessage. We did this for a dental clinic: GPT answers questions, Mini App shows available slots and accepts VK Pay payments.
from vkbottle_types.objects import MessagesKeyboard, MessagesKeyboardButton
keyboard = MessagesKeyboard(
one_time=True,
buttons=[[
{"action": {"type": "text", "label": "Order"}},
{"action": {"type": "text", "label": "Information"}},
]]
)
Marketing campaigns: legal and effective
VK allows sending messages only to those who first wrote to the bot or subscribed to the newsletter. Open rate for such messages is 30-50% (higher than email). We use interest-based segmentation: for a user who asked about a specific product, we send a personalized offer after 3 days via Messaging API. No spam — every dialog with consent.
10+ years of experience in VK bot development, certified AI engineers. Get a consultation.
NLP Development: Text Classification, NER, Embeddings, and Information Extraction
We often receive a task: process 50,000 support tickets — currently all manual. Dataset — 3,000 labeled examples, 12 categories, imbalance: one category occupies 40% of the sample, three at 1-2% each. Baseline accuracy — 78%. Sounds decent until you look at recall for rare classes: 0.31, 0.44, 0.28. These classes — complaints and churn threats — are most important to the business.
This is a typical NLP development project. The problem is not the algorithm but that accuracy is the wrong metric. Our experience across 30+ projects shows: we start by analyzing business metrics and only then choose the model.
Why accuracy is not the right metric for rare classes?
Accuracy ignores imbalance. If the "churn" class appears in 2% of cases, the model can predict "all good" and get 98% accuracy — but the business loses clients. Solution: F1 macro (averaged over all classes) or weighted F1. For NER — strict entity F1 (exact matches only). We guarantee: after choosing the correct metric, model quality becomes measurable and predictable.
Text Classification: From BERT to Distillation
BERT-like models are the standard for classification. ruBERT-base or ruBERT-large from DeepPavlov for Russian. multilingual-e5-large — for multiple languages in one pipeline. XLM-RoBERTa-large — a strong multilingual backbone.
Fine-tuning for classification: add a classification head on top of the [CLS] token, train for 3-5 epochs with lr=2e-5, weight decay=0.01. For imbalance — weighted CrossEntropyLoss or focal loss with gamma=2.0. Contact us — we will show a code snippet.
Imbalance case study. Dataset — 3,000 examples, imbalance 1:20. Solution: class_weight via sklearn + CrossEntropyLoss. Additionally — augmentation of rare classes via backtranslation (ru→en→ru through MarianMT). Recall for rare classes rose from 0.31 to 0.67 with a slight drop in accuracy (76%→74%). Full NLP development end-to-end took 3 weeks.
Distillation for production. BERT-large gives F1 0.89, but inference on CPU — 180ms. Distillation into DistilBERT or ruBERT-tiny2 reduces latency to 25ms with F1 0.84. Export to ONNX Runtime provides an additional 1.5-2x speedup. DistilBERT achieves 7x lower latency than BERT-large with only a 5% drop in macro F1 – a typical production trade-off.
| Model |
F1 macro |
Latency (CPU) |
Size |
| BERT-large |
0.89 |
180 ms |
1.3 GB |
| DistilBERT |
0.84 |
25 ms |
250 MB |
| ruBERT-tiny2 |
0.81 |
12 ms |
120 MB |
| DistilBERT + ONNX |
0.84 |
14 ms |
150 MB |
How to choose between BERT and LLM for your task?
For most classification and extraction tasks, BERT-sized models offer the best trade-off between cost and performance. Shift to LLMs only when the task demands generation, complex reasoning, or zero-shot generalization.
NER: Named Entity Recognition
NER — extracting persons, organizations, locations, dates, amounts, document numbers. For general categories (PER, ORG, LOC), pre-trained models work well. For specialized ones (medical terms, legal concepts) — fine-tuning is needed.
Data annotation. The main cost of an NER project. For a quality model — 500-2,000 labeled sentences per entity type. Tools: Label Studio (open source) or Prodigy (by spaCy creators). IOB2 format — standard.
Architecture. Token classification on top of BERT: each token gets a label (B-PER, I-PER, O). spaCy 3.x with transformer pipeline — a convenient production choice.
Nested entities. Standard IOB models cannot handle nested entities (organization inside an address). For such tasks — span-based NER: SpanBERT or SpERT. More complex but correct.
Post-processing is mandatory. The model predicts tokens — normalized entities are needed. Date — dateparser. Amounts — regex + validation. Names — deduplication via rapidfuzz. Included in our standard delivery.
Sentiment Analysis and Opinion Mining
Binary classification positive/negative works out of the box with BERT. Complexity — aspect-based sentiment analysis (ABSA): "the restaurant has good food but terrible service." For ABSA: aspect extraction (NER) + sentiment per aspect. Joint models BERT-for-ABSA — quality on Russian data is lower due to dataset scarcity. RuSentiment, SentiRuEval — main resources.
For production with simple positive/negative/neutral: distil models are enough. Three classes, balanced dataset, 2,000+ examples — F1 macro 0.82-0.87 in 1-2 days.
Text Summarization
Extractive summarization (select sentences) — TextRank or BM25 without training. Fast, no hallucinations. Good for long documents.
Abstractive (generates new text) — seq2seq: mT5, mBART, FRED-T5, ruT5-large. For production via LLM API (GPT-4, Claude) — often the best cost/quality/speed trade-off.
Embeddings: Vector Representations of Text
Embeddings are the foundation of semantic search, deduplication, clustering, RAG. Quality critically affects downstream tasks.
Models. E5-large-v2, BGE-M3, multilingual-e5-large — strong multilingual embedders. sentence-transformers/paraphrase-multilingual-mpnet-base-v2 — fast option. For Russian: ru-en-RoSBERTa (Skoltech) performs well on semantic textual similarity.
Embedding quality evaluation uses the MTEB benchmark as standard. But top results on MTEB don't guarantee success on a domain dataset — we build domain-specific eval.
Fine-tuning embeddings. If standard models don't give the required Recall@k — contrastive learning on domain pairs with MultipleNegativesRankingLoss. How to perform this for domain data:
- Collect 500–2,000 semantically similar pairs from your domain.
- Apply MultipleNegativesRankingLoss with a batch size of 32–64.
- Train for 1–3 epochs using AdamW (lr=2e-5).
- Evaluate Recall@k on a held-out domain test set.
This approach yields a 5–15% improvement in Recall@k in practice.
Dimensionality and storage. E5-large: 1024 dim, float32 — 4KB per vector. For 10M documents — 40GB. Quantization int8 reduces to 10GB. FAISS IVF_PQ — more compact but with losses. Included in our deployment recommendations.
Information Extraction
Structured extraction is a frequent task. Examples: key contract terms, technical characteristics, dates and amounts from invoices.
- Regex + rule-based. For INN, OGRN, amounts, dates — more reliable than neural networks. No data required.
- NER + post-processing. For variable formats.
- LLM with structured output. GPT‑4 / Claude with JSON schema — for complex documents. Cost: minimal per document. For 10k+ documents/day — we calculate the economics.
We guarantee a hybrid: regex/NER for typical fields + LLM for edge cases. Our guarantee is backed by years of production experience and more than 30 projects.
Work Stages
| Stage |
Duration |
What's included |
| Data and metric analysis |
3-5 days |
Class distribution, text lengths, baseline |
| Baseline (TF‑IDF + LogReg) |
1 day |
Quick estimate of gap with deep models |
| Training and validation |
1-2 weeks |
k‑fold, early stopping, error analysis |
| Deployment (ONNX + FastAPI) |
1-2 weeks |
REST API, batching, monitoring |
| Documentation and training |
2-3 days |
Model card, API docs, team training |
Prototype on existing data — 1-3 weeks. Production system with CI/CD — 1.5–2.5 months. Cost is calculated individually — get a consultation for a project estimate.
What's Included
- Model and pipeline architecture documentation
- Access to the model via REST API (FastAPI + ONNX)
- Client team training (2-hour webinar + Q&A)
- Accuracy guarantee on the agreed test set
- Months of post-delivery support (bug fixes, adaptation to new data)
Our Experience
Years of NLP projects from classification to RAG systems. The team includes ML engineers experienced with Hugging Face, spaCy, LangChain, MLOps. We use vLLM, Kubeflow, Weights & Biases — a production stack, not toys. Contact us to evaluate your NLP project within two days — request a free consultation on your text processing pipeline.