Development of AI Workflows with Branching and Conditional Logic
When processing 500+ incoming documents daily—invoices, contracts, complaints, requests—it is critical to instantly route each document to the right specialist. A linear pipeline here is ineffective: classification errors reach 15%, and processing time per document is up to 45 minutes. As a result, urgent invoices get lost, complaints stall, and contracts are signed with delays. This problem is especially acute in companies with high document loads: retail, logistics, finance. Every day, dozens of specialists manually sort documents—which is not only slow but also costly: sorting takes up to 30% of working time. AI Workflow (also machine learning workflow) combining document processing AI, document automation AI, and conditional branching AI based on LangGraph eliminates these losses. We design such systems turnkey with a guaranteed result: auto-processing of up to 80% of the flow, classification accuracy from 90%, routing time—seconds instead of hours. Typical project investment ranges from $15,000 to $40,000, delivering an average of $200,000 in annual operational savings. By integrating conditional branching AI and LLM routing, LangGraph enables smart document processing AI with RAG pipeline support.
Choosing branching type for your workflow
All branching methods (conditional branching AI) fall into three categories. Comparison:
| Type |
Basis |
Example |
Control |
Flexibility |
Typical Accuracy |
| Deterministic |
Hard rules (if/else) |
if amount > 1000: |
Full |
Low |
95-99% (on structured data) |
| LLM-based |
Language model decision |
Sentiment classification |
Medium |
High |
85-92% (on unstructured) |
| Hybrid |
Combination |
Structured data - code, text - LLM |
High |
Maximum |
92-97% |
Choice of type depends on the task: for predictable scenarios, deterministic is sufficient; for complex unstructured ones, hybrid provides 2-3x accuracy improvement compared to a pure LLM approach. You can also combine with a RAG pipeline for context enrichment.
What is hybrid branching and when to use it?
Hybrid branching combines deterministic rules and LLM-based decisions in one workflow. This approach leverages LLM-based decisions for complex cases. For example, when processing an invoice, a deterministic node checks the format and numeric fields, while the LLM extracts unstructured data (product names, terms). If the LLM cannot find the amount, a retry cycle with a modified prompt kicks in. This approach reduces latency by 30% compared to a pure LLM solution and ensures 95% accuracy even on noisy data. Hybrid branching is indispensable when both structured and unstructured documents appear in a single stream. It achieves 2x higher accuracy than purely deterministic systems on unstructured data.
How to implement cyclic branching with retry?
In real scenarios, data may be incomplete or of low quality. We add a retry loop machine learning with a limit on attempts. The retry loops machine learning technique ensures robust handling of errors.
Example code for cyclic branching with retry
MAX_RETRIES = 3
def check_quality_and_retry(state: WorkflowState) -> str:
"""Decides: accept the result or send for rework"""
if state.get("retry_count", 0) >= MAX_RETRIES:
return "accept" # Accept even imperfect result
quality = assess_output_quality(state["output"])
if quality < 0.8:
return "retry"
return "accept"
def increment_retry(state: WorkflowState) -> WorkflowState:
return {**state, "retry_count": state.get("retry_count", 0) + 1}
# Add cycle to graph
graph.add_conditional_edges("quality_check", check_quality_and_retry, {
"retry": "processing_node",
"accept": "finalize",
})
Advantages of hybrid branching
From our practice: a workflow for processing incoming correspondence. The client received 500+ documents per day (email attachments, portal uploads). Deterministic branching couldn't handle it—documents of the same type often required different actions. We implemented a hybrid system where code handles structured fields (numbers, dates) and the LLM classifies ambiguous cases. Result: auto-processing without manual intervention—71%, routing time reduced from 45 minutes to instant, classification accuracy—94%, routing errors—2.1%. Our solution processes documents 10x faster than manual routing. LLM routing decides document path based on content. This approach allowed the client to significantly reduce manual processing costs.
Typical mistakes in AI Workflow development
| Mistake |
Reason |
Solution |
| Insufficient edge case testing |
Ignoring non-standard inputs |
Cover tests: empty documents, broken files, unexpected formats |
| No fallback routes |
Relying on constant model availability |
Configure switch to backup model or manual mode on failures |
| Ignoring latency |
Not accounting for LLM response time |
Use timeout and async calls, track p99 latency (<1 sec) |
| Infinite retry loops |
No limit on number of attempts |
Introduce MAX_RETRIES and quality exit condition |
Process of work
- Analytics—study input data, branching types, latency and accuracy requirements. Fix target metrics: accuracy, p99 latency, throughput (1000+ documents/day).
- Graph design—draw a diagram, define nodes and conditional edges. Choose branching type for each node.
- Implementation—write code in LangGraph, integrate with LLM (GPT-4, Claude) and external APIs. Add logging and monitoring.
- Testing—cover each node with unit tests, check edge cases (empty inputs, model errors). Conduct load testing on real data.
- Deployment—deploy under load, configure alerts for metrics: accuracy, latency p99, throughput. Ensure hot-reload for quick fixes.
Deliverables package
- Project documentation: graph description, node and conditional edge specifications, use cases.
- Source code in LangGraph with comments and deployment instructions.
- Access to a version-controlled repository (Git) with CI/CD.
- Team training (up to 2 online sessions) on configuring and modifying the workflow.
- Technical support for 3 months after implementation.
Guarantees and experience
Our engineers have 10+ years of experience in AI/ML, implemented over 50 AI automation projects. 5 years on the market. We guarantee: classification accuracy not lower than 90%, auto-processing of 70%+ of the flow, system response time—less than 1 second per document, 99.9% uptime. We record results in an SLA with monthly audits. Typical annual savings: $200,000.
Timelines
- Design: from 1 week.
- Implementation of nodes and branches: 2–3 weeks.
- Edge case testing: 1–2 weeks.
- Total turnkey: 4–6 weeks.
Get a consultation from an engineer—we will evaluate your case and propose a solution. Contact us for a project cost and timeline estimate.
LLM Development: Fine-Tuning, RAG, Agents, and Production Deployment
Using GPT‑4 or Claude 3.5 Sonnet through a public API is not a solution — it's just a tool. When the requirement is to "make it like ChatGPT, but on our data," there is a real engineering challenge behind it: from prompt engineering to training a 70B model on your own infrastructure. End-to-end LLM solution development is a complex stack, and we have been doing it for over 5 years. During this time, we have completed over 20 projects in generative AI: from RAG systems for legal departments to custom support agents. Where exactly your task falls depends on data, latency requirements, budget, and how critical confidentiality is.
A typical situation: the client has already tried ChatGPT, but results are unstable — sometimes accurate, sometimes hallucinating. Or they need integration into a corporate portal while complying with security policies. Let's break down each layer of the stack in detail — from RAG to production deployment.
Why Do RAG Systems Break and How to Fix It?
RAG (Retrieval-Augmented Generation) looks simple: find relevant documents, put them in context, get an answer. In practice, it fails in several places.
Chunking without overlap. Classic mistake: chunk_size=512, overlap=0. If the answer lies across two chunks, retrieval won't find either with sufficient confidence. Solution: overlap 15–25% of chunk_size, or better yet, sentence-aware splitting with spaCy or NLTK instead of naive character splitting.
Poor embedder. text-embedding-ada-002 is good for general use, but on legal or medical texts, specialized models like E5-large-v2, BGE-M3, or fine-tuned sentence-transformers on domain data outperform it. Recall@5 differences can be 15–25%.
No re-ranking. Vector search optimizes for speed, not relevance. A cross-encoder re-ranker (ms-marco-MiniLM-L-6-v2, bge-reranker-large) after initial retrieval improves top-3 accuracy with acceptable latency (+50–150ms). This is often more impactful than improving the embedding model.
Hybrid search. Dense vectors alone work poorly on exact queries: names, SKUs, codes. BM25 (sparse) finds exact matches but misses semantics. Hybrid via RRF (Reciprocal Rank Fusion) is the optimal compromise. Qdrant, Weaviate, and pgvector 0.7+ support hybrid search natively.
Typical production architecture for a corporate knowledge base
- Documents → preprocessing (PyMuPDF, Unstructured)
- Chunking → embedding (BGE-M3)
- Qdrant (hybrid dense+sparse)
- Cross-encoder re-ranking
- Context → LLM (vLLM or OpenAI API)
- Answer with sources (RAGAS for quality evaluation)
When to Fine-Tune Instead of Prompt Engineering?
Prompt engineering solves ~70% of LLM adaptation tasks for a domain. The remaining 30% require fine-tuning. Three indicators: the model ignores a specific output format even with detailed prompting; the task requires deep knowledge of specialized vocabulary (medicine, law); you need to significantly reduce token costs by replacing a large model with a smaller specialized one.
LoRA and QLoRA are the standard for SFT. LoRA adds trainable low-rank matrices to attention layers. A typical configuration for Llama-3 8B: r=64, lora_alpha=128, target_modules=["q_proj","v_proj","k_proj","o_proj"] yields ~0.8% trainable parameters, training on one A100 40GB. QLoRA adds 4-bit quantization (NF4) and allows fine-tuning 70B models on two A100 40GB, though speed drops by half compared to bf16.
DPO instead of RLHF. Direct Preference Optimization requires only (chosen, rejected) pairs, not scalar reward signals. DPOTrainer from the trl library (Hugging Face) implements it in a few dozen lines.
Common mistake. A dataset of 500 examples, 5 epochs, validation loss 0.8 — seems fine. But on test, the model degrades on general instructions. Cause: catastrophic forgetting. Solution: add 10–20% general instruction-following examples (Alpaca, FLAN) to the training set to preserve original capabilities.
How to Choose a Base Model: 8B or 70B?
| Model |
Parameters |
Strengths |
Context |
| Llama-3.1 8B |
8B |
Quality/speed balance |
128k |
| Llama-3.1 70B |
70B |
Complex reasoning |
128k |
| Mistral 7B / Mixtral 8x7B |
7B / 47B |
Efficiency for size |
32k |
| Qwen2.5 72B |
72B |
Code, multilingual |
128k |
| Gemma 2 27B |
27B |
Open license |
8k |
For most tasks, fine-tuning an 8B model is sufficient. 70B is needed when deep reasoning is required or the 8B baseline does not reach the required quality even after fine-tuning. Inference cost for Llama-3 8B via vLLM on A100 is efficient; the exact cost depends on volume.
What Does PagedAttention Bring to Production?
vLLM is the first choice for serving open-source models. PagedAttention is the key technical innovation: KV-cache is managed like virtual memory in an OS, without fragmentation. This yields 2–4x higher throughput compared to naive HuggingFace Transformers inference. The vLLM documentation confirms that continuous batching and PagedAttention are the standard for high-load LLM services.
Typical numbers on A100 80GB for Llama-3 8B (bf16): 400–600 req/s, P50 latency 200–400ms, P99 latency 600–900ms at concurrency 64. For 70B on two A100 with tensor parallelism: 80–120 req/s, P99 latency 1.5–2.5s. AWQ or GPTQ quantization reduces memory consumption by 2x with quality loss within 1–3%.
Multi-Agent Systems
Agents are LLMs with access to tools: search, code execution, API calls, database interaction. Common patterns:
- ReAct (Reason + Act): the model reasons → chooses a tool → observes the result → reasons again. LangChain and LlamaIndex implement it out of the box.
- Multi-agent orchestration: multiple specialized agents with a coordinator on top. Example: coordinator → researcher (search + summarization) → coder (code generation and execution) → critic (verification). Tools: AutoGen (Microsoft), CrewAI, custom implementation on LangGraph.
In production, agent systems are non-deterministic. Essential: guardrails, step limits, logging of each step, human-in-the-loop for critical actions.
How We Work: Stages, Timeline, Deliverables
| Stage |
Duration |
What You Get |
| Audit and data collection |
1–2 weeks |
Eval dataset of 100+ examples, task formalization |
| Baseline (prompt + RAG) |
1–2 weeks |
Working prototype, quality metrics |
| Fine-tuning (if needed) |
2–4 weeks |
Trained model, LoRA weights, model card |
| Deployment and monitoring |
1–2 weeks |
vLLM server, Grafana + Prometheus |
| Documentation and training |
1 week |
API documentation, team training |
What Is Included
We deliver:
- Technical documentation (model card, configs, deployment instructions)
- Access to infrastructure (code repository, trained weights)
- 1 month of post-deployment support (consultations, bug fixes)
- Customer team training (2–3 sessions on system operation)
Timeline: basic RAG prototype — 1–2 weeks. Fine-tuning with customer data — 3–6 weeks (including data preparation). Production system with monitoring and retraining — 2–4 months. Cost is calculated individually based on data volume, model complexity, and infrastructure requirements.
We guarantee the quality of the final model with performance benchmarks and ongoing monitoring. Our engineers have hands‑on experience with dozens of production LLM systems.
Want to evaluate your project? Leave a request — we will prepare a preliminary summary within 1–2 business days. Or get a consultation on choosing the approach: RAG, fine-tuning, or hybrid — we will tell you what works best for you. Contact us to discuss your LLM development needs. Schedule a free consultation today.