Integrating LangGraph for Graph-Based AI Agents

We design and deploy artificial intelligence systems: from prototype to production-ready solutions. Our team combines expertise in machine learning, data engineering and MLOps to make AI work not in the lab, but in real business.
Showing 1 of 1All 1564 services
Integrating LangGraph for Graph-Based AI Agents
Medium
from 1 week to 3 months
Frequently Asked Questions

AI Development Areas

AI Solution Development Stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1351
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1247
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    950
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1186
  • image_logo-advance_0.webp
    B2B Advance company logo design
    642
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    922

When developing complex AI agents, we often encounter situations where linear LangChain Expression Language (LCEL) chains stop being effective. When cyclical processing, conditional transitions based on intermediate step results, or pauses for human confirmation are required, LCEL falls short. That's exactly where LangGraph comes in—a library that extends LangChain and allows building agents as directed graphs with explicit state. Our experience deploying LangGraph in production includes projects for fintech, logistics, and legal services. We guarantee a fault-tolerant architecture and full documentation.

LangGraph implements the StateGraph concept, where each node is a function and edges are transitions. The agent's state is described with a TypedDict and can automatically merge messages. This makes it easy to implement loops and multi-agent systems.

Why agents need a graph, not a chain?

Linear LCEL chains are fine for simple pipelines: take input, apply a sequence of steps, get output. But real agent scenarios often require going back to a previous step, running parallel checks, or pausing execution for manual control. LangGraph's graph model solves these problems naturally: cycles are edges leading back; parallelism is multiple nodes executing simultaneously; human-in-the-loop is built-in interruptions.

Basic graph structure

from langgraph.graph import StateGraph, END
from langgraph.checkpoint.memory import MemorySaver
from langgraph.prebuilt import ToolNode
from langchain_openai import ChatOpenAI
from langchain_core.messages import HumanMessage, AIMessage
from typing import TypedDict, Annotated
import operator

class AgentState(TypedDict):
    messages: Annotated[list, operator.add]  # Automatically concatenated
    user_id: str
    iteration_count: int

llm = ChatOpenAI(model="gpt-4o")

def agent_node(state: AgentState) -> AgentState:
    response = llm.bind_tools(tools).invoke(state["messages"])
    return {"messages": [response], "iteration_count": state["iteration_count"] + 1}

def should_continue(state: AgentState) -> str:
    last_msg = state["messages"][-1]
    if last_msg.tool_calls:
        return "tools"
    return END

# Build the graph
graph = StateGraph(AgentState)
graph.add_node("agent", agent_node)
graph.add_node("tools", ToolNode(tools))

graph.set_entry_point("agent")
graph.add_conditional_edges("agent", should_continue, {"tools": "tools", END: END})
graph.add_edge("tools", "agent")  # Loop: after tools, back to agent

app = graph.compile(checkpointer=MemorySaver())

Persistent state and interrupts

LangGraph supports checkpoint saving between runs and pauses for human approval:

from langgraph.checkpoint.postgres import PostgresSaver
from psycopg import Connection

# Persistence in PostgreSQL
conn = Connection.connect("postgresql://user:pass@localhost/langgraph_db")
checkpointer = PostgresSaver(conn)

# Interrupt: graph stops before the specified node
app = graph.compile(
    checkpointer=checkpointer,
    interrupt_before=["execute_payment"],  # Requires human confirmation
)

config = {"configurable": {"thread_id": "order_12345"}}

# Run until the break point
result = app.invoke({"messages": [HumanMessage("Pay the invoice for 50000 RUB")]}, config)
# Graph stopped before execute_payment

# After human review, continue
app.invoke(None, config)  # None = resume from current state

Multi-agent: Supervisor pattern

from langgraph.graph import StateGraph, END
from typing import Literal

class SupervisorState(TypedDict):
    messages: Annotated[list, operator.add]
    next_agent: str

AGENTS = ["researcher", "analyst", "writer"]

supervisor_prompt = f"""You are a supervisor of a multi-agent system.
Based on the request and current progress, choose the next agent: {AGENTS}
Or return FINISH if the task is complete.
"""

def supervisor_node(state: SupervisorState):
    response = llm.with_structured_output(
        {"next": {"type": "string", "enum": AGENTS + ["FINISH"]}}
    ).invoke([{"role": "system", "content": supervisor_prompt}] + state["messages"])
    return {"next_agent": response["next"]}

def route_to_agent(state: SupervisorState) -> str:
    if state["next_agent"] == "FINISH":
        return END
    return state["next_agent"]

# Create agents
def make_agent_node(name: str, system_prompt: str):
    agent_llm = ChatOpenAI(model="gpt-4o").bind_tools(get_tools_for(name))
    def node(state):
        result = agent_llm.invoke(
            [{"role": "system", "content": system_prompt}] + state["messages"]
        )
        return {"messages": [result]}
    return node

graph = StateGraph(SupervisorState)
graph.add_node("supervisor", supervisor_node)
graph.add_node("researcher", make_agent_node("researcher", "Research the topic and find facts"))
graph.add_node("analyst", make_agent_node("analyst", "Analyze data and draw conclusions"))
graph.add_node("writer", make_agent_node("writer", "Formulate the final answer"))

graph.set_entry_point("supervisor")
graph.add_conditional_edges("supervisor", route_to_agent)
for agent in AGENTS:
    graph.add_edge(agent, "supervisor")

multi_agent = graph.compile()

Streaming and real-time output

# Streaming events from the graph
async for event in app.astream_events(
    {"messages": [HumanMessage("Analyze Q1 sales")]},
    config={"configurable": {"thread_id": "analysis_001"}},
    version="v2",
):
    kind = event["event"]
    if kind == "on_chat_model_stream":
        print(event["data"]["chunk"].content, end="", flush=True)
    elif kind == "on_tool_start":
        print(f"\n[Tool call: {event['name']}]")
    elif kind == "on_tool_end":
        print(f"[Tool result received]")

How nested graphs simplify modularity?

LangGraph supports SubGraphs—nested graphs. You can define a graph for document processing and then include it as a regular node in the parent graph. This allows decomposing complex systems into reusable components.

# Subgraph for document processing
doc_graph = StateGraph(DocumentState)
doc_graph.add_node("extract", extract_text)
doc_graph.add_node("classify", classify_document)
doc_graph.add_node("validate", validate_structure)
# ... build the subgraph

doc_subgraph = doc_graph.compile()

# Include subgraph in parent
main_graph = StateGraph(MainState)
main_graph.add_node("process_document", doc_subgraph)  # Subgraph as node
main_graph.add_node("send_result", send_to_crm)
main_graph.add_edge("process_document", "send_result")

Practical case: contract review system (from our practice)

Task: one of our clients—a legal department of a large company—received 30–50 contracts daily. Each contract required 1–2 hours of lawyer time. We built a graph agent on LangGraph that automated the review.

Graph:

  1. extract_node — parse PDF, extract structure
  2. classify_node — contract type (supply, services, lease, NDA)
  3. risk_check_node — parallel checks: financial terms, duration, liability
  4. legal_rules_node — check against corporate list of prohibited clauses
  5. human_review — interrupt for contracts with risk_score > 7
  6. finalize_node — generate conclusion and recommendations
app = graph.compile(
    checkpointer=PostgresSaver(conn),
    interrupt_before=["human_review"],  # Pause only for risky ones
)

Routing: low risk → automatic approval; high risk → pause for lawyer with agent's draft conclusion.

Results:

  • Standard contract review time: 90 min → 8 min
  • Automatic approval without lawyer: 61% of contracts
  • Missed non-standard clauses: 0 (vs ~3% manual due to fatigue)
  • Legal department workload: -58%

Comparison of LangGraph and LCEL

Criteria LCEL LangGraph
Structure Linear chain Arbitrary graph
Loops No Yes
State Passed via pipe TypedDict with merge strategy
Checkpoint No PostgreSQL, Redis, SQLite
Human-in-the-loop No interrupt_before/after
Use case Simple pipelines Agents, multi-agent systems

LangGraph components

Component Purpose Example
StateGraph Defines state and nodes StateGraph(AgentState)
Node Processing function agent_node
Edge Connection between nodes graph.add_edge("tools", "agent")
Conditional Edge Conditional transition should_continue
Checkpointer State persistence MemorySaver, PostgresSaver
Interrupt Pause for HITL interrupt_before=["node"]

What our LangGraph integration includes

  • Architecture session: analyze the task, select graph topology, define human-in-the-loop points.
  • Code development: implement nodes, edges, state, integrate with your LLM and tools.
  • Persistence setup: PostgreSQL, Redis, or SQLite for checkpoints.
  • Environment integration: deploy via Docker/Kubernetes, connect monitoring (LangSmith).
  • Testing: unit tests, load testing with latency recording (p99).
  • Documentation: full graph description, API, operation guide.
  • Team training: workshop on LangGraph for your developers.

Estimated timelines

  • Basic ReAct agent on LangGraph: 3 to 5 days.
  • Multi-agent system with supervisor: 2 to 3 weeks.
  • Human-in-the-loop workflow with persistence: 1 to 2 weeks.
  • Production integration with PostgreSQL checkpoint: +3–5 days.

Pricing is determined individually—we assess the project free of charge within one business day. Contact us for a consultation: our certified engineers will help choose the architecture for your task. Request a preliminary assessment, and we will prepare a commercial proposal with a quality guarantee and transparent timelines.

LLM Development: Fine-Tuning, RAG, Agents, and Production Deployment

Using GPT‑4 or Claude 3.5 Sonnet through a public API is not a solution — it's just a tool. When the requirement is to "make it like ChatGPT, but on our data," there is a real engineering challenge behind it: from prompt engineering to training a 70B model on your own infrastructure. End-to-end LLM solution development is a complex stack, and we have been doing it for over 5 years. During this time, we have completed over 20 projects in generative AI: from RAG systems for legal departments to custom support agents. Where exactly your task falls depends on data, latency requirements, budget, and how critical confidentiality is.

A typical situation: the client has already tried ChatGPT, but results are unstable — sometimes accurate, sometimes hallucinating. Or they need integration into a corporate portal while complying with security policies. Let's break down each layer of the stack in detail — from RAG to production deployment.

Why Do RAG Systems Break and How to Fix It?

RAG (Retrieval-Augmented Generation) looks simple: find relevant documents, put them in context, get an answer. In practice, it fails in several places.

Chunking without overlap. Classic mistake: chunk_size=512, overlap=0. If the answer lies across two chunks, retrieval won't find either with sufficient confidence. Solution: overlap 15–25% of chunk_size, or better yet, sentence-aware splitting with spaCy or NLTK instead of naive character splitting.

Poor embedder. text-embedding-ada-002 is good for general use, but on legal or medical texts, specialized models like E5-large-v2, BGE-M3, or fine-tuned sentence-transformers on domain data outperform it. Recall@5 differences can be 15–25%.

No re-ranking. Vector search optimizes for speed, not relevance. A cross-encoder re-ranker (ms-marco-MiniLM-L-6-v2, bge-reranker-large) after initial retrieval improves top-3 accuracy with acceptable latency (+50–150ms). This is often more impactful than improving the embedding model.

Hybrid search. Dense vectors alone work poorly on exact queries: names, SKUs, codes. BM25 (sparse) finds exact matches but misses semantics. Hybrid via RRF (Reciprocal Rank Fusion) is the optimal compromise. Qdrant, Weaviate, and pgvector 0.7+ support hybrid search natively.

Typical production architecture for a corporate knowledge base
  1. Documents → preprocessing (PyMuPDF, Unstructured)
  2. Chunking → embedding (BGE-M3)
  3. Qdrant (hybrid dense+sparse)
  4. Cross-encoder re-ranking
  5. Context → LLM (vLLM or OpenAI API)
  6. Answer with sources (RAGAS for quality evaluation)

When to Fine-Tune Instead of Prompt Engineering?

Prompt engineering solves ~70% of LLM adaptation tasks for a domain. The remaining 30% require fine-tuning. Three indicators: the model ignores a specific output format even with detailed prompting; the task requires deep knowledge of specialized vocabulary (medicine, law); you need to significantly reduce token costs by replacing a large model with a smaller specialized one.

LoRA and QLoRA are the standard for SFT. LoRA adds trainable low-rank matrices to attention layers. A typical configuration for Llama-3 8B: r=64, lora_alpha=128, target_modules=["q_proj","v_proj","k_proj","o_proj"] yields ~0.8% trainable parameters, training on one A100 40GB. QLoRA adds 4-bit quantization (NF4) and allows fine-tuning 70B models on two A100 40GB, though speed drops by half compared to bf16.

DPO instead of RLHF. Direct Preference Optimization requires only (chosen, rejected) pairs, not scalar reward signals. DPOTrainer from the trl library (Hugging Face) implements it in a few dozen lines.

Common mistake. A dataset of 500 examples, 5 epochs, validation loss 0.8 — seems fine. But on test, the model degrades on general instructions. Cause: catastrophic forgetting. Solution: add 10–20% general instruction-following examples (Alpaca, FLAN) to the training set to preserve original capabilities.

How to Choose a Base Model: 8B or 70B?

Model Parameters Strengths Context
Llama-3.1 8B 8B Quality/speed balance 128k
Llama-3.1 70B 70B Complex reasoning 128k
Mistral 7B / Mixtral 8x7B 7B / 47B Efficiency for size 32k
Qwen2.5 72B 72B Code, multilingual 128k
Gemma 2 27B 27B Open license 8k

For most tasks, fine-tuning an 8B model is sufficient. 70B is needed when deep reasoning is required or the 8B baseline does not reach the required quality even after fine-tuning. Inference cost for Llama-3 8B via vLLM on A100 is efficient; the exact cost depends on volume.

What Does PagedAttention Bring to Production?

vLLM is the first choice for serving open-source models. PagedAttention is the key technical innovation: KV-cache is managed like virtual memory in an OS, without fragmentation. This yields 2–4x higher throughput compared to naive HuggingFace Transformers inference. The vLLM documentation confirms that continuous batching and PagedAttention are the standard for high-load LLM services.

Typical numbers on A100 80GB for Llama-3 8B (bf16): 400–600 req/s, P50 latency 200–400ms, P99 latency 600–900ms at concurrency 64. For 70B on two A100 with tensor parallelism: 80–120 req/s, P99 latency 1.5–2.5s. AWQ or GPTQ quantization reduces memory consumption by 2x with quality loss within 1–3%.

Multi-Agent Systems

Agents are LLMs with access to tools: search, code execution, API calls, database interaction. Common patterns:

  • ReAct (Reason + Act): the model reasons → chooses a tool → observes the result → reasons again. LangChain and LlamaIndex implement it out of the box.
  • Multi-agent orchestration: multiple specialized agents with a coordinator on top. Example: coordinator → researcher (search + summarization) → coder (code generation and execution) → critic (verification). Tools: AutoGen (Microsoft), CrewAI, custom implementation on LangGraph.

In production, agent systems are non-deterministic. Essential: guardrails, step limits, logging of each step, human-in-the-loop for critical actions.

How We Work: Stages, Timeline, Deliverables

Stage Duration What You Get
Audit and data collection 1–2 weeks Eval dataset of 100+ examples, task formalization
Baseline (prompt + RAG) 1–2 weeks Working prototype, quality metrics
Fine-tuning (if needed) 2–4 weeks Trained model, LoRA weights, model card
Deployment and monitoring 1–2 weeks vLLM server, Grafana + Prometheus
Documentation and training 1 week API documentation, team training

What Is Included

We deliver:

  • Technical documentation (model card, configs, deployment instructions)
  • Access to infrastructure (code repository, trained weights)
  • 1 month of post-deployment support (consultations, bug fixes)
  • Customer team training (2–3 sessions on system operation)

Timeline: basic RAG prototype — 1–2 weeks. Fine-tuning with customer data — 3–6 weeks (including data preparation). Production system with monitoring and retraining — 2–4 months. Cost is calculated individually based on data volume, model complexity, and infrastructure requirements.

We guarantee the quality of the final model with performance benchmarks and ongoing monitoring. Our engineers have hands‑on experience with dozens of production LLM systems.

Want to evaluate your project? Leave a request — we will prepare a preliminary summary within 1–2 business days. Or get a consultation on choosing the approach: RAG, fine-tuning, or hybrid — we will tell you what works best for you. Contact us to discuss your LLM development needs. Schedule a free consultation today.