How ChatDev accelerates prototyping?
Imagine compressing a sprint into one day: a product manager describes an idea, and a developer spends half a day cutting a prototype. A week later the hypothesis is validated, but the prototype goes into the bin — 80% of the time went into routine. ChatDev is a framework from Tsinghua University where AI agents with roles of CEO, CTO, Programmer, Reviewer, and Tester mimic a waterfall process, generating working code from a natural language description. We use it in our project workflow to obtain a PoC in 30–90 minutes instead of 1–2 days.
How it works?
ChatDev builds a chain of dialogues: CEO sets the task, CTO chooses architecture, Programmer writes code, Reviewer and Tester find errors. Each phase is a separate dialogue with clear input/output. It is not a chatbot but a managed process where the result of one phase feeds into the next.
Why chat agents are more effective than a single LLM prompt?
A single run through GPT-4 with a prompt "create a scraper" often produces non-working code — mixes libraries, misses edge cases. ChatDev splits the task: CTO specifies architecture, Programmer writes to spec, Reviewer checks compliance. Each agent sees only its part of the context, reducing the load on the model's attention window. Practice shows: the percentage of immediately deployable code is ~15% for simple utilities, but time to first working commit decreases by 60–70%.
Examples of ChatDev integration
Basic launch and Python API
# Installation
git clone https://github.com/OpenBMB/ChatDev.git
cd ChatDev
pip install -r requirements.txt
# Launch development
python run.py \
--task "Develop a web scraper to extract prices from e-commerce sites. Python + BeautifulSoup + requests. Save to CSV." \
--name "PriceScraperApp" \
--model GPT_4_TURBO
from chatdev.chat_chain import ChatChain
import os
os.environ["OPENAI_API_KEY"] = "sk-..."
chat_chain = ChatChain(
config_path="CompanyConfig/Default",
config_phase_path="PhaseConfig/Default",
config_role_path="RoleConfig/Default",
task_prompt="Create a utility for converting data formats (CSV, JSON, XML, YAML)",
project_name="DataConverter",
org_name="TechTeam",
model_type="GPT_4_TURBO",
)
chat_chain.pre_processing()
chat_chain.make_recruitment()
chat_chain.execute_chain()
chat_chain.post_processing()
Custom phases
{
"chain": [
{"phase": "DemandAnalysis", "phaseType": "SimplePhase"},
{"phase": "LanguageChoose", "phaseType": "SimplePhase"},
{"phase": "Coding", "phaseType": "SimplePhase"},
{"phase": "SecurityAudit", "phaseType": "SimplePhase"},
{"phase": "CodeCompleteAll", "phaseType": "ComposedPhase",
"cycleNum": 3,
"Composition": [
{"phase": "CodeReviewComment", "phaseType": "SimplePhase"},
{"phase": "CodeReviewModification", "phaseType": "SimplePhase"},
{"phase": "TestErrorSummary", "phaseType": "SimplePhase"},
{"phase": "TestModification", "phaseType": "SimplePhase"}
]
},
{"phase": "Manual", "phaseType": "SimplePhase"}
]
}
Iterative improvement via ExperiencePool
from chatdev.experience_pool import ExperiencePool
experience_pool = ExperiencePool.load("./experience_pool")
chat_chain = ChatChain(
task_prompt="...",
experience_pool=experience_pool,
use_experience=True,
)
experience_pool.update(chat_chain.get_experience())
experience_pool.save("./experience_pool")
Practical results and limitations
Case: rapid prototyping
Problem: a client team generated 2–3 prototypes per week for hypothesis validation. Each prototype required 1–2 developer days. We introduced ChatDev as a draft generator.
Usage: prototypes of simple utilities (converters, scrapers, report generators), CLI tools for internal use, basic CRUD APIs for PoCs.
Results from practice:
- Prototype creation time: 1–2 days → 30–90 minutes
- Review and refinement: 2–4 hours (vs. full write from scratch)
- Production readiness without refinement: ~15% (only the simplest utilities)
- Main value: fast concept validation — the team tested 12 hypotheses in a quarter instead of 6
- Time savings: ~15 person-days per month
Limitations: maximum efficiency on projects up to 500 lines of code. Complex multi-file projects with dependencies require significant post-processing. No native support for existing codebases.
Comparison with MetaGPT
| Aspect |
ChatDev |
MetaGPT |
| Approach |
Dialogue-based (roles talk) |
SOP-based (formal documents) |
| Code quality |
Satisfactory for PoC |
Higher for production code |
| Customisation |
JSON configs |
Python API + custom roles |
| Project size |
Up to ~500 lines |
Up to several thousand lines |
| Research-oriented |
Yes |
Less |
What the integration includes
When ordering ChatDev integration, you get: agent configuration for your typical tasks, a set of custom phases (SecurityAudit, PerformanceCheck, StyleLint), CI/CD integration (trigger generation from a button or from Jira), a basic ExperiencePool assembly with patterns from your projects, documentation (architecture, launch examples, troubleshooting) and a team training session (4 hours, remote).
Working process:
- Analytics: we analyse typical team tasks, determine the suitable volume for ChatDev (up to 500 lines).
- Design: configure roles, phases, ExperiencePool for your tech stack (Python, Go, TypeScript — any).
- Implementation: write configs, integrate API, add CI triggers.
- Testing: run through 10–15 real tickets, record generation success rate.
- Deployment: deploy on a dev stand or in the cloud, hand over refined templates.
Timelines: basic launch and setup — from 1 day; custom phases aligned with team processes — 3–5 days; full integration with CI and training — from 1 week. Cost is calculated individually — contact us to discuss details. Get a consultation: we will send a sample config for your project and estimate timelines.
Typical mistakes and ChatDev's role
- Overly long task descriptions (>200 words): agents lose focus — split the task into subtasks.
- Using for legacy refactoring: ChatDev does not see the code, generates from scratch. An agent-analyser is needed.
- Skipping the SecurityAudit phase: generated code may contain vulnerabilities (SQL injections, hardcoded tokens).
ChatDev will not replace developers but will accelerate them: it takes over routine (writing boilerplate, tests, documentation). 85% of results require refinement, but that refinement takes hours, not days. If you generate 20 prototypes per month, ChatDev saves ~15 person-days. That is enough to clear the backlog or launch a parallel experiment. Contact us — we will show how to embed ChatDev into your pipeline.
LLM Development: Fine-Tuning, RAG, Agents, and Production Deployment
Using GPT‑4 or Claude 3.5 Sonnet through a public API is not a solution — it's just a tool. When the requirement is to "make it like ChatGPT, but on our data," there is a real engineering challenge behind it: from prompt engineering to training a 70B model on your own infrastructure. End-to-end LLM solution development is a complex stack, and we have been doing it for over 5 years. During this time, we have completed over 20 projects in generative AI: from RAG systems for legal departments to custom support agents. Where exactly your task falls depends on data, latency requirements, budget, and how critical confidentiality is.
A typical situation: the client has already tried ChatGPT, but results are unstable — sometimes accurate, sometimes hallucinating. Or they need integration into a corporate portal while complying with security policies. Let's break down each layer of the stack in detail — from RAG to production deployment.
Why Do RAG Systems Break and How to Fix It?
RAG (Retrieval-Augmented Generation) looks simple: find relevant documents, put them in context, get an answer. In practice, it fails in several places.
Chunking without overlap. Classic mistake: chunk_size=512, overlap=0. If the answer lies across two chunks, retrieval won't find either with sufficient confidence. Solution: overlap 15–25% of chunk_size, or better yet, sentence-aware splitting with spaCy or NLTK instead of naive character splitting.
Poor embedder. text-embedding-ada-002 is good for general use, but on legal or medical texts, specialized models like E5-large-v2, BGE-M3, or fine-tuned sentence-transformers on domain data outperform it. Recall@5 differences can be 15–25%.
No re-ranking. Vector search optimizes for speed, not relevance. A cross-encoder re-ranker (ms-marco-MiniLM-L-6-v2, bge-reranker-large) after initial retrieval improves top-3 accuracy with acceptable latency (+50–150ms). This is often more impactful than improving the embedding model.
Hybrid search. Dense vectors alone work poorly on exact queries: names, SKUs, codes. BM25 (sparse) finds exact matches but misses semantics. Hybrid via RRF (Reciprocal Rank Fusion) is the optimal compromise. Qdrant, Weaviate, and pgvector 0.7+ support hybrid search natively.
Typical production architecture for a corporate knowledge base
- Documents → preprocessing (PyMuPDF, Unstructured)
- Chunking → embedding (BGE-M3)
- Qdrant (hybrid dense+sparse)
- Cross-encoder re-ranking
- Context → LLM (vLLM or OpenAI API)
- Answer with sources (RAGAS for quality evaluation)
When to Fine-Tune Instead of Prompt Engineering?
Prompt engineering solves ~70% of LLM adaptation tasks for a domain. The remaining 30% require fine-tuning. Three indicators: the model ignores a specific output format even with detailed prompting; the task requires deep knowledge of specialized vocabulary (medicine, law); you need to significantly reduce token costs by replacing a large model with a smaller specialized one.
LoRA and QLoRA are the standard for SFT. LoRA adds trainable low-rank matrices to attention layers. A typical configuration for Llama-3 8B: r=64, lora_alpha=128, target_modules=["q_proj","v_proj","k_proj","o_proj"] yields ~0.8% trainable parameters, training on one A100 40GB. QLoRA adds 4-bit quantization (NF4) and allows fine-tuning 70B models on two A100 40GB, though speed drops by half compared to bf16.
DPO instead of RLHF. Direct Preference Optimization requires only (chosen, rejected) pairs, not scalar reward signals. DPOTrainer from the trl library (Hugging Face) implements it in a few dozen lines.
Common mistake. A dataset of 500 examples, 5 epochs, validation loss 0.8 — seems fine. But on test, the model degrades on general instructions. Cause: catastrophic forgetting. Solution: add 10–20% general instruction-following examples (Alpaca, FLAN) to the training set to preserve original capabilities.
How to Choose a Base Model: 8B or 70B?
| Model |
Parameters |
Strengths |
Context |
| Llama-3.1 8B |
8B |
Quality/speed balance |
128k |
| Llama-3.1 70B |
70B |
Complex reasoning |
128k |
| Mistral 7B / Mixtral 8x7B |
7B / 47B |
Efficiency for size |
32k |
| Qwen2.5 72B |
72B |
Code, multilingual |
128k |
| Gemma 2 27B |
27B |
Open license |
8k |
For most tasks, fine-tuning an 8B model is sufficient. 70B is needed when deep reasoning is required or the 8B baseline does not reach the required quality even after fine-tuning. Inference cost for Llama-3 8B via vLLM on A100 is efficient; the exact cost depends on volume.
What Does PagedAttention Bring to Production?
vLLM is the first choice for serving open-source models. PagedAttention is the key technical innovation: KV-cache is managed like virtual memory in an OS, without fragmentation. This yields 2–4x higher throughput compared to naive HuggingFace Transformers inference. The vLLM documentation confirms that continuous batching and PagedAttention are the standard for high-load LLM services.
Typical numbers on A100 80GB for Llama-3 8B (bf16): 400–600 req/s, P50 latency 200–400ms, P99 latency 600–900ms at concurrency 64. For 70B on two A100 with tensor parallelism: 80–120 req/s, P99 latency 1.5–2.5s. AWQ or GPTQ quantization reduces memory consumption by 2x with quality loss within 1–3%.
Multi-Agent Systems
Agents are LLMs with access to tools: search, code execution, API calls, database interaction. Common patterns:
- ReAct (Reason + Act): the model reasons → chooses a tool → observes the result → reasons again. LangChain and LlamaIndex implement it out of the box.
- Multi-agent orchestration: multiple specialized agents with a coordinator on top. Example: coordinator → researcher (search + summarization) → coder (code generation and execution) → critic (verification). Tools: AutoGen (Microsoft), CrewAI, custom implementation on LangGraph.
In production, agent systems are non-deterministic. Essential: guardrails, step limits, logging of each step, human-in-the-loop for critical actions.
How We Work: Stages, Timeline, Deliverables
| Stage |
Duration |
What You Get |
| Audit and data collection |
1–2 weeks |
Eval dataset of 100+ examples, task formalization |
| Baseline (prompt + RAG) |
1–2 weeks |
Working prototype, quality metrics |
| Fine-tuning (if needed) |
2–4 weeks |
Trained model, LoRA weights, model card |
| Deployment and monitoring |
1–2 weeks |
vLLM server, Grafana + Prometheus |
| Documentation and training |
1 week |
API documentation, team training |
What Is Included
We deliver:
- Technical documentation (model card, configs, deployment instructions)
- Access to infrastructure (code repository, trained weights)
- 1 month of post-deployment support (consultations, bug fixes)
- Customer team training (2–3 sessions on system operation)
Timeline: basic RAG prototype — 1–2 weeks. Fine-tuning with customer data — 3–6 weeks (including data preparation). Production system with monitoring and retraining — 2–4 months. Cost is calculated individually based on data volume, model complexity, and infrastructure requirements.
We guarantee the quality of the final model with performance benchmarks and ongoing monitoring. Our engineers have hands‑on experience with dozens of production LLM systems.
Want to evaluate your project? Leave a request — we will prepare a preliminary summary within 1–2 business days. Or get a consultation on choosing the approach: RAG, fine-tuning, or hybrid — we will tell you what works best for you. Contact us to discuss your LLM development needs. Schedule a free consultation today.